Sync: local changes as of 2026-07-15T16:19:22+02:00

This commit is contained in:
Michał Kopeć 2026-07-15 16:19:25 +02:00
parent 6af7b388f8
commit aa819daabd
43 changed files with 4 additions and 6493 deletions

View file

@ -1,149 +0,0 @@
---
name: check-my-timesheet
description: >
Two behaviours in one skill:
(1) CHECK — show the current user's timesheet gaps for a given period. Compares
Tempo worklogs against the working calendar and flags missing or partial days.
(2) LOG — help the user log time to Tempo. User provides a description or issue key;
Claude searches Jira if needed, confirms the entry, then posts it.
TRIGGER check when user says: "check my timesheet", "timesheet gaps", "where am I missing
hours", "did I log all hours", "show my missing time".
TRIGGER log when user says: "log time", "add time entry", "log hours", "report hours",
"book time on", "add worklog".
---
# Check My Timesheet
Personal timesheet checker and time logger for the user identified by `JIRA_EMAIL`.
## Prerequisites
All configuration lives in the `.env` file in the project root:
| Variable | Purpose | Default |
|---|---|---|
| `TEMPO_API_TOKEN` | Read and write Tempo worklogs | — |
| `JIRA_BASE_URL` | Resolve account ID and search issues | — |
| `JIRA_EMAIL` | Identifies whose timesheet to read / who is logging | — |
| `JIRA_API_TOKEN` | Authenticate with Jira | — |
| `MY_TIMESHEET_COUNTRY` | Country code for public holiday calendar (`PL`, `GB`, `DE`…) | `PL` |
| `MY_TIMESHEET_HOURS_PER_DAY` | Expected working hours per day | `8` |
## Setup
Determine `SKILL_DIR` from the location of this file.
---
## Behaviour 1 — Check timesheet gaps
### Trigger phrases
"check my timesheet", "timesheet gaps", "where am I missing hours", "did I log all hours",
"show my time", "show my missing time"
### Workflow
**Step 1 — Ask for period.** Always ask — never assume a default:
> "Which period should I check? (e.g. this week, last week, June 2026, or a date range)"
Convert natural language to `YYYY-MM-DD:YYYY-MM-DD`:
- "this week" → Monday of current week to yesterday
- "last week" → previous MondayFriday
- "June 2026" → `2026-06-01:2026-06-30`
- "today" → today:today
**Step 2 — Run check:**
```bash
python3 {SKILL_DIR}/scripts/check.py \
--period {YYYY-MM-DD:YYYY-MM-DD} \
--output-dir ./timesheet-output
```
**Step 3 — Present results.** Read the output JSON and show:
1. Summary: `Logged X h / Y h expected (Z working days)`
2. If `missing_hours == 0`: "✓ Timesheet complete for this period."
3. Otherwise, a table of all working days — highlight missing and partial:
```
| Date | Logged | Status |
| 2026-06-23 | 8.0 h | ✓ OK |
| 2026-06-24 | 4.0 h | ⚠ Partial (4h missing) |
| 2026-06-25 | 0.0 h | ✗ Missing |
```
4. List only the problem days prominently if there are many OK days.
---
## Behaviour 2 — Log time
### Trigger phrases
"log time", "add time entry", "log hours", "report hours", "book time on", "add worklog",
"log X hours on", "log X hours for"
### Workflow
**Step 1 — Collect what is needed:**
From the user's message extract or ask for:
- **Issue** — Jira issue key (e.g. `IAA-42`) or a description to search by
- **Hours** — number of hours (decimals OK: 1.5 = 1h30m)
- **Date** — which day to log on (default: today if not mentioned)
- **Description** — optional note for the worklog
If issue key is not provided but a description is, run a search:
```bash
python3 {SKILL_DIR}/scripts/log.py \
--mode search \
--query "{description}"
```
Present the results (up to 5 issues) and ask the user to pick one or confirm.
If no results found, ask the user to provide the issue key directly.
**Step 2 — Preview (run without --confirmed).** Always run the preview step first:
```bash
python3 {SKILL_DIR}/scripts/log.py \
--mode log \
--issue {KEY} \
--hours {N} \
--date {YYYY-MM-DD} \
--description "{description}"
```
The script prints the pending worklog and exits without writing anything.
If hours exceed `MY_TIMESHEET_HOURS_PER_DAY` it also prints a `⚠ WARNING` line.
**Step 3 — Show preview to user and ask for confirmation.** Present the full script output
including any warnings. Wait for explicit yes/no. Do NOT proceed without explicit confirmation.
**Step 4 — Log (only after yes, with --confirmed):**
```bash
python3 {SKILL_DIR}/scripts/log.py \
--mode log \
--issue {KEY} \
--hours {N} \
--date {YYYY-MM-DD} \
--description "{description}" \
--confirmed
```
**Step 5 — Report result.** Show the worklog ID and a confirmation message.
Offer to check the timesheet again to verify the entry is reflected.
---
## Guardrails
- Never log time without explicit user confirmation.
- Always run the preview step first — the `--confirmed` flag must only be added after the user says yes.
- If the preview shows a `⚠ WARNING` (hours exceed daily limit), make sure the user acknowledges it before confirming.
- If `my_timesheet.md` is missing, stop and show the user the template.
- If `JIRA_EMAIL` is not set, stop and ask the user to set it in `.env`.
- If any script step fails, show the error and stop.
- Do not guess the issue key — always search or ask if unsure.

View file

@ -1,250 +0,0 @@
#!/usr/bin/env python3
"""Check the current user's timesheet for gaps against the working calendar."""
from __future__ import annotations
import argparse
import base64
import calendar
import json
import os
import re
import sys
import urllib.error
import urllib.request
from datetime import date, timedelta
from pathlib import Path
from typing import Any
# ── .env loading ──────────────────────────────────────────────────────────────
def load_dotenv(*paths: Path) -> Path | None:
for path in paths:
if not path.exists():
continue
for raw in path.read_text(encoding="utf-8").splitlines():
line = raw.strip()
if not line or line.startswith("#") or "=" not in line:
continue
key, _, value = line.partition("=")
key = key.strip()
value = value.strip().strip('"').strip("'")
if key and key not in os.environ:
os.environ[key] = value
return path
return None
# ── Config loading ────────────────────────────────────────────────────────────
def load_config() -> dict[str, Any]:
"""Read timesheet config from env vars MY_TIMESHEET_COUNTRY and MY_TIMESHEET_HOURS_PER_DAY."""
country = os.environ.get("MY_TIMESHEET_COUNTRY", "PL").strip().upper()
raw_hpd = os.environ.get("MY_TIMESHEET_HOURS_PER_DAY", "8").strip()
try:
hours_per_day = float(raw_hpd)
except ValueError:
print(f"Warning: invalid MY_TIMESHEET_HOURS_PER_DAY '{raw_hpd}' — using 8.", file=sys.stderr)
hours_per_day = 8.0
return {"country": country, "hours_per_day": hours_per_day}
# ── Period parsing ────────────────────────────────────────────────────────────
def parse_period(period_str: str) -> tuple[date, date]:
s = period_str.strip()
if ":" in s:
a, b = s.split(":", 1)
return date.fromisoformat(a.strip()), date.fromisoformat(b.strip())
raise ValueError(f"Invalid period '{period_str}'. Use YYYY-MM-DD:YYYY-MM-DD.")
# ── Public holidays ───────────────────────────────────────────────────────────
def get_holidays(country: str, years: set[int]) -> set[date]:
holidays: set[date] = set()
for year in sorted(years):
url = f"https://date.nager.at/api/v3/PublicHolidays/{year}/{country.upper()}"
req = urllib.request.Request(url, headers={"User-Agent": "check-my-timesheet/1.0"})
try:
with urllib.request.urlopen(req, timeout=10) as r:
holidays |= {date.fromisoformat(i["date"]) for i in json.loads(r.read())}
except Exception as exc:
print(f"Warning: could not fetch holidays for {country} {year}: {exc}", file=sys.stderr)
return holidays
def working_days(start: date, end: date, holidays: set[date]) -> list[date]:
result, cur = [], start
while cur <= end:
if cur.weekday() < 5 and cur not in holidays:
result.append(cur)
cur += timedelta(days=1)
return result
# ── Jira account lookup ───────────────────────────────────────────────────────
def resolve_account_id(email: str, jira_base_url: str,
jira_email: str, jira_token: str) -> str:
creds = base64.b64encode(f"{jira_email}:{jira_token}".encode()).decode()
headers = {"Authorization": f"Basic {creds}", "Accept": "application/json"}
url = f"{jira_base_url.rstrip('/')}/rest/api/3/user/search?query={email}&maxResults=5"
req = urllib.request.Request(url, headers=headers)
try:
with urllib.request.urlopen(req, timeout=10) as r:
users = json.loads(r.read())
except urllib.error.HTTPError as exc:
if exc.code == 401:
print("Error: Jira 401. Check JIRA_EMAIL and JIRA_API_TOKEN.", file=sys.stderr)
sys.exit(1)
raise
match = next((u for u in users if u.get("emailAddress", "").lower() == email.lower()), None)
if not match:
print(f"Error: Jira account not found for {email}.", file=sys.stderr)
sys.exit(1)
return match["accountId"]
# ── Tempo worklogs ────────────────────────────────────────────────────────────
def fetch_worklogs(account_id: str, start: date, end: date,
token: str, base_url: str = "https://api.tempo.io/4") -> list[dict]:
results: list[dict] = []
headers = {"Authorization": f"Bearer {token}", "Accept": "application/json"}
next_url: str | None = (
f"{base_url}/worklogs/user/{account_id}?from={start}&to={end}&limit=1000"
)
while next_url:
req = urllib.request.Request(next_url, headers=headers)
try:
with urllib.request.urlopen(req, timeout=30) as r:
data = json.loads(r.read())
except urllib.error.HTTPError as exc:
body = exc.read().decode("utf-8", errors="replace")
if exc.code == 401:
print("Error: Tempo 401. Check TEMPO_API_TOKEN.", file=sys.stderr)
sys.exit(1)
raise RuntimeError(f"Tempo API {exc.code}: {body[:200]}") from exc
results.extend(data.get("results", []))
next_url = data.get("metadata", {}).get("next")
return results
# ── Main ──────────────────────────────────────────────────────────────────────
def run(args: argparse.Namespace) -> int:
skill_dir = Path(__file__).parent.parent
loaded = load_dotenv(Path.cwd() / ".env", skill_dir / ".env")
if loaded:
print(f"Loaded credentials from {loaded}")
tempo_token = os.environ.get("TEMPO_API_TOKEN", "")
jira_base_url = os.environ.get("JIRA_BASE_URL", "")
jira_email = os.environ.get("JIRA_EMAIL", "")
jira_token = os.environ.get("JIRA_API_TOKEN", "")
missing = [n for n, v in [("TEMPO_API_TOKEN", tempo_token), ("JIRA_BASE_URL", jira_base_url),
("JIRA_EMAIL", jira_email), ("JIRA_API_TOKEN", jira_token)] if not v]
if missing:
print(f"Error: missing credentials: {', '.join(missing)}", file=sys.stderr)
return 1
cfg = load_config()
print(f"Config: country={cfg['country']}, hours_per_day={cfg['hours_per_day']}")
try:
start, end = parse_period(args.period)
except ValueError as exc:
print(f"Error: {exc}", file=sys.stderr)
return 1
print(f"Period: {start} {end}")
print(f"User: {jira_email}")
print("Resolving Jira account...")
account_id = resolve_account_id(jira_email, jira_base_url, jira_email, jira_token)
print(f" {jira_email}{account_id}")
print("Fetching public holidays...")
holidays = get_holidays(cfg["country"], {start.year, end.year})
wdays = working_days(start, end, holidays)
expected_hours = len(wdays) * cfg["hours_per_day"]
print(f" Working days: {len(wdays)} Expected: {expected_hours:.0f}h")
print("Fetching Tempo worklogs...")
worklogs = fetch_worklogs(account_id, start, end, tempo_token)
print(f" Worklogs: {len(worklogs)}")
# Aggregate hours per date
hours_by_date: dict[str, float] = {}
for wl in worklogs:
d = wl.get("startDate", "")
if d:
hours_by_date[d] = hours_by_date.get(d, 0.0) + wl.get("timeSpentSeconds", 0) / 3600
# Classify each working day
hpd = cfg["hours_per_day"]
days_out = []
for d in wdays:
logged = round(hours_by_date.get(d.isoformat(), 0.0), 2)
if logged >= hpd:
status = "ok"
entry = {"date": d.isoformat(), "logged_hours": logged, "status": "ok"}
elif logged > 0:
missing_h = round(hpd - logged, 2)
status = "partial"
entry = {"date": d.isoformat(), "logged_hours": logged,
"status": "partial", "missing": missing_h}
else:
status = "missing"
entry = {"date": d.isoformat(), "logged_hours": 0.0,
"status": "missing", "missing": hpd}
days_out.append(entry)
total_logged = round(sum(hours_by_date.get(d.isoformat(), 0.0) for d in wdays), 2)
missing_hours = round(max(0.0, expected_hours - total_logged), 2)
output = {
"period": {"start": str(start), "end": str(end)},
"config": cfg,
"working_days": len(wdays),
"expected_hours": expected_hours,
"total_logged_hours": total_logged,
"missing_hours": missing_hours,
"days": days_out,
}
output_dir = Path(args.output_dir)
output_dir.mkdir(parents=True, exist_ok=True)
out_path = output_dir / f"my_check_{start}_{end}.json"
out_path.write_text(json.dumps(output, indent=2, ensure_ascii=False), encoding="utf-8")
print(f"Wrote {out_path}")
# Quick summary
problem_days = [d for d in days_out if d["status"] != "ok"]
print(f"\nLogged: {total_logged:.1f}h / {expected_hours:.0f}h expected")
if not problem_days:
print("✓ Timesheet complete.")
else:
missing_days = [d for d in problem_days if d["status"] == "missing"]
partial_days = [d for d in problem_days if d["status"] == "partial"]
if missing_days:
print(f"✗ Missing ({len(missing_days)} day(s)): "
+ ", ".join(d["date"] for d in missing_days))
if partial_days:
print(f"⚠ Partial ({len(partial_days)} day(s)): "
+ ", ".join(f"{d['date']} ({d['logged_hours']}h)" for d in partial_days))
return 0
def build_parser() -> argparse.ArgumentParser:
p = argparse.ArgumentParser(description="Check timesheet gaps for the current user.")
p.add_argument("--period", required=True, help="YYYY-MM-DD:YYYY-MM-DD")
p.add_argument("--output-dir", required=True)
return p
if __name__ == "__main__":
sys.exit(run(build_parser().parse_args()))

View file

@ -1,280 +0,0 @@
#!/usr/bin/env python3
"""
Search Jira issues or log a time entry to Tempo.
Modes:
--mode search --query "text"
Search Jira for matching issues. Returns JSON list of matches.
--mode log --issue KEY --hours N --date YYYY-MM-DD [--description "..."]
Post a worklog to Tempo. Requires confirmation from the caller.
"""
from __future__ import annotations
import argparse
import base64
import json
import os
import re
import sys
import urllib.error
import urllib.request
from pathlib import Path
# ── .env loading ──────────────────────────────────────────────────────────────
def load_dotenv(*paths: Path) -> Path | None:
for path in paths:
if not path.exists():
continue
for raw in path.read_text(encoding="utf-8").splitlines():
line = raw.strip()
if not line or line.startswith("#") or "=" not in line:
continue
key, _, value = line.partition("=")
key = key.strip()
value = value.strip().strip('"').strip("'")
if key and key not in os.environ:
os.environ[key] = value
return path
return None
# ── Jira helpers ──────────────────────────────────────────────────────────────
def jira_headers(jira_email: str, jira_token: str) -> dict:
creds = base64.b64encode(f"{jira_email}:{jira_token}".encode()).decode()
return {"Authorization": f"Basic {creds}", "Accept": "application/json",
"Content-Type": "application/json"}
def resolve_account_id(email: str, jira_base_url: str,
jira_email: str, jira_token: str) -> str:
url = f"{jira_base_url.rstrip('/')}/rest/api/3/user/search?query={email}&maxResults=5"
req = urllib.request.Request(url, headers=jira_headers(jira_email, jira_token))
try:
with urllib.request.urlopen(req, timeout=10) as r:
users = json.loads(r.read())
except urllib.error.HTTPError as exc:
if exc.code == 401:
print("Error: Jira 401. Check JIRA_EMAIL and JIRA_API_TOKEN.", file=sys.stderr)
sys.exit(1)
raise
match = next((u for u in users if u.get("emailAddress", "").lower() == email.lower()), None)
if not match:
print(f"Error: Jira account not found for {email}.", file=sys.stderr)
sys.exit(1)
return match["accountId"]
def search_issues(query: str, jira_base_url: str,
jira_email: str, jira_token: str) -> list[dict]:
"""Search Jira for issues matching the query. Uses POST /search/jql (v3)."""
body = json.dumps({
"jql": f'text ~ "{query}" ORDER BY updated DESC',
"fields": ["key", "summary", "status", "assignee"],
"maxResults": 5,
}).encode()
url = f"{jira_base_url.rstrip('/')}/rest/api/3/search/jql"
req = urllib.request.Request(url, data=body,
headers=jira_headers(jira_email, jira_token),
method="POST")
try:
with urllib.request.urlopen(req, timeout=15) as r:
data = json.loads(r.read())
except urllib.error.HTTPError as exc:
body_txt = exc.read().decode("utf-8", errors="replace")
raise RuntimeError(f"Jira search failed ({exc.code}): {body_txt[:300]}") from exc
results = []
for issue in data.get("issues", []):
fields = issue.get("fields", {})
results.append({
"key": issue["key"],
"summary": fields.get("summary", ""),
"status": fields.get("status", {}).get("name", ""),
"assignee": (fields.get("assignee") or {}).get("displayName", "unassigned"),
})
return results
def get_issue(issue_key: str, jira_base_url: str,
jira_email: str, jira_token: str) -> dict:
"""Fetch id and summary of a Jira issue. Returns {"id": int, "summary": str}."""
url = f"{jira_base_url.rstrip('/')}/rest/api/3/issue/{issue_key}?fields=summary,id"
req = urllib.request.Request(url, headers=jira_headers(jira_email, jira_token))
try:
with urllib.request.urlopen(req, timeout=10) as r:
data = json.loads(r.read())
return {
"id": int(data.get("id", 0)),
"summary": data.get("fields", {}).get("summary", ""),
}
except Exception:
return {"id": 0, "summary": ""}
# ── Tempo log ─────────────────────────────────────────────────────────────────
def log_worklog(account_id: str, issue_id: int, hours: float,
log_date: str, description: str, tempo_token: str,
base_url: str = "https://api.tempo.io/4") -> dict:
"""POST a new worklog to Tempo. Returns the created worklog data."""
body = {
"issueId": issue_id,
"timeSpentSeconds": int(hours * 3600),
"startDate": log_date,
"startTime": "09:00:00",
"authorAccountId": account_id,
}
if description:
body["description"] = description
data = json.dumps(body).encode()
headers = {
"Authorization": f"Bearer {tempo_token}",
"Content-Type": "application/json",
"Accept": "application/json",
}
req = urllib.request.Request(f"{base_url}/worklogs", data=data,
headers=headers, method="POST")
try:
with urllib.request.urlopen(req, timeout=30) as r:
return json.loads(r.read())
except urllib.error.HTTPError as exc:
body_txt = exc.read().decode("utf-8", errors="replace")
if exc.code == 401:
print("Error: Tempo 401. Check TEMPO_API_TOKEN.", file=sys.stderr)
sys.exit(1)
raise RuntimeError(f"Tempo API {exc.code}: {body_txt[:300]}") from exc
# ── Main ──────────────────────────────────────────────────────────────────────
def run(args: argparse.Namespace) -> int:
skill_dir = Path(__file__).parent.parent
loaded = load_dotenv(Path.cwd() / ".env", skill_dir / ".env")
if loaded:
print(f"Loaded credentials from {loaded}")
tempo_token = os.environ.get("TEMPO_API_TOKEN", "")
jira_base_url = os.environ.get("JIRA_BASE_URL", "")
jira_email = os.environ.get("JIRA_EMAIL", "")
jira_token = os.environ.get("JIRA_API_TOKEN", "")
missing = [n for n, v in [("JIRA_BASE_URL", jira_base_url),
("JIRA_EMAIL", jira_email),
("JIRA_API_TOKEN", jira_token)] if not v]
if args.mode == "log" and not tempo_token:
missing.append("TEMPO_API_TOKEN")
if missing:
print(f"Error: missing credentials: {', '.join(missing)}", file=sys.stderr)
return 1
# ── Search mode ───────────────────────────────────────────────────────────
if args.mode == "search":
if not args.query:
print("Error: --query is required for search mode.", file=sys.stderr)
return 1
print(f"Searching Jira for: {args.query!r}")
try:
results = search_issues(args.query, jira_base_url, jira_email, jira_token)
except RuntimeError as exc:
print(f"Error: {exc}", file=sys.stderr)
return 1
print(json.dumps(results, indent=2, ensure_ascii=False))
if not results:
print("No issues found.", file=sys.stderr)
return 0
# ── Log mode ──────────────────────────────────────────────────────────────
if not args.issue:
print("Error: --issue is required for log mode.", file=sys.stderr)
return 1
if not args.hours or args.hours <= 0:
print("Error: --hours must be a positive number.", file=sys.stderr)
return 1
if not args.date:
print("Error: --date is required for log mode (YYYY-MM-DD).", file=sys.stderr)
return 1
issue_key = args.issue.strip().upper()
if not re.fullmatch(r"[A-Z][A-Z0-9]+-\d+", issue_key):
print(f"Error: '{issue_key}' is not a valid Jira issue key (expected e.g. IAA-42).",
file=sys.stderr)
return 1
# Read configured daily limit for over-hours warning
try:
hours_per_day = float(os.environ.get("MY_TIMESHEET_HOURS_PER_DAY", "8"))
except ValueError:
hours_per_day = 8.0
print(f"Resolving Jira account for {jira_email}...")
account_id = resolve_account_id(jira_email, jira_base_url, jira_email, jira_token)
# Fetch issue ID and summary (Tempo v4 requires numeric issueId, not issueKey)
issue_info = get_issue(issue_key, jira_base_url, jira_email, jira_token)
issue_id = issue_info["id"]
summary = issue_info["summary"]
if not issue_id:
print(f"Error: could not resolve numeric ID for {issue_key}.", file=sys.stderr)
return 1
# ── Guardrail: over-hours warning ─────────────────────────────────────────
if args.hours > hours_per_day:
print(
f"\n⚠ WARNING: {args.hours}h exceeds the configured daily limit "
f"of {hours_per_day}h (MY_TIMESHEET_HOURS_PER_DAY).",
file=sys.stderr,
)
# ── Preview (always shown) ────────────────────────────────────────────────
print(f"\nPending worklog:")
print(f" Issue: {issue_key}" + (f"{summary}" if summary else ""))
print(f" Hours: {args.hours}h")
print(f" Date: {args.date}")
if args.description:
print(f" Note: {args.description}")
# ── Guardrail: require explicit confirmation ───────────────────────────────
if not args.confirmed:
print(
"\n⏸ Not logged. Re-run with --confirmed to post this worklog to Tempo."
)
return 0
try:
result = log_worklog(
account_id=account_id,
issue_id=issue_id,
hours=args.hours,
log_date=args.date,
description=args.description or "",
tempo_token=tempo_token,
)
worklog_id = result.get("tempoWorklogId") or result.get("id", "?")
print(f"\n✓ Worklog logged. Tempo ID: {worklog_id}")
return 0
except RuntimeError as exc:
print(f"\nError: {exc}", file=sys.stderr)
return 1
def build_parser() -> argparse.ArgumentParser:
p = argparse.ArgumentParser(description="Search Jira or log time to Tempo.")
p.add_argument("--mode", required=True, choices=["search", "log"])
p.add_argument("--query", help="Search text (search mode)")
p.add_argument("--issue", help="Jira issue key, e.g. IAA-42 (log mode)")
p.add_argument("--hours", type=float, help="Hours to log, e.g. 2.5 (log mode)")
p.add_argument("--date", help="Date YYYY-MM-DD (log mode)")
p.add_argument("--description", default="", help="Optional worklog note (log mode)")
p.add_argument("--confirmed", action="store_true",
help="Must be passed to actually post the worklog. Without it the script "
"only shows a preview.")
return p
if __name__ == "__main__":
sys.exit(run(build_parser().parse_args()))

Binary file not shown.

View file

@ -1,141 +0,0 @@
---
name: clouddrift-docx
description: Convert any Markdown file into a Cloud Drift branded .docx or .pdf using pandoc and a custom reference template (fonts, colors, logo, footer pagination pulled from Cloud Drift's brand guide and case-study docx). Use when the user asks to "export as a Cloud Drift branded doc/pdf", "convert this markdown to a Cloud Drift docx", "make this look like our case study", or wants a client-ready styled document from a markdown file.
---
# Cloud Drift branded docx/pdf export
## Purpose
Turn any Markdown file into a document that looks like it came from Cloud
Drift's own brand: Open Sans / Open Sans Light typography, the Fire
Opal / White Coffee / Raisin Black color system, the "Cloud Drift." logo in
a running header, and page-numbered footers. Output can be `.docx`,
`.pdf`, or both.
The brand spec was extracted directly from `tmp/Cloud_Drift_CI_v1b 2.pdf`
(the CI/brand guide) and cross-checked against the real embedded fonts/
colors/logo inside `tmp/Branded [case study] ... .docx` — this isn't a
guess at "corporate-looking" styling, it replicates the actual brand:
- **Font:** Open Sans Light (body, title, H1), Open Sans regular/bold (H2/H3, table headers)
- **Colors:** Fire Opal `#E45249` (primary/accent — title, H2, hyperlinks, blockquote bar), White Coffee `#E8DCD0` (table shading/borders), Raisin Black `#252525` (body text, H1)
- **Logo:** "Cloud Drift." wordmark, placed in the page header on every page
- **Layout:** A4, 1" margins, footer with "Page X of Y"
## How this fits together
1. `assets/clouddrift-reference.docx` — the pandoc reference-doc. Pandoc
clones this file's styles (Normal, Title, Heading 1-6, Block Text,
Hyperlink, Table), page setup, and header/footer (including the logo)
into whatever it generates. The Open Sans / Open Sans Light fonts are
embedded inside it (unobfuscated TTF parts, same scheme the original
Cloud Drift case-study docx uses), so Word will render it correctly even
on a machine that doesn't have the fonts installed.
2. `assets/fonts/*.ttf` — the same 8 font files, for reinstalling the fonts
locally if needed (see below).
3. `assets/clouddrift-logo.png` — the extracted logo, already baked into
the reference doc's header; kept here for reuse elsewhere if needed.
4. `scripts/convert.py` — the conversion driver (see usage below).
5. `scripts/build_reference.py` + `scripts/embed_fonts.py` — the one-time
scripts that built `clouddrift-reference.docx` from pandoc's own default
reference doc. Only needed again if the brand changes (new colors, new
logo, different heading scheme) — see "Regenerating the template" below.
## Usage
```bash
python3 .claude/skills/clouddrift-docx/scripts/convert.py INPUT.md --format docx
python3 .claude/skills/clouddrift-docx/scripts/convert.py INPUT.md --format pdf
python3 .claude/skills/clouddrift-docx/scripts/convert.py INPUT.md --format both --output outputs/documents/my-doc
```
- `--format``docx` (default), `pdf`, or `both`.
- `--output` — path without extension; defaults to the input's own path/name.
- `--reference-doc` — override the template (rarely needed).
The `.docx` step is pure pandoc (`pandoc INPUT.md -o OUTPUT.docx --reference-doc=...`).
PDF export works by asking macOS **Pages** to open that generated `.docx`
and export it to PDF — this guarantees the PDF is pixel-identical to the
branded docx rather than a second, independently-maintained template. This
means:
- **macOS only.** There's no PDF engine pandoc can drive directly in this
environment, so this is the practical path rather than building a
parallel LaTeX/CSS template.
- **First run may need a permission grant.** macOS will prompt to let the
automating process control Pages (System Settings → Privacy & Security →
Automation). Approve it once.
- **If invoked through Claude Code's Bash tool**, the PDF step needs
`dangerouslyDisableSandbox: true` — sandboxed Bash can't send Apple
Events to GUI apps like Pages. The docx-only step does not need this.
- If Pages returns "Connection is invalid" on the very first call, it
usually means Pages hadn't finished launching yet — retry once.
**Why not `pandoc --pdf-engine=...` directly?** Tried this on 2026-07-14 —
installed `tectonic` (a self-contained LaTeX engine) specifically so pandoc
could produce PDF natively. It failed: tectonic fetches its TeX resource
bundle from `relay.fullyjustified.net` on first use, and that domain
resolves to `0.0.0.0` on this network (a DNS-level filter, not something to
route around). General internet access otherwise works fine — it's specific
to that host. Asked the user how to proceed; they chose to keep the
Pages-based PDF path rather than switch to a bigger `BasicTeX` install or
allowlist the domain. `tectonic` was uninstalled again. If this is
revisited later, either option is viable — see git/workload history for
2026-07-14 for the tradeoffs discussed.
## Font install (one-time, already done as of 2026-07-14)
Open Sans / Open Sans Light aren't system fonts on macOS by default. They've
already been installed to `~/Library/Fonts/` from `assets/fonts/` so Pages/
Word render them correctly instead of falling back to a serif substitute.
If this is ever run on a different machine, install them first:
```bash
cp .claude/skills/clouddrift-docx/assets/fonts/*.ttf ~/Library/Fonts/
```
(Open Sans is SIL Open Font License — freely redistributable. These exact
files came from the Cloud Drift case-study docx in `tmp/`.)
## Regenerating the template
Only needed if the brand changes. From a scratch directory:
```bash
pandoc -o pandoc-default-reference.docx --print-default-data-file reference.docx
cp <skill>/assets/fonts/*.ttf ./fonts/
cp <skill>/assets/clouddrift-logo.png ./clouddrift-logo.png
python3 <skill>/scripts/build_reference.py # edit colors/fonts/sizes at the top first if rebranding
python3 <skill>/scripts/embed_fonts.py clouddrift-reference.docx
cp clouddrift-reference.docx <skill>/assets/clouddrift-reference.docx
```
Then sanity-check visually: convert a test markdown file and render page 1
with `qlmanage -t -s 1600 -o <dir> file.pdf` (no poppler/pdftoppm needed) —
or install `poppler` (`brew install poppler`) for `pdftoppm` to check
arbitrary pages of a multi-page doc, which is what caught the page-break
issue below.
## Known fixes
- **2026-07-14 — empty page before large tables.** Pandoc's default
reference doc sets `keepNext`/`keepLines` on every Heading style
(standard Word behavior: never strand a heading alone at the bottom of a
page). Under Pages specifically, this backfired when a heading was
immediately followed by a large table: the heading got stranded alone on
a page and the *entire* table got pushed to the next page, leaving a
near-empty page in between. `build_reference.py` now explicitly strips
`keepNext`/`keepLines` from Heading 19 (see `disable_keep_with_next`) so
pagination flows naturally — worst case a heading ends up as the last
line on a page, which is a far smaller cosmetic cost than a blank page.
If a similar gap ever reappears with some other block type, check the
relevant style's `pPr` for `keepNext`/`keepLines`/`pageBreakBefore` first.
## Known limitations
- Bullet markers use the default (black) bullet glyph, not the Fire-Opal-red
bullet dot seen in the case study — pandoc generates its own numbering
definitions per document rather than inheriting the reference doc's, so
this isn't controllable through the reference-doc mechanism alone.
- PDF export is macOS/Pages-only; there is no cross-platform fallback in
this environment.

Binary file not shown.

Before

Width:  |  Height:  |  Size: 16 KiB

View file

@ -1,250 +0,0 @@
#!/usr/bin/env python3
"""Build the Cloud Drift pandoc reference.docx from pandoc's default reference doc."""
import copy
from docx import Document
from docx.shared import Pt, Inches, RGBColor, Emu
from docx.enum.text import WD_ALIGN_PARAGRAPH
from docx.oxml.ns import qn
from docx.oxml import OxmlElement
SRC = "pandoc-default-reference.docx"
OUT = "clouddrift-reference.docx"
LOGO = "clouddrift-logo.png"
FIRE_OPAL = RGBColor(0xE4, 0x52, 0x49)
WHITE_COFFEE = RGBColor(0xE8, 0xDC, 0xD0)
RAISIN = RGBColor(0x25, 0x25, 0x25)
MUTED = RGBColor(0x59, 0x59, 0x59)
LIGHT = "Open Sans Light"
REGULAR = "Open Sans"
def strip_theme_attrs(rpr_or_font_element):
"""Remove theme-based color/font attrs so explicit values always win."""
if rpr_or_font_element is None:
return
for tag in ("color",):
el = rpr_or_font_element.find(qn(f"w:{tag}"))
if el is not None:
for attr in ("themeColor", "themeTint", "themeShade"):
if el.get(qn(f"w:{attr}")) is not None:
del el.attrib[qn(f"w:{attr}")]
rfonts = rpr_or_font_element.find(qn("w:rFonts"))
if rfonts is not None:
for attr in ("asciiTheme", "hAnsiTheme", "eastAsiaTheme", "cstheme"):
if rfonts.get(qn(f"w:{attr}")) is not None:
del rfonts.attrib[qn(f"w:{attr}")]
def set_style(doc, name, font_name=None, size=None, color=None, bold=None, italic=None):
style = doc.styles[name]
f = style.font
if font_name:
f.name = font_name
rpr = style.element.get_or_add_rPr()
rfonts = rpr.find(qn("w:rFonts"))
if rfonts is None:
rfonts = OxmlElement("w:rFonts")
rpr.append(rfonts)
rfonts.set(qn("w:ascii"), font_name)
rfonts.set(qn("w:hAnsi"), font_name)
rfonts.set(qn("w:cs"), font_name)
rfonts.set(qn("w:eastAsia"), font_name)
if size:
f.size = Pt(size)
if color:
f.color.rgb = color
if bold is not None:
f.bold = bold
if italic is not None:
f.italic = italic
strip_theme_attrs(style.element.get_or_add_rPr())
def disable_keep_with_next(doc, names):
"""Remove keepNext/keepLines from these styles' pPr.
Pandoc's default reference doc sets keepNext+keepLines on every Heading
style (standard Word behavior: never orphan a heading alone at the
bottom of a page). But when the very next block is a large table, this
can backfire badly under some renderers (observed in Pages): the
heading gets stranded alone on a page and the entire table is pushed to
the following page, leaving a mostly-blank page in between. Dropping
keepNext/keepLines lets pagination flow naturally instead - worst case
a heading ends up as the last line on a page, which is a far smaller
cosmetic cost than a near-empty page.
"""
for name in names:
style = doc.styles[name]
pPr = style.element.get_or_add_pPr()
for tag in ("w:keepNext", "w:keepLines"):
el = pPr.find(qn(tag))
if el is not None:
pPr.remove(el)
def add_field(paragraph, field_code):
run = paragraph.add_run()
r = run._r
fld_begin = OxmlElement("w:fldChar")
fld_begin.set(qn("w:fldCharType"), "begin")
instr = OxmlElement("w:instrText")
instr.set(qn("xml:space"), "preserve")
instr.text = f" {field_code} "
fld_sep = OxmlElement("w:fldChar")
fld_sep.set(qn("w:fldCharType"), "separate")
fld_end = OxmlElement("w:fldChar")
fld_end.set(qn("w:fldCharType"), "end")
r.append(fld_begin)
r2 = paragraph.add_run()._r
r2.append(instr)
r3 = paragraph.add_run()._r
r3.append(fld_sep)
r4 = paragraph.add_run()._r
r4.append(fld_end)
def main():
doc = Document(SRC)
# --- Base body text ---
set_style(doc, "Normal", LIGHT, 11, RAISIN)
set_style(doc, "Body Text", LIGHT, 11, RAISIN)
set_style(doc, "Compact", LIGHT, 11, RAISIN)
set_style(doc, "First Paragraph", LIGHT, 11, RAISIN)
set_style(doc, "Default Paragraph Font", LIGHT, 11, RAISIN)
# --- Title / subtitle ---
set_style(doc, "Title", LIGHT, 30, FIRE_OPAL, bold=False)
doc.styles["Title"].paragraph_format.space_after = Pt(4)
set_style(doc, "Subtitle", LIGHT, 14, MUTED, bold=False, italic=False)
# --- Headings ---
set_style(doc, "Heading 1", LIGHT, 22, RAISIN, bold=False)
doc.styles["Heading 1"].paragraph_format.space_before = Pt(20)
doc.styles["Heading 1"].paragraph_format.space_after = Pt(8)
set_style(doc, "Heading 2", REGULAR, 16, FIRE_OPAL, bold=True)
doc.styles["Heading 2"].paragraph_format.space_before = Pt(16)
doc.styles["Heading 2"].paragraph_format.space_after = Pt(6)
set_style(doc, "Heading 3", REGULAR, 13, RAISIN, bold=True)
doc.styles["Heading 3"].paragraph_format.space_before = Pt(12)
for lvl, sz in ((4, 11.5), (5, 11), (6, 11)):
name = f"Heading {lvl}"
set_style(doc, name, REGULAR, sz, MUTED, bold=True, italic=(lvl == 6))
# Avoid huge empty-page gaps when a heading is immediately followed by
# a large table (see disable_keep_with_next docstring).
disable_keep_with_next(doc, [f"Heading {n}" for n in range(1, 10)])
# --- Quotes / block text ---
set_style(doc, "Block Text", LIGHT, 11, MUTED, italic=True)
bt_pPr = doc.styles["Block Text"].element.get_or_add_pPr()
pbdr = OxmlElement("w:pBdr")
left = OxmlElement("w:left")
left.set(qn("w:val"), "single")
left.set(qn("w:sz"), "18")
left.set(qn("w:space"), "8")
left.set(qn("w:color"), "E45249")
pbdr.append(left)
bt_pPr.append(pbdr)
# --- Hyperlinks ---
set_style(doc, "Hyperlink", LIGHT, None, FIRE_OPAL)
hl_rpr = doc.styles["Hyperlink"].element.get_or_add_rPr()
u = OxmlElement("w:u")
u.set(qn("w:val"), "single")
hl_rpr.append(u)
# --- Table: shaded header row, light borders ---
table_style = doc.styles["Table"]
tbl_pr = table_style.element.find(qn("w:tblPr"))
if tbl_pr is None:
tbl_pr = OxmlElement("w:tblPr")
table_style.element.append(tbl_pr)
borders = OxmlElement("w:tblBorders")
for edge in ("top", "left", "bottom", "right", "insideH", "insideV"):
el = OxmlElement(f"w:{edge}")
el.set(qn("w:val"), "single")
el.set(qn("w:sz"), "4")
el.set(qn("w:space"), "0")
el.set(qn("w:color"), "E8DCD0")
borders.append(el)
tbl_pr.append(borders)
style_pr = table_style.element.find(qn("w:tblStylePr"))
if style_pr is None:
style_pr = OxmlElement("w:tblStylePr")
style_pr.set(qn("w:type"), "firstRow")
table_style.element.append(style_pr)
tc_pr = style_pr.find(qn("w:tcPr"))
if tc_pr is None:
tc_pr = OxmlElement("w:tcPr")
style_pr.append(tc_pr)
shd = OxmlElement("w:shd")
shd.set(qn("w:val"), "clear")
shd.set(qn("w:color"), "auto")
shd.set(qn("w:fill"), "E8DCD0")
tc_pr.append(shd)
rpr_fr = style_pr.find(qn("w:rPr"))
if rpr_fr is None:
rpr_fr = OxmlElement("w:rPr")
style_pr.append(rpr_fr)
b_el = OxmlElement("w:b")
rpr_fr.append(b_el)
color_el = OxmlElement("w:color")
color_el.set(qn("w:val"), "252525")
rpr_fr.append(color_el)
# --- Verbatim / code ---
try:
set_style(doc, "Verbatim Char", None, 10, RAISIN)
except KeyError:
pass
# --- Page setup: A4, 1 inch margins ---
section = doc.sections[0]
section.page_width = Inches(8.27)
section.page_height = Inches(11.69)
section.left_margin = Inches(1)
section.right_margin = Inches(1)
section.top_margin = Inches(1)
section.bottom_margin = Inches(1)
section.header_distance = Inches(0.4)
section.footer_distance = Inches(0.4)
# --- Header: Cloud Drift logo ---
header = section.header
header.is_linked_to_previous = False
hp = header.paragraphs[0]
hp.text = ""
hp.alignment = WD_ALIGN_PARAGRAPH.LEFT
run = hp.add_run()
run.add_picture(LOGO, width=Inches(0.85))
# --- Footer: page number, right aligned, muted ---
footer = section.footer
footer.is_linked_to_previous = False
fp = footer.paragraphs[0]
fp.text = ""
fp.alignment = WD_ALIGN_PARAGRAPH.RIGHT
run = fp.add_run("Page ")
run.font.name = REGULAR
run.font.size = Pt(9)
run.font.color.rgb = MUTED
add_field(fp, "PAGE")
run2 = fp.add_run(" of ")
run2.font.name = REGULAR
run2.font.size = Pt(9)
run2.font.color.rgb = MUTED
add_field(fp, "NUMPAGES")
doc.save(OUT)
print("Saved", OUT)
if __name__ == "__main__":
main()

View file

@ -1,112 +0,0 @@
#!/usr/bin/env python3
"""Convert a Markdown file into a Cloud Drift branded .docx and/or .pdf.
Usage:
python3 convert.py INPUT.md [--format docx|pdf|both] [--output PATH]
[--reference-doc PATH]
docx is produced by pandoc using the bundled Cloud Drift reference.docx
(fonts, colors, logo header, footer pagination). pdf is produced by asking
macOS Pages to open that docx and export it to PDF, so the PDF is pixel-for-
pixel the same branded layout, not a second independent template.
"""
import argparse
import shutil
import subprocess
import sys
from pathlib import Path
SKILL_DIR = Path(__file__).resolve().parent.parent
REFERENCE_DOC = SKILL_DIR / "assets" / "clouddrift-reference.docx"
def convert_to_docx(input_md: Path, output_docx: Path, reference_doc: Path):
if shutil.which("pandoc") is None:
sys.exit("pandoc is not installed or not on PATH.")
cmd = [
"pandoc",
str(input_md),
"-o",
str(output_docx),
f"--reference-doc={reference_doc}",
"--standalone",
]
subprocess.run(cmd, check=True)
print(f"Wrote {output_docx}")
def convert_docx_to_pdf(input_docx: Path, output_pdf: Path):
"""Use macOS Pages (via AppleScript) to export the docx to PDF, preserving
the exact branded layout produced by the reference.docx styles."""
if sys.platform != "darwin":
sys.exit("PDF export currently requires macOS (uses Pages via AppleScript).")
script = f'''
try
tell application "Pages"
set theDoc to open POSIX file "{input_docx.resolve()}"
delay 2
export theDoc to POSIX file "{output_pdf.resolve()}" as PDF
close theDoc saving no
end tell
return "SUCCESS"
on error errMsg number errNum
return "ERROR: " & errMsg & " (" & errNum & ")"
end try
'''
result = subprocess.run(["osascript", "-e", script], capture_output=True, text=True)
out = result.stdout.strip()
if out != "SUCCESS":
sys.exit(
"Pages PDF export failed: "
f"{out or result.stderr.strip()}\n"
"If this is the first run, macOS may need you to grant automation "
"permission for controlling Pages (System Settings > Privacy & "
"Security > Automation), or Pages may need a moment after "
"launching — try again."
)
print(f"Wrote {output_pdf}")
def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("input", type=Path, help="Input Markdown file")
parser.add_argument(
"--format", choices=["docx", "pdf", "both"], default="docx",
help="Output format (default: docx)",
)
parser.add_argument(
"--output", type=Path, default=None,
help="Output path (without extension needed for --format both). "
"Defaults to the input filename next to the input file.",
)
parser.add_argument(
"--reference-doc", type=Path, default=REFERENCE_DOC,
help="Override the Cloud Drift reference.docx template",
)
args = parser.parse_args()
if not args.input.exists():
sys.exit(f"Input file not found: {args.input}")
if not args.reference_doc.exists():
sys.exit(f"Reference doc not found: {args.reference_doc}")
stem = args.output if args.output else args.input.with_suffix("")
docx_path = stem.with_suffix(".docx")
pdf_path = stem.with_suffix(".pdf")
if args.format in ("docx", "both"):
convert_to_docx(args.input, docx_path, args.reference_doc)
if args.format in ("pdf", "both"):
if not docx_path.exists():
convert_to_docx(args.input, docx_path, args.reference_doc)
convert_docx_to_pdf(docx_path, pdf_path)
if args.format == "pdf" and docx_path.exists() and args.output is None:
# pdf-only was requested and we only made the docx as an
# intermediate step; clean it up unless the caller named an
# explicit --output (in which case leave both, they may want it).
docx_path.unlink()
if __name__ == "__main__":
main()

View file

@ -1,127 +0,0 @@
#!/usr/bin/env python3
"""Embed Open Sans / Open Sans Light TTFs into a docx so it renders correctly
even on machines that don't have the fonts installed. Mirrors the (unobfuscated)
embedding scheme found in the Cloud Drift case-study docx: fontKey all-zeros,
plain .ttf parts referenced directly."""
import shutil
import zipfile
import re
import sys
DOCX = sys.argv[1] if len(sys.argv) > 1 else "clouddrift-reference.docx"
FONT_DIR = "fonts"
FONTS = {
"Open Sans Light": {
"regular": "OpenSansLight-regular.ttf",
"bold": "OpenSansLight-bold.ttf",
"italic": "OpenSansLight-italic.ttf",
"boldItalic": "OpenSansLight-boldItalic.ttf",
},
"Open Sans": {
"regular": "OpenSans-regular.ttf",
"bold": "OpenSans-bold.ttf",
"italic": "OpenSans-italic.ttf",
"boldItalic": "OpenSans-boldItalic.ttf",
},
}
NS_R = "http://schemas.openxmlformats.org/officeDocument/2006/relationships"
FONT_REL_TYPE = f"{NS_R}/font".replace(NS_R, "http://schemas.openxmlformats.org/officeDocument/2006/relationships") + "/font"
FONT_REL_TYPE = "http://schemas.openxmlformats.org/officeDocument/2006/relationships/font"
def build_font_table_and_rels():
rid = 1
font_entries = []
rels = []
file_map = {}
for name, variants in FONTS.items():
embeds = []
for kind, fname in variants.items():
tag = {
"regular": "w:embedRegular",
"bold": "w:embedBold",
"italic": "w:embedItalic",
"boldItalic": "w:embedBoldItalic",
}[kind]
rId = f"rIdFont{rid}"
embeds.append(
f'<{tag} r:id="{rId}" w:fontKey="{{00000000-0000-0000-0000-000000000000}}" w:subsetted="0"/>'
)
rels.append(
f'<Relationship Id="{rId}" Type="{FONT_REL_TYPE}" Target="fonts/{fname}"/>'
)
file_map[rId] = fname
rid += 1
font_entries.append(f'<w:font w:name="{name}">{"".join(embeds)}</w:font>')
font_table_xml = (
'<?xml version="1.0" encoding="UTF-8" standalone="yes"?>'
'<w:fonts xmlns:r="http://schemas.openxmlformats.org/officeDocument/2006/relationships" '
'xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main">'
+ "".join(font_entries)
+ "</w:fonts>"
)
rels_xml = (
'<?xml version="1.0" encoding="UTF-8" standalone="yes"?>'
'<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
+ "".join(rels)
+ "</Relationships>"
)
return font_table_xml, rels_xml, file_map
def patch_settings(xml_text):
if "embedTrueTypeFonts" in xml_text:
return xml_text
return re.sub(
r"(<w:settings[^>]*>)",
r'\1<w:embedTrueTypeFonts w:val="1"/>',
xml_text,
count=1,
)
def patch_content_types(xml_text):
if 'Extension="ttf"' in xml_text:
return xml_text
return xml_text.replace(
"<Types ",
"<Types ",
).replace(
'<Default Extension="rels"',
'<Default Extension="ttf" ContentType="application/x-font-ttf"/><Default Extension="rels"',
)
def main():
font_table_xml, rels_xml, file_map = build_font_table_and_rels()
tmp = DOCX + ".tmp"
with zipfile.ZipFile(DOCX, "r") as zin:
names = zin.namelist()
with zipfile.ZipFile(tmp, "w", zipfile.ZIP_DEFLATED) as zout:
for item in zin.infolist():
data = zin.read(item.filename)
if item.filename == "word/fontTable.xml":
data = font_table_xml.encode("utf-8")
elif item.filename == "word/settings.xml":
data = patch_settings(data.decode("utf-8")).encode("utf-8")
elif item.filename == "[Content_Types].xml":
data = patch_content_types(data.decode("utf-8")).encode("utf-8")
zout.writestr(item, data)
if "word/_rels/fontTable.xml.rels" not in names:
zout.writestr("word/_rels/fontTable.xml.rels", rels_xml)
for rId, fname in file_map.items():
with open(f"{FONT_DIR}/{fname}", "rb") as f:
zout.writestr(f"word/fonts/{fname}", f.read())
shutil.move(tmp, DOCX)
print("Embedded fonts into", DOCX)
if __name__ == "__main__":
main()

View file

@ -1,84 +0,0 @@
---
name: extract-transcript
description: Extract the meeting transcript from a saved Fireflies.ai (or similar) .mhtml page capture into a Markdown file. Use when the user asks to "extract the transcript", "convert this mhtml to a transcript", "get the transcript out of this mhtml file", or provides a `.mhtml` meeting recording file and wants the transcript pulled out. Takes one filename argument.
---
# Extract transcript skill
## Purpose
Fireflies.ai (and similar tools) render meeting transcripts client-side,
behind login - `WebFetch` and plain `curl`/headless-browser scraping only
see the app shell, never the transcript. The reliable workaround is for the
user to save the fully-rendered meeting page as a browser `.mhtml` page
capture (File > Save Page As > Webpage, Single File, or the browser's
"Save as MHTML" option). This skill turns that `.mhtml` capture into a
clean Markdown transcript, deterministically - no model reasoning needed.
## How to run this skill
Run the bundled script with Bash, passing the `.mhtml` file path as the
only argument:
```bash
python3 "<skill-dir>/scripts/extract_transcript.py" "<path/to/file.mhtml>"
```
Resolve `<skill-dir>` to this skill's own directory (the directory
containing this SKILL.md) and `<path/to/file.mhtml>` to the file the user
gave you - accept it whether they pasted an absolute path, a relative
path, or just referenced an open/attached file.
The script:
- Parses the `.mhtml` as a MIME multipart message (Python's `email`
module) and finds the rendered HTML part(s).
- Locates the transcript panel (`id` ending in `content-transcript`, or
falls back to the largest `ScrollArea-styled__Root` element) and walks
its structured DOM - per-paragraph speaker name, timestamp, and sentence
text - rather than flattening all text, which would lose speaker
boundaries.
- Extracts the page `<title>` and a best-effort meeting date/time from the
page text.
- Writes the output next to the input file, same directory and basename,
with a `.md` extension (e.g. `Team sync.mhtml` -> `Team sync.md`),
overwriting any existing file at that path.
If `beautifulsoup4` isn't installed, the script exits with the exact
`pip3 install beautifulsoup4` command to run - run it, then retry.
## After running
Report the output path and the number of transcript lines extracted (the
script prints both). Do not summarize or otherwise process the transcript
unless the user separately asks for that (e.g. ingesting it into the
wiki) - this skill's job ends at producing the `.md` file.
The script also detects and reports two common capture mistakes rather
than silently producing bad output:
- **Wrong tab active:** if the page was saved while a tab other than
Transcript was open (commonly Notes), Fireflies never mounted the
transcript panel into the DOM at all. The script detects this and tells
the user to reopen the meeting, click Transcript, and re-save.
- **Incomplete scroll:** Fireflies virtualizes the transcript list, so if
the user didn't scroll all the way through it before saving, only the
visible portion is captured and the rest is silently missing (not a
quiet stretch of the meeting). The script still writes the `.md` file in
this case but prints a warning to stderr for any gap of 90+ seconds
between consecutive lines, with the exact timestamp range of each gap,
and the same "scroll to the end, then re-save" guidance. Pass this
warning on to the user rather than treating a successful "Wrote N lines"
message as automatically complete.
## Limitations
- Built against Fireflies.ai's current page structure (styled-components
class names change on redeploys, so a Fireflies UI change could break
the selectors - if extraction fails, inspect the `.mhtml`'s HTML part
for the new transcript container structure and update
`scripts/extract_transcript.py` accordingly).
- Only captures speakers and sentences visible in the saved page. If the
transcript panel wasn't fully scrolled/loaded before saving, only the
loaded portion will be present.
- Attendee list is inferred purely from who has transcript lines - silent
attendees who never spoke won't appear.

View file

@ -1,233 +0,0 @@
#!/usr/bin/env python3
"""Extract a Fireflies.ai (or similar) meeting transcript from a saved .mhtml page.
Usage:
python3 extract_transcript.py <path/to/page.mhtml>
Writes <path/to/page>.md next to the input file (same directory, same
basename, .md extension), containing the meeting title/date if found and
the speaker-by-speaker transcript with timestamps.
"""
import email
import re
import sys
from pathlib import Path
try:
from bs4 import BeautifulSoup
except ImportError:
sys.exit(
"Missing dependency 'beautifulsoup4'. Install it with:\n"
" pip3 install beautifulsoup4"
)
MONTHS = "Jan|Feb|Mar|Apr|May|Jun|Jul|Aug|Sep|Oct|Nov|Dec"
DATE_RE = re.compile(rf"({MONTHS})[a-z]* \d{{1,2}},? \d{{4}}(?:, \d{{1,2}}:\d{{2}} ?[AP]M)?")
# Fireflies virtualizes the transcript list - if the user didn't scroll
# through the whole thing before saving the page, only the visible portion
# ends up in the DOM/mhtml, leaving silent gaps. A gap this long between
# consecutive timestamps is a strong signal that's what happened.
GAP_WARNING_SECONDS = 90
def parse_timestamp(ts):
if not ts:
return None
parts = ts.split(":")
try:
parts = [int(p) for p in parts]
except ValueError:
return None
seconds = 0
for p in parts:
seconds = seconds * 60 + p
return seconds
def find_large_gaps(lines):
gaps = []
prev_ts = None
prev_idx = None
for idx, (ts, _, _) in enumerate(lines):
secs = parse_timestamp(ts)
if secs is None:
continue
if prev_ts is not None and secs - prev_ts >= GAP_WARNING_SECONDS:
gaps.append((prev_idx, idx, secs - prev_ts))
prev_ts, prev_idx = secs, idx
return gaps
def load_html_parts(mhtml_path):
with open(mhtml_path, "rb") as f:
msg = email.message_from_binary_file(f)
for part in msg.walk():
if part.get_content_type() != "text/html":
continue
payload = part.get_payload(decode=True)
if not payload:
continue
charset = part.get_content_charset() or "utf-8"
yield payload.decode(charset, errors="replace")
def find_transcript_container(html):
soup = BeautifulSoup(html, "html.parser")
container = soup.find(id=lambda i: i and i.endswith("content-transcript"))
if container is not None:
return soup, container
# Fallback: some captures use the ScrollArea class directly without the
# radix id being present in this particular MIME part.
candidates = soup.find_all(
class_=lambda c: c and any("ScrollArea-styled__Root" in x for x in (c if isinstance(c, list) else [c]))
)
if candidates:
# The transcript panel is reliably the largest ScrollArea on the page.
best = max(candidates, key=lambda el: len(el.get_text()))
if len(best.get_text(strip=True)) > 0:
return soup, best
return soup, None
def extract_paragraphs(container):
paragraphs = container.find_all(
"div", id=lambda i: i and i.startswith("transcript-paragraph-")
)
lines = []
last_name = None
for p in paragraphs:
name_span = p.find("span", class_="name")
name = name_span.get_text(strip=True) if name_span else (last_name or "Unknown speaker")
last_name = name
ts_span = p.find("span", attrs={"text-decoration": "underline"})
ts = ts_span.get_text(strip=True) if ts_span else ""
content_div = p.find("div", class_=lambda c: c and "ContentPost-styled__Content" in c)
text = content_div.get_text(" ", strip=True) if content_div else p.get_text(" ", strip=True)
if not text:
continue
lines.append((ts, name, text))
return lines
def find_inactive_transcript_tab(html_parts):
"""Detect the common failure case: the page was saved with a different
tab (usually Notes) active, so Fireflies never mounted the transcript
panel into the DOM at all - there's nothing to extract, not a selector
mismatch."""
for html in html_parts:
soup = BeautifulSoup(html, "html.parser")
tab = soup.find(
attrs={"role": "tab"},
id=lambda i: i and i.endswith("trigger-transcript"),
)
if tab is not None:
active_tab = soup.find(attrs={"role": "tab", "data-state": "active"})
active_name = active_tab.get_text(strip=True) if active_tab else "another tab"
return tab.get("data-state") != "active", active_name
return False, None
def extract_title_and_date(html_parts):
title = None
date = None
for html in html_parts:
soup = BeautifulSoup(html, "html.parser")
if title is None and soup.title and soup.title.get_text(strip=True):
title = soup.title.get_text(strip=True)
if date is None:
m = DATE_RE.search(soup.get_text(" ", strip=True))
if m:
date = m.group(0)
if title and date:
break
return title, date
def main():
if len(sys.argv) != 2:
sys.exit(f"Usage: {sys.argv[0]} <path/to/page.mhtml>")
src = Path(sys.argv[1]).expanduser()
if not src.is_file():
sys.exit(f"File not found: {src}")
html_parts = list(load_html_parts(src))
if not html_parts:
sys.exit("No text/html part found in this .mhtml file - is it a valid MIME HTML capture?")
lines = []
for html in html_parts:
_, container = find_transcript_container(html)
if container is None:
continue
lines = extract_paragraphs(container)
if lines:
break
if not lines:
tab_inactive, active_name = find_inactive_transcript_tab(html_parts)
if tab_inactive:
sys.exit(
f"No transcript found - this page was saved with the "
f"'{active_name}' tab open, not 'Transcript'. Fireflies only "
f"renders the active tab's content into the page, so the "
f"transcript panel isn't in this file at all. Reopen the "
f"meeting, click the 'Transcript' tab, wait for it to load "
f"and scroll to the end (so the full transcript renders), "
f"then re-save the page and try again."
)
sys.exit(
"Could not find a transcript panel in this file. This script targets "
"Fireflies.ai-style pages (a ScrollArea containing "
"'#*-content-transcript'). If the page structure differs, the "
"selectors in extract_transcript.py need updating."
)
title, date = extract_title_and_date(html_parts)
speakers = sorted(set(name for _, name, _ in lines))
out_path = src.with_suffix(".md")
body = []
body.append(f"# {title or src.stem}")
body.append("")
if date:
body.append(f"- **Date:** {date}")
body.append(f"- **Attendees (speakers detected):** {', '.join(speakers)}")
body.append(f"- **Source:** `{src.name}`")
body.append("")
body.append("## Transcript")
body.append("")
# Blank line between entries (not a single "\n") so each timestamp
# starts its own paragraph in Markdown preview - a lone newline is a
# soft break that most renderers collapse into one running paragraph.
entries = []
for ts, name, text in lines:
prefix = f"[{ts}] " if ts else ""
entries.append(f"{prefix}{name}: {text}")
body.append("\n\n".join(entries))
out_path.write_text("\n".join(body) + "\n", encoding="utf-8")
print(f"Wrote {len(lines)} transcript lines to {out_path}")
gaps = find_large_gaps(lines)
if gaps:
print(
f"WARNING: {len(gaps)} gap(s) of {GAP_WARNING_SECONDS}s or more "
f"between consecutive lines - Fireflies virtualizes the transcript "
f"list, so this usually means the page was saved before scrolling "
f"through the whole transcript, and content in between is simply "
f"missing (not silence). Re-open the meeting, scroll the "
f"Transcript tab all the way to the end first, then re-save and "
f"re-run this script:",
file=sys.stderr,
)
for start_idx, end_idx, gap_secs in gaps:
start_ts = lines[start_idx][0]
end_ts = lines[end_idx][0]
print(f" - {start_ts} -> {end_ts} ({gap_secs}s gap)", file=sys.stderr)
if __name__ == "__main__":
main()

View file

@ -1,196 +0,0 @@
---
name: fireflies-relanguage
description: Fix a wrong-language Fireflies transcript on a meeting you don't own (can't hit Reprocess yourself) by re-uploading its audio under the correct language via the Fireflies API, recovering real speaker names into the result with transcript-speaker-fill, then deleting the duplicate meeting. Use when the user has a link to a Fireflies recording they aren't the host/owner of, the transcript is in the wrong language, and they want a corrected transcript without waiting for the owner to act or leaving a duplicate meeting behind. Trigger phrases: "I'm not the owner of this meeting", "can't reprocess this transcript", "fix this Fireflies transcript without the owner", "re-transcribe a shared meeting".
---
# Fireflies relanguage skill
## Purpose
Fireflies' own "Update Language + Reprocess" fix only works for the meeting's
host/owner (see `guide.fireflies.ai` docs). If you were just sent a link to
someone else's recording and it came out in the wrong language, you have no
UI path to fix it and can't wait for the owner. This skill routes around
that: pull the meeting's audio via the API, re-upload it as a **new** meeting
under your own account with the correct language forced, recover the
original's real speaker names into the new, correctly-transcribed text using
the [[transcript-speaker-fill]](../transcript-speaker-fill/SKILL.md) skill,
then delete the duplicate meeting the re-upload created.
## Before you start: real constraints, not hypothetical ones
**This may simply not be possible on your account.** `audio_url`/`video_url`
on a Fireflies transcript require a Pro+ plan, and it is undocumented whether
they populate at all for a meeting merely *shared* with you rather than
owned by you. [Michał's own wiki notes](/wiki/entities/fireflies-transcript-handling.md)
record hitting exactly this wall before (no Pro seat → transcript download
tooling broke). Step 1 below (`fetch`) is also the diagnostic: if it prints
the "audio_url is null" warning, stop here — this workflow cannot proceed
without one of the workarounds it suggests (a Pro+ seat's API key, the actual
owner reprocessing it, or you manually downloading the audio via the web UI
and hosting it at some public HTTPS URL yourself).
**Every re-upload consumes real transcription quota/minutes** on the account
whose API key you use, same as any other Fireflies upload — this isn't free
just because it's automated. **Deleting a transcript is irreversible.** Don't
run this on a meeting you actually might need to keep two copies of, and
don't skip the confirmation gates below.
Also worth knowing up front: the Fireflies API has thin rate limits (Free:
50 requests/day, Pro: 500/day, Business/Enterprise: 60/min) and
`deleteTranscript` specifically is capped at 10/min. This workflow only
needs a handful of calls plus some polling, but don't loop the `wait` step
aggressively.
## Requirements
- `FIREFLIES_API_KEY` set in the environment. Get one from the Fireflies web
app under **Settings > API**. The script exits with this exact instruction
if it's missing — don't hardcode the key anywhere.
- The `transcript-speaker-fill` skill installed alongside this one (it is,
in this repo) — this skill hands off the actual name-recovery merge to it
rather than duplicating that logic.
## Workflow
Resolve `<skill-dir>` to this skill's own directory throughout.
### 1. Fetch the original (broken-language) transcript
```bash
python3 "<skill-dir>/scripts/fireflies_client.py" fetch --ref "<link-or-id-the-user-gave-you>" --out original.md
```
Accepts a bare transcript id or a full share link like
`https://app.fireflies.ai/view/Some-Title::abcDEF123`. Writes `original.md`
(real speaker names, garbled text — the shape `transcript-speaker-fill`
expects for its "broken" input) and `original.md.meta.json` (id, title,
`audio_url`, etc., needed by later steps).
**Check the printed output for the audio_url warning before continuing.** If
it's null, stop and follow the guidance it prints instead of proceeding to
step 2 — do not attempt step 2 anyway "just to see."
### 2. Confirm the target language with the user
Don't guess the correct language from context alone unless it's unambiguous
(e.g. the user already told you). Ask if unclear. Use the language code
Fireflies expects — check
[Learn about Fireflies supported languages](https://guide.fireflies.ai/articles/2973706448-learn-about-fireflies-supported-languages)
for the exact code if you're not sure it matches (e.g. `en`, `pl`, `es`).
### 3. Re-upload the audio under the correct language
```bash
python3 "<skill-dir>/scripts/fireflies_client.py" upload --meta original.md.meta.json --language <code>
```
This calls `uploadAudio` with `custom_language` set, tagging the new
meeting's title with `[relang:<original-id>]` so it can be found
unambiguously afterward (override with `--title` if you want a cleaner
name, but then pass that exact same string to step 4). Note the exact title
printed — you need it verbatim for the next step.
### 4. Wait for the re-upload to finish processing
```bash
python3 "<skill-dir>/scripts/fireflies_client.py" wait --title "<exact title from step 3>"
```
Polls every 30s (default) up to 15 minutes (default) for a transcript with
that exact title to appear with content. Long recordings can take longer —
if it times out, just re-run `wait` again rather than assuming failure.
Prints the new transcript's `id` once ready.
### 5. Fetch the new (correct-language) transcript
```bash
python3 "<skill-dir>/scripts/fireflies_client.py" fetch --ref "<id from step 4>" --out new.md
```
This one should come back with correct text but generic `Speaker 1`,
`Speaker 2`, ... labels (no calendar/roster context on a bare re-upload) —
exactly the gap `transcript-speaker-fill` closes.
### 6. Recover real speaker names
Hand off to the `transcript-speaker-fill` skill exactly per its own
SKILL.md, using `original.md` and `new.md` as the two inputs — dry run
first, present the resolution report, get the user's confirmation/manual
overrides, then `--apply`. Do not skip its confirmation gate just because
you're mid-pipeline; it's exactly as load-bearing here as when invoked
standalone.
Also pass `--recording-url "https://app.fireflies.ai/view/<id from step 4>"`
on every `fill_speakers.py` call in this handoff (both the dry-run and the
`--apply` run) — the bare-id form of the URL works without needing to
slug-encode the title, and gives the user a direct link to a recording they
definitely have full access to (they own the re-uploaded meeting) right
next to every label's example lines, for the spot-check
`transcript-speaker-fill` asks them to do before confirming. Use the new
transcript's id, not the original's — do this *before* step 8's cleanup,
while that meeting (and its recording) still exists.
### 7. Ask where to save the corrected transcript
Before touching the duplicate meeting, ask the user where the final
`--apply`'d output (from step 6) should end up — don't default to leaving
it in a scratch/temp location without asking. Common answers: a specific
path they name, this project's own `tmp/` (if this skill is being run from
inside a Cascade Knowledge Base repo like this one — gitignored, agent-
managed), or `raw/inbox/` if they want it ingested into a wiki afterward.
Move (don't copy) the file there once they've told you.
### 8. Delete the duplicate meeting — only after explicit confirmation
Once the user has confirmed the merged transcript (from step 6) looks right
and it's been saved wherever they wanted (step 7), delete the meeting the
re-upload created (**not** the original — you don't own that one anyway,
and couldn't delete it if you tried):
```bash
python3 "<skill-dir>/scripts/fireflies_client.py" delete --id <id from step 4> --yes
```
The script refuses to run without `--yes`. Never pass `--yes` without the
user having explicitly confirmed they're ready — this is irreversible and
removes a real meeting from their Fireflies account. Show them the printed
`{id, title, date, duration}` of what was deleted as final confirmation.
## Failure modes and what they mean
- **`fetch` on the original prints the audio_url warning** — see "Before you
start" above. This is the expected failure mode when the API key's
account isn't Pro+, or when a merely-shared meeting doesn't expose audio
via the API. Not a bug to work around silently.
- **`upload` mutation succeeds (`success: true`) but `wait` never finds it**
— the title match is exact-string, so a `--title` override that doesn't
exactly match what you pass to `wait` will never resolve; double check
you used the identical string in both commands.
- **`wait` times out** — normal for long recordings. Re-run it; don't
assume the upload failed.
- **Rate limit errors (`too_many_requests`)** — the error message includes
`retryAfter`; wait that long before retrying, especially on a Free-tier
key (50 requests/day total).
- **`delete` fails with `require_elevated_privilege`** — you're trying to
delete a transcript you don't own (e.g. you accidentally passed the
*original* id instead of the new one from step 4). Only the re-uploaded
meeting is yours to delete.
## Edge cases
- **Multiple speakers with a name-recovery gap** (new transcript has more
distinct `Speaker N` labels than the original has real names) — this is
`transcript-speaker-fill`'s "roster gap" warning, not something this
skill's own steps can fix; it means someone's voice wasn't distinctly
captured with a real name in the original either.
- **Original meeting has `video_url` but not `audio_url`**`upload` only
accepts `--audio-url`; there's no video re-upload path here. If you truly
only have video access, extract audio from it yourself first and host
that as a public URL, then pass it via `--audio-url`.
- **The user wants the corrected transcript ingested into this repo's
wiki**, not just saved as a file — step 7 already asks where to save it;
if the answer is "ingest it," save it to `raw/inbox/` there, then run the
normal "Ingest" workflow from the root `CLAUDE.md` on it after step 8's
cleanup (the duplicate meeting is Fireflies-side bookkeeping, unrelated
to whether the wiki ingest has happened yet).

View file

@ -1,337 +0,0 @@
#!/usr/bin/env python3
"""Minimal Fireflies.ai GraphQL client for the fireflies-relanguage skill.
Handles the four network steps of the workflow: pull a transcript you can
view (owned or shared) into the Markdown shape transcript-speaker-fill
expects, kick off a re-upload of its audio under the correct language,
poll until that re-upload shows up as a finished transcript, and delete
the resulting duplicate meeting once you're done with it.
No third-party dependencies - stdlib only (urllib), so this runs with a
bare `python3` on any machine that has the transcript-speaker-fill skill
installed.
Requires FIREFLIES_API_KEY in the environment (Settings > API in the
Fireflies web app to generate one).
"""
import argparse
import json
import os
import re
import sys
import time
import urllib.error
import urllib.request
from pathlib import Path
API_URL = "https://api.fireflies.ai/graphql"
GENERIC_RE = re.compile(r'^Speaker\s+\d+$')
def parse_transcript_ref(ref):
"""Accept a bare id, or a shared link like
https://app.fireflies.ai/view/Some-Title::abcDEF123?channelSource=mine-shared
(query string / fragment and the ::title part are both optional) and
return just the transcript id."""
ref = ref.strip()
if '://' in ref:
ref = ref.split('?', 1)[0].split('#', 1)[0]
ref = ref.rstrip('/').rsplit('/', 1)[-1]
if '::' in ref:
ref = ref.rsplit('::', 1)[-1]
return ref
def gql(api_key, query, variables=None):
body = json.dumps({"query": query, "variables": variables or {}}).encode('utf-8')
req = urllib.request.Request(
API_URL,
data=body,
method='POST',
headers={
'Content-Type': 'application/json',
'Authorization': f'Bearer {api_key}',
},
)
try:
with urllib.request.urlopen(req, timeout=60) as resp:
payload = json.loads(resp.read().decode('utf-8'))
except urllib.error.HTTPError as e:
raw = e.read().decode('utf-8', errors='replace')
sys.exit(f"HTTP {e.code} from Fireflies API: {raw}")
if payload.get('errors'):
msgs = []
for err in payload['errors']:
ext = err.get('extensions') or {}
code = ext.get('code', '')
msg = err.get('message', '')
extra = f" (code={code}" + (f", retryAfter={ext['retryAfter']}" if 'retryAfter' in ext else '') + ')' if code else ''
msgs.append(f"{msg}{extra}")
sys.exit("Fireflies API returned error(s): " + " | ".join(msgs))
return payload.get('data') or {}
def require_api_key():
key = os.environ.get('FIREFLIES_API_KEY')
if not key:
sys.exit(
"FIREFLIES_API_KEY is not set. Generate a key in the Fireflies web "
"app under Settings > API, then export it, e.g.:\n"
" export FIREFLIES_API_KEY=your-key-here"
)
return key
def format_ts(seconds):
seconds = float(seconds)
total = int(round(seconds))
h, rem = divmod(total, 3600)
m, s = divmod(rem, 60)
if h:
return f"{h}:{m:02d}:{s:02d}"
return f"{m}:{s:02d}"
TRANSCRIPT_FIELDS = """
id
title
dateString
date
duration
audio_url
video_url
sentences {
speaker_name
speaker_id
start_time
end_time
text
raw_text
}
"""
def cmd_fetch(args):
api_key = require_api_key()
tid = parse_transcript_ref(args.ref)
data = gql(
api_key,
f"query Transcript($id: String!) {{ transcript(id: $id) {{ {TRANSCRIPT_FIELDS} }} }}",
{"id": tid},
)
t = data.get('transcript')
if not t:
sys.exit(
f"No transcript returned for id '{tid}'. Either the id is wrong, or this "
f"API key's account doesn't have access to it (not shared with you, or "
f"the workspace/plan doesn't expose it via API)."
)
sentences = t.get('sentences') or []
if not sentences:
sys.exit(
f"Transcript '{t.get('title')}' ({tid}) has no sentences yet - it may "
f"still be processing, or your account tier doesn't return transcript "
f"content via the API for this meeting."
)
out_path = Path(args.out)
lines = []
for s in sentences:
speaker = s.get('speaker_name') or f"Speaker {s.get('speaker_id', '?')}"
text = s.get('text') or s.get('raw_text') or ''
ts = format_ts(s.get('start_time', 0))
lines.append(f"**{speaker}** *[{ts}]*: {text}")
out_path.write_text("\n".join(lines) + "\n", encoding='utf-8')
meta = {
'id': t['id'],
'title': t.get('title'),
'dateString': t.get('dateString'),
'date': t.get('date'),
'duration': t.get('duration'),
'audio_url': t.get('audio_url'),
'video_url': t.get('video_url'),
'sentence_count': len(sentences),
}
meta_path = out_path.with_suffix(out_path.suffix + '.meta.json')
meta_path.write_text(json.dumps(meta, indent=2), encoding='utf-8')
print(f"Wrote {len(lines)} lines to {out_path}")
print(f"Wrote metadata to {meta_path}")
print(f"Title: {meta['title']}")
print(f"audio_url present: {bool(meta['audio_url'])}")
print(f"video_url present: {bool(meta['video_url'])}")
if not meta['audio_url']:
print(
"\nWARNING: audio_url is null. This is the expected failure mode when the "
"querying account isn't on a Fireflies Pro+ seat, or when audio access "
"isn't granted to a merely-shared (non-owned) meeting. Re-uploading this "
"meeting's audio via the API is NOT possible until this is resolved - "
"either use an API key belonging to a Pro+ seat, ask the meeting owner to "
"reprocess it directly, or manually download the audio via the Fireflies "
"web UI (if the share settings allow it) and host it at a public HTTPS URL "
"yourself before using the `upload` command's --audio-url override.",
file=sys.stderr,
)
def cmd_upload(args):
api_key = require_api_key()
meta = json.loads(Path(args.meta).read_text(encoding='utf-8'))
audio_url = args.audio_url or meta.get('audio_url')
if not audio_url:
sys.exit(
"No audio_url available (neither in the meta file nor via --audio-url). "
"See the WARNING printed by the `fetch` command for why, and how to work "
"around it."
)
original_id = meta['id']
title = args.title or f"{meta.get('title', 'Meeting')} [relang-{original_id}]"
variables = {
"input": {
"url": audio_url,
"title": title,
"custom_language": args.language,
"client_reference_id": original_id,
}
}
if args.bypass_size_check:
variables["input"]["bypass_size_check"] = True
data = gql(
api_key,
"""
mutation UploadAudio($input: AudioUploadInput!) {
uploadAudio(input: $input) {
success
title
message
}
}
""",
variables,
)
result = data.get('uploadAudio') or {}
echoed_title = result.get('title') or title
print(json.dumps({
"submitted_title": title,
"language": args.language,
"source_transcript_id": original_id,
"api_response": result,
}, indent=2))
if not result.get('success'):
sys.exit("uploadAudio did not report success - check the message above.")
if echoed_title != title:
print(
f"\nNOTE: Fireflies echoed back a different title than submitted "
f"(likely stripped/altered some character) - use the ECHOED one below "
f"for `wait`, not what you submitted."
)
print(
f"\nQueued. Use the `wait` command with --title '{echoed_title}' to find the "
f"new transcript once processing finishes (this can take several minutes for "
f"a long recording)."
)
def cmd_wait(args):
api_key = require_api_key()
deadline = time.time() + args.max_wait
attempt = 0
while True:
attempt += 1
data = gql(
api_key,
"""
query Transcripts($limit: Int) {
transcripts(mine: true, limit: $limit) {
id
title
dateString
sentences { speaker_name }
}
}
""",
{"limit": args.list_limit},
)
candidates = [t for t in (data.get('transcripts') or []) if t.get('title') == args.title]
ready = [t for t in candidates if t.get('sentences')]
if ready:
t = ready[0]
print(json.dumps({"id": t['id'], "title": t['title'], "dateString": t.get('dateString'), "ready": True}, indent=2))
return
if candidates:
print(f"[attempt {attempt}] Found the meeting but it's still processing (no sentences yet)...", file=sys.stderr)
else:
print(f"[attempt {attempt}] Not found yet...", file=sys.stderr)
if time.time() >= deadline:
sys.exit(
f"Gave up after {args.max_wait}s without finding a ready transcript "
f"titled '{args.title}'. Long recordings can take longer than that to "
f"process - re-run `wait` again with a fresh --max-wait, or check the "
f"Fireflies web UI directly for a meeting with that title."
)
time.sleep(args.interval)
def cmd_delete(args):
if not args.yes:
sys.exit("Refusing to delete without --yes (this is irreversible).")
api_key = require_api_key()
data = gql(
api_key,
"""
mutation DeleteTranscript($id: String!) {
deleteTranscript(id: $id) {
id
title
date
duration
}
}
""",
{"id": args.id},
)
print(json.dumps(data.get('deleteTranscript') or {}, indent=2))
def main():
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
sub = ap.add_subparsers(dest='command', required=True)
p = sub.add_parser('fetch', help="Fetch a transcript (owned or shared) into transcript-speaker-fill-compatible Markdown + a metadata sidecar.")
p.add_argument('--ref', required=True, help='Transcript id, or a Fireflies share URL/link')
p.add_argument('--out', required=True, help='Output .md path')
p.set_defaults(func=cmd_fetch)
p = sub.add_parser('upload', help="Re-upload a transcript's audio under a different language via uploadAudio.")
p.add_argument('--meta', required=True, help="Path to the .meta.json produced by `fetch` for the ORIGINAL (broken-language) transcript")
p.add_argument('--language', required=True, help='Target language code, e.g. "en", "pl", "es"')
p.add_argument('--title', help='Override the title used for the new upload (default: original title + a [relang-<id>] tag)')
p.add_argument('--audio-url', help='Override the audio URL instead of using the one from --meta (e.g. a self-hosted fallback URL)')
p.add_argument('--bypass-size-check', action='store_true')
p.set_defaults(func=cmd_upload)
p = sub.add_parser('wait', help='Poll until the re-uploaded transcript shows up as fully processed.')
p.add_argument('--title', required=True, help='Exact title echoed back by `upload` (not necessarily what you submitted - Fireflies can alter it, e.g. stripping colons)')
p.add_argument('--interval', type=int, default=30, help='Seconds between polls (default 30)')
p.add_argument('--max-wait', type=int, default=900, help='Give up after this many seconds (default 900 = 15 min)')
p.add_argument('--list-limit', type=int, default=20, help="How many of your most recent transcripts to scan each poll for an exact title match (default 20; server 'keyword' search was found unreliable against bracket-tagged titles, so this lists recent transcripts client-side instead of filtering server-side)")
p.set_defaults(func=cmd_wait)
p = sub.add_parser('delete', help='Delete a transcript by id (irreversible).')
p.add_argument('--id', required=True)
p.add_argument('--yes', action='store_true', help='Required confirmation flag')
p.set_defaults(func=cmd_delete)
args = ap.parse_args()
args.func(args)
if __name__ == '__main__':
main()

View file

@ -1,13 +0,0 @@
# Invoice Checker — credentials template
# Copy this file to .env (in the project root) and fill in the values.
# .env is gitignored and never committed.
# Your company NIP — 10 digits, no dashes or spaces
KSEF_NIP=1111111111
# KSeF authorization token — generate in the KSeF portal under Zarządzanie tokenami
# Grant it at minimum: "Odczyt faktur" (invoice read) permissions.
KSEF_TOKEN=20260601-EC-25684C5000-A33DFDC343-5F|nip-1111111111|4b6eaeb9027d45428fe714ff2a32a8f75bae7b2b83c14fb1bd64a60543767df5
# Environment: "test" (sandbox) or "prod" (production)
KSEF_ENV=test

View file

@ -1 +0,0 @@
.env

View file

@ -1,116 +0,0 @@
---
name: invoice-checker
description: >
Pull and check invoices from KSeF (Krajowy System e-Faktur) for a selected period.
For every new invoice checks whether the contractor is in the contractors file and
whether the contractor's bank account is on the MF white list (Biała lista podatników VAT).
TRIGGER this skill automatically whenever the user mentions pulling, checking, reviewing,
or auditing invoices from KSeF — especially when they mention a period like last week,
last month, or a date range. Trigger phrases include: "check invoices", "pull invoices",
"invoice audit", "KSeF invoices", "check KSeF", "white list check", "biała lista".
---
# Invoice Checker
Use this skill when someone needs to download purchase invoices from KSeF for a given period, verify contractors, and check bank accounts against the MF white list.
**In Claude Code, you run the full pipeline automatically.** The user provides a period — you handle everything else.
## Prerequisites
This skill reads credentials from the `.env` file in the **project root** (the working directory where Claude Code is open). The same file is shared with the timesheet-checker skill — just add the KSEF variables to it.
Add these lines to your `.env`:
```env
# KSeF — Invoice Checker
KSEF_NIP= # 10-digit company NIP, no dashes
KSEF_TOKEN= # Authorization token from KSeF portal → Zarządzanie tokenami
KSEF_ENV=prod # "test" for sandbox, "prod" for production
```
The variable reference:
| Variable | What it is | Where to get it |
|---|---|---|
| `KSEF_NIP` | Company NIP (10 digits, no dashes) | Your company's tax ID |
| `KSEF_TOKEN` | KSeF authorization token | KSeF portal → Zarządzanie tokenami |
| `KSEF_ENV` | `test` or `prod` | Use `test` for the sandbox, `prod` for production |
The script also requires the `cryptography` Python package:
```bash
pip3 install cryptography
```
## Setup
Determine `SKILL_DIR` from the location of this file. Set the output directory to `./invoices-output/` in the current working directory. Create it if it does not exist.
## Load First
Read these reference files:
1. The contractors config — check `./contractors.md` in the current working directory first; if not present, fall back to `{SKILL_DIR}/references/contractors.md`. Tell the user which one is being used.
2. `{SKILL_DIR}/references/output-format.md` — the structure of `invoices_data.json` and the final report.
## Extracting the period
The user may express the period in various ways. Convert to `YYYY-MM-DD:YYYY-MM-DD` before running:
| User says | Start | End |
|---|---|---|
| "last week" | Last Monday | Last Friday |
| "last month" | 1st of previous month | Last day of previous month |
| "this month" | 1st of current month | Yesterday |
| "this year" | 1 January of current year | Yesterday |
| "June", "June 2026" | 2026-06-01 | 2026-06-30 |
| Explicit dates | As given | As given |
If the user has not specified a period, ask: "Który okres sprawdzić? (np. zeszły tydzień, zeszły miesiąc, zakres dat)"
## Workflow
**Step 1 — Fetch and check.** Run this command:
```bash
python3 {SKILL_DIR}/scripts/preprocess.py \
--period {YYYY-MM-DD:YYYY-MM-DD} \
--contractors {contractors_path} \
--output-dir ./invoices-output
```
This authenticates with KSeF, downloads all purchase invoices for the period, compares against `invoices_cache.json` to identify new invoices, checks each contractor against `contractors.md`, checks bank accounts against the MF white list, and writes `invoices_data.json` to `./invoices-output/`.
**Step 2 — Interpret.** Read `./invoices-output/invoices_data.json`. Note:
- **New vs known invoices** — only new ones (not in cache) are actively checked.
- **Unknown contractors** — contractors not in `contractors.md`. Show these to the user and ask if they should be added.
- **White list failures** — bank accounts not found on the white list. These are high priority — flag them prominently.
- **White list errors** — API call failed (network issue, etc.). Note that these could not be verified.
**Step 3 — Generate report.**
```bash
python3 {SKILL_DIR}/scripts/postprocess.py \
--data ./invoices-output/invoices_data.json \
--output-dir ./invoices-output
```
This writes `invoices_report.md`.
**Step 4 — Present results.** Show the user:
- Summary block: total invoices, new invoices, contractor status, white list status.
- Table of new invoices with contractor and white list status.
- Any white list failures or unknown contractors — highlighted prominently.
- "Full report saved to `./invoices-output/invoices_report.md`."
**Step 5 — Handle unknown contractors.** If the user confirms any unknown contractors should be added, add them to `contractors.md`. Use the NIP and name from the invoice; leave `bank_accounts` blank if they want to fill it in later, or populate it from the invoice bank account.
## Guardrails
- Never fabricate invoice data. All data comes from KSeF.
- White list checks are always performed against the invoice issue date (not today).
- If `KSEF_NIP` or `KSEF_TOKEN` are not set, stop and ask the user for them.
- If the `cryptography` package is missing, tell the user to run `pip3 install cryptography` and stop.
- If any script step fails, show the error and stop. Do not proceed with partial output.
- Always terminate the KSeF session even if an error occurs (the script handles this internally).

View file

@ -1,15 +0,0 @@
# Contractors
Add a row for each contractor. Use their NIP as the primary identifier — the script matches invoices by NIP.
`bank_accounts` is a semicolon-separated list of IBAN numbers (with or without the PL prefix, spaces ignored).
Leave it blank if you do not want to maintain expected accounts — the white list check will still run against the account from the invoice.
`max_invoice_net_value` — maximum net total (PLN) allowed for a single invoice. Leave blank for no limit.
`max_position_net_rate` — maximum unit price (net) allowed on any single line item. Leave blank for no limit.
`max_total_quantity` — maximum sum of quantities across all line items on a single invoice. Leave blank for no limit.
`email` — contact email address for the contractor (used for notifications).
`check_with_jira` — set to `yes` (or `true` / `1`) to indicate invoices from this contractor should be cross-checked with Jira.
| name | nip | bank_accounts | max_invoice_net_value | max_position_net_rate | max_total_quantity | email | check_with_jira | notes |
| -----| ---| --------------| ----------------------| ----------------------| -------------------| ------| ----------------| ------|

View file

@ -1,70 +0,0 @@
# Output Format
## invoices_data.json (written by preprocess.py)
```json
{
"period": {
"start": "2026-06-01",
"end": "2026-06-30"
},
"summary": {
"total_invoices": 12,
"new_invoices": 4,
"known_contractors": 3,
"unknown_contractors": 1,
"whitelist_ok": 3,
"whitelist_failed": 1,
"whitelist_error": 0,
"account_mismatch": 0
},
"invoices": [
{
"ksef_reference": "1234567890-20260601-ABC123",
"invoice_number": "FV/123/06/2026",
"issue_date": "2026-06-01",
"seller_nip": "1234567890",
"seller_name": "Firma ABC Sp. z o.o.",
"gross_amount": 1230.00,
"currency": "PLN",
"bank_account": "PL61109010140000071219812874",
"is_new": true,
"contractor_known": true,
"account_in_contractors": true,
"whitelist_status": "ok",
"whitelist_account_assigned": true,
"whitelist_error": null
}
],
"known_invoices_count": 8
}
```
### Field notes
- `is_new` — true if the KSeF reference was not in `invoices_cache.json` before this run. The cache is updated at the end of a successful run.
- `contractor_known` — true if the seller NIP is in `contractors.md`.
- `account_in_contractors` — true if the invoice bank account matches one of the accounts listed in `contractors.md` for this NIP. `null` if the contractor is unknown or has no accounts configured.
- `whitelist_status``"ok"` (account found on white list), `"failed"` (account NOT on white list), `"no_account"` (invoice has no bank account), `"error"` (API call failed).
- `whitelist_account_assigned` — the raw `accountAssigned` value from the MF API (`true`/`false`). `null` if not checked.
- `whitelist_error` — error message string if the API call failed, otherwise `null`.
## invoices_report.md (written by postprocess.py)
Manager-readable markdown report containing:
- Period and run timestamp
- Summary block
- Table of new invoices with contractor and white list status
- Section highlighting white list failures (if any)
- Section listing unknown contractors (if any)
### Status icons used in the report
| Icon | Meaning |
|---|---|
| ✓ | White list OK |
| ✗ | White list FAILED — do not pay until resolved |
| ? | Could not verify (API error) |
| — | No bank account on invoice |
| ★ | Known contractor |
| ✦ | Unknown contractor |

View file

@ -1,421 +0,0 @@
#!/usr/bin/env python3
"""Format invoice check data into a markdown report."""
from __future__ import annotations
import argparse
import json
import sys
from datetime import datetime, timezone
from pathlib import Path
from typing import Any
def fmt_amount(amount: Any, currency: str = "PLN") -> str:
if amount is None:
return ""
try:
return f"{float(amount):,.2f} {currency}"
except (TypeError, ValueError):
return str(amount)
def contractor_badge(inv: dict) -> str:
if inv.get("contractor_known"):
mismatch = inv.get("account_in_contractors") is False
return "" + (" ⚠acc" if mismatch else "")
return "✦ NEW"
def whitelist_badge(inv: dict) -> str:
status = inv.get("whitelist_status", "")
if status == "ok":
return ""
if status == "failed":
return "✗ FAILED"
if status == "no_account":
return ""
if status == "error":
return "? err"
if status == "no_nip":
return "? no NIP"
return status
def rules_badge(inv: dict) -> str:
violations = inv.get("rule_violations") or []
return f"{len(violations)}" if violations else ""
def duplicate_badge(inv: dict) -> str:
if inv.get("potential_duplicate"):
return "⚠ DUPE?"
if inv.get("second_in_month"):
return "2nd"
return ""
def jira_badge(inv: dict) -> str:
if not inv.get("jira_checked"):
return ""
if inv.get("jira_compliant"):
return ""
return ""
def _invoice_table_row(inv: dict) -> str:
return (
f"| {inv.get('invoice_number') or inv.get('ksef_reference', '')} "
f"| {inv.get('issue_date') or ''} "
f"| {inv.get('seller_name') or ''} "
f"| {inv.get('seller_nip') or ''} "
f"| {fmt_amount(inv.get('gross_amount'), inv.get('currency', 'PLN'))} "
f"| {contractor_badge(inv)} "
f"| {whitelist_badge(inv)} "
f"| {rules_badge(inv)} "
f"| {duplicate_badge(inv)} "
f"| {jira_badge(inv)} |"
)
def _invoice_table_header() -> list[str]:
return [
"| Invoice | Date | Seller | NIP | Amount | Contractor | White list | Rules | Dup. | Jira |",
"| ---| ---| ---| ---| ---:| ---| ---| ---| ---| ---|",
]
def _legend() -> list[str]:
return [
"",
"---",
"",
"**Legend:** ★ = known contractor · ✦ NEW = unknown contractor · "
"✓ = white list OK · ✗ FAILED = not on white list · "
"— = no bank account · ? = could not verify · ⚠acc = account not in contractors.md · "
"⚠ N = N rule violation(s) · ✓ = rules OK · "
"2nd = 2nd invoice from seller this month · ⚠ DUPE? = same amount as earlier invoice · "
"✓/✗ Jira = Jira compliance check result",
"",
"_White list source: [wl-api.mf.gov.pl](https://wl-api.mf.gov.pl)_",
"_Invoice source: KSeF (Ministerstwo Finansów)_",
]
def build_report(data: dict[str, Any]) -> str:
"""Build the per-run report covering only new invoices."""
period = data["period"]
summary = data["summary"]
all_invoices = data.get("invoices", [])
new_invoices = [i for i in all_invoices if i.get("is_new")]
generated_at = datetime.now(timezone.utc).strftime("%Y-%m-%d %H:%M UTC")
env = data.get("ksef_env", "prod")
lines = [
"# Invoice Check Report",
"",
f"Generated: {generated_at}",
f"Period: {period['start']} to {period['end']}",
f"KSeF environment: {env}",
"",
"## Summary",
"",
f"- Total invoices in period: **{summary['total_invoices']}**",
f"- New (checked this run): **{summary['new_invoices']}**",
f"- Known contractors: {summary['known_contractors']}",
f"- Unknown contractors: **{summary['unknown_contractors']}**",
f"- White list OK: {summary['whitelist_ok']}",
f"- White list FAILED: **{summary['whitelist_failed']}**",
f"- White list errors: {summary['whitelist_error']}",
f"- Account mismatch (not in contractors.md): {summary['account_mismatch']}",
f"- Rule violations (limit exceeded): **{summary.get('rule_violations', 0)}**",
f"- 2nd invoice from same seller this month: **{summary.get('second_invoices', 0)}**",
f"- Potential duplicates (same seller, same month, same amount): **{summary.get('potential_duplicates', 0)}**",
f"- Jira non-compliant: **{summary.get('jira_non_compliant', 0)}**",
]
# Jira non-compliance — prominent
jira_bad = [i for i in new_invoices if i.get("jira_checked") and not i.get("jira_compliant")]
if jira_bad:
lines += [
"",
"## ✗ Jira Non-Compliant Invoices",
"",
"> Quantities on these invoices do not match hours logged in Jira/Tempo.",
"",
]
for inv in jira_bad:
lines += [
f"**{inv.get('invoice_number') or inv.get('ksef_reference', '')}** "
f"{inv.get('seller_name') or ''} ({inv.get('issue_date') or ''})",
"",
]
violations = inv.get("jira_violations") or []
no_proj = inv.get("jira_no_project_positions") or []
if violations:
lines += [
"| Project | Invoice qty | Jira hours | Month |",
"| ---| ---:| ---:| ---|",
]
for v in violations:
lines.append(
f"| {v['project']} | {v['invoice_qty']} | {v['jira_hours']} | {v['month']} |"
)
lines.append("")
if no_proj:
lines.append(
"Positions without a Jira project code: "
+ ", ".join(
f"#{p['line_number']} ({p['description'] or 'no description'})"
for p in no_proj
)
)
lines.append("")
# Potential duplicates — prominent
dupes = [i for i in new_invoices if i.get("potential_duplicate")]
if dupes:
lines += [
"",
"## ⚠ Potential Duplicates",
"",
"> These invoices have the same seller and gross amount as another invoice in the same month.",
"> Verify they are not duplicate payments before approving.",
"",
"| Invoice | Date | Seller | NIP | Amount |",
"| ---| ---| ---| ---| ---:|",
]
for inv in dupes:
lines.append(
f"| {inv.get('invoice_number') or inv.get('ksef_reference', '')} "
f"| {inv.get('issue_date') or ''} "
f"| {inv.get('seller_name') or ''} "
f"| {inv.get('seller_nip') or ''} "
f"| {fmt_amount(inv.get('gross_amount'), inv.get('currency', 'PLN'))} |"
)
# Rule violations — prominent
rule_violated = [i for i in new_invoices if i.get("rule_violations")]
if rule_violated:
lines += [
"",
"## ⚠ Rule Violations",
"",
"> These invoices exceed limits configured in `contractors.md`.",
"",
"| Invoice | Seller | Rule | Limit | Actual | Position |",
"| ---| ---| ---| ---:| ---:| ---|",
]
for inv in rule_violated:
for v in inv.get("rule_violations", []):
rule_label = {
"max_invoice_net_value": "Max invoice net",
"max_position_net_rate": "Max position rate",
"max_total_quantity": "Max total qty",
}.get(v["rule"], v["rule"])
pos_info = f"#{v['position']} {v.get('description', '')}" if v.get("position") else ""
lines.append(
f"| {inv.get('invoice_number') or inv.get('ksef_reference', '')} "
f"| {inv.get('seller_name') or ''} "
f"| {rule_label} "
f"| {v['limit']:,.2f} "
f"| {v['actual']:,.2f} "
f"| {pos_info} |"
)
# White list failures — prominent
failed = [i for i in new_invoices if i.get("whitelist_status") == "failed"]
if failed:
lines += [
"",
"## ✗ White List Failures — DO NOT PAY",
"",
"> These bank accounts are **not registered** on the MF white list for the given NIP.",
"> Verify manually before making any payment.",
"",
"| Invoice | Seller | NIP | Bank account | Amount |",
"| ---| ---| ---| ---| ---:|",
]
for inv in failed:
lines.append(
f"| {inv.get('invoice_number') or inv['ksef_reference']} "
f"| {inv.get('seller_name') or ''} "
f"| {inv.get('seller_nip') or ''} "
f"| {inv.get('bank_account') or ''} "
f"| {fmt_amount(inv.get('gross_amount'), inv.get('currency', 'PLN'))} |"
)
# New invoices table
if new_invoices:
lines += ["", "## New Invoices", ""] + _invoice_table_header()
for inv in new_invoices:
lines.append(_invoice_table_row(inv))
else:
lines += ["", "## New Invoices", "", "_No new invoices found for this period._"]
# Unknown contractors
unknown = [i for i in new_invoices if not i.get("contractor_known")]
if unknown:
lines += [
"",
"## ✦ Unknown Contractors",
"",
"These sellers are not in `contractors.md`. Consider adding them.",
"",
"| Seller name | NIP | Bank account |",
"| ---| ---| ---|",
]
seen_nips: set[str] = set()
for inv in unknown:
nip = inv.get("seller_nip") or ""
if nip in seen_nips:
continue
seen_nips.add(nip)
lines.append(
f"| {inv.get('seller_name') or ''} | {nip} | {inv.get('bank_account') or ''} |"
)
# Account mismatches
mismatches = [i for i in new_invoices if i.get("account_in_contractors") is False]
if mismatches:
lines += [
"",
"## ⚠ Account Mismatches",
"",
"These invoices carry a bank account **not listed** in `contractors.md` for the seller.",
"",
"| Invoice | Seller | NIP | Invoice account |",
"| ---| ---| ---| ---|",
]
for inv in mismatches:
lines.append(
f"| {inv.get('invoice_number') or inv['ksef_reference']} "
f"| {inv.get('seller_name') or ''} "
f"| {inv.get('seller_nip') or ''} "
f"| {inv.get('bank_account') or ''} |"
)
# White list errors
errors = [i for i in new_invoices if i.get("whitelist_status") == "error"]
if errors:
lines += [
"",
"## ? White List Check Errors",
"",
"| Invoice | Seller | NIP | Error |",
"| ---| ---| ---| ---|",
]
for inv in errors:
lines.append(
f"| {inv.get('invoice_number') or inv['ksef_reference']} "
f"| {inv.get('seller_name') or ''} "
f"| {inv.get('seller_nip') or ''} "
f"| {inv.get('whitelist_error') or ''} |"
)
lines += _legend()
return "\n".join(lines) + "\n"
def build_all_report(all_data: dict[str, Any]) -> str:
"""Build the cumulative report covering all invoices ever processed."""
updated_at = all_data.get("updated_at", "")[:19].replace("T", " ") + " UTC"
invoices_map: dict[str, dict] = all_data.get("invoices", {})
# Sort by issue_date descending, then by ksef_reference for stability
invoices = sorted(
invoices_map.values(),
key=lambda i: (i.get("issue_date") or "0000-00-00", i.get("ksef_reference", "")),
reverse=True,
)
total = len(invoices)
wl_ok = sum(1 for i in invoices if i.get("whitelist_status") == "ok")
wl_failed = sum(1 for i in invoices if i.get("whitelist_status") == "failed")
unknown_count = sum(1 for i in invoices if not i.get("contractor_known"))
gross_total = sum(float(i["gross_amount"]) for i in invoices if i.get("gross_amount") is not None)
lines = [
"# All Invoices",
"",
f"Last updated: {updated_at}",
f"Total: **{total}** invoice(s)",
"",
"## Summary",
"",
f"- Total invoices: **{total}**",
f"- Total gross value: **{gross_total:,.2f} PLN**",
f"- White list OK: {wl_ok}",
f"- White list FAILED: **{wl_failed}**",
f"- Unknown contractors: **{unknown_count}**",
]
# White list failures — prominent
failed = [i for i in invoices if i.get("whitelist_status") == "failed"]
if failed:
lines += [
"",
"## ✗ White List Failures — DO NOT PAY",
"",
"> These bank accounts are **not registered** on the MF white list.",
"",
"| Invoice | Seller | NIP | Bank account | Amount |",
"| ---| ---| ---| ---| ---:|",
]
for inv in failed:
lines.append(
f"| {inv.get('invoice_number') or inv.get('ksef_reference', '')} "
f"| {inv.get('seller_name') or ''} "
f"| {inv.get('seller_nip') or ''} "
f"| {inv.get('bank_account') or ''} "
f"| {fmt_amount(inv.get('gross_amount'), inv.get('currency', 'PLN'))} |"
)
# Full invoices table
lines += ["", "## All Invoices (newest first)", ""] + _invoice_table_header()
for inv in invoices:
lines.append(_invoice_table_row(inv))
lines += _legend()
return "\n".join(lines) + "\n"
def run(args: argparse.Namespace) -> int:
data_path = Path(args.data)
if not data_path.exists():
print(f"Error: {data_path} not found. Run preprocess.py first.", file=sys.stderr)
return 1
data = json.loads(data_path.read_text(encoding="utf-8"))
output_dir = Path(args.output_dir)
output_dir.mkdir(parents=True, exist_ok=True)
# Per-run report (new invoices only)
report_path = output_dir / "invoices_report.md"
report_path.write_text(build_report(data), encoding="utf-8")
print(f"Wrote {report_path}")
# Cumulative report (all invoices ever processed)
all_data_path = output_dir / "invoices_all.json"
if all_data_path.exists():
all_data = json.loads(all_data_path.read_text(encoding="utf-8"))
all_report_path = output_dir / "invoices_all_report.md"
all_report_path.write_text(build_all_report(all_data), encoding="utf-8")
print(f"Wrote {all_report_path}")
return 0
def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
description="Format invoice check data into markdown reports."
)
parser.add_argument("--data", required=True,
help="Path to invoices_data.json produced by preprocess.py.")
parser.add_argument("--output-dir", required=True,
help="Directory for invoices_report.md and invoices_all_report.md.")
return parser
if __name__ == "__main__":
sys.exit(run(build_parser().parse_args()))

File diff suppressed because it is too large Load diff

View file

@ -1,84 +0,0 @@
---
name: invoice-prep
description: >
Show what to put on an invoice for a selected month by checking the user's
own Tempo timesheets. Lists every Jira project code and the total hours logged
against it. Warns if there are missing hours or days with no logged time.
TRIGGER when user asks what to put on an invoice, what to invoice, invoice
preparation, timesheet summary for invoicing, or similar phrases.
Trigger phrases: "what should I put on an invoice", "prepare my invoice",
"invoice for [month]", "what to invoice for", "check my TS for invoice".
---
# Invoice Preparation
Use this skill when someone wants to know what project codes and hours to put on their invoice for a given month. It reads their own Tempo worklogs, groups hours by Jira project, and warns about any missing time.
**In Claude Code, you run the pipeline automatically.** The user provides a month — you handle everything else.
## Prerequisites
Uses the same credentials as the timesheet-checker — no extra setup needed:
| Variable | Purpose |
|---|---|
| `TEMPO_API_TOKEN` | Fetch worklogs from Tempo |
| `JIRA_BASE_URL` | Resolve the user's account ID |
| `JIRA_EMAIL` | Identifies whose timesheets to read |
| `JIRA_API_TOKEN` | Authenticate with Jira |
All four must be present in the `.env` file in the project root.
## Extracting the month
Convert natural language to `YYYY-MM` before running:
| User says | CLI value |
|---|---|
| "last month" | Previous calendar month |
| "this month" | Current calendar month |
| "June", "June 2026" | `2026-06` |
| "May" | `2026-05` (current year) |
If no month is specified, ask: "Which month should I check? (e.g. last month, June 2026)"
## Workflow
**Step 1 — Run the script:**
```bash
python3 {SKILL_DIR}/scripts/run.py \
--month {YYYY-MM} \
--output-dir ./timesheet-output
```
This resolves the user's Jira account, fetches their Tempo worklogs for the full calendar month, groups hours by Jira project code, checks completeness against the Polish working calendar, and writes `invoice_prep_{YYYY-MM}.json` to `./timesheet-output/`.
**Step 2 — Present results.** Read the output JSON and show the user:
1. A table of project codes and hours — formatted ready to copy onto an invoice.
Include **every project in the `projects` array**. Do NOT label any project as "non-billable" or omit it from the table — the script already excludes ignored issues; everything remaining is billable:
```
| Project | Hours |
| IAA | 45.5 |
| PROJ | 22.0 |
```
Below the table, note ignored issues and their hours as a parenthetical: `(INTERNAL-1: 40h ignored per config.)` — do not mix them into the billable table.
2. The **total billable hours** from `total_billable_hours` (not `total_logged_hours`).
3. If `missing_hours > 0`: a clear warning showing:
- How many hours are missing vs the expected total for the month
- Which specific dates have no logged time (up to 10; if more, show count)
- Reminder to complete the timesheet before issuing the invoice
4. If `total_logged_hours == 0`: warn that no hours were found at all for the month.
## Guardrails
- Never fabricate or estimate hours. All data comes from Tempo.
- If `JIRA_EMAIL` is not set, stop and tell the user.
- If any script step fails, show the error and stop.
- Always remind the user that incomplete timesheets will cause Jira compliance failures when the invoice is checked.

View file

@ -1,353 +0,0 @@
#!/usr/bin/env python3
"""
Generate invoice line items from the user's own Tempo worklogs for a given month.
Groups hours by Jira project code and reports missing time vs the Polish
working calendar.
"""
from __future__ import annotations
import argparse
import base64
import calendar
import json
import os
import re
import sys
import urllib.error
import urllib.request
from datetime import date, timedelta
from pathlib import Path
from typing import Any
# ── .env loading ──────────────────────────────────────────────────────────────
def load_dotenv(*paths: Path) -> Path | None:
for path in paths:
if not path.exists():
continue
for raw in path.read_text(encoding="utf-8").splitlines():
line = raw.strip()
if not line or line.startswith("#") or "=" not in line:
continue
key, _, value = line.partition("=")
key = key.strip()
value = value.strip().strip('"').strip("'")
if key and key not in os.environ:
os.environ[key] = value
return path
return None
# ── Period helpers ────────────────────────────────────────────────────────────
def parse_month(month_str: str) -> tuple[date, date]:
"""Parse 'YYYY-MM' into (first_day, last_day) of that month."""
s = month_str.strip()
if not re.fullmatch(r"\d{4}-\d{2}", s):
raise ValueError(f"Invalid month '{month_str}'. Use YYYY-MM.")
year, month = int(s[:4]), int(s[5:7])
first = date(year, month, 1)
last = date(year, month, calendar.monthrange(year, month)[1])
return first, last
# ── Public holidays ───────────────────────────────────────────────────────────
def get_public_holidays(country: str, year: int) -> set[date]:
url = f"https://date.nager.at/api/v3/PublicHolidays/{year}/{country.upper()}"
req = urllib.request.Request(url, headers={"User-Agent": "invoice-prep/1.0"})
try:
with urllib.request.urlopen(req, timeout=10) as resp:
return {date.fromisoformat(item["date"]) for item in json.loads(resp.read())}
except Exception as exc:
print(f"Warning: could not fetch holidays for {country} {year}: {exc}", file=sys.stderr)
return set()
def working_days(start: date, end: date, holidays: set[date]) -> list[date]:
result, cur = [], start
while cur <= end:
if cur.weekday() < 5 and cur not in holidays:
result.append(cur)
cur += timedelta(days=1)
return result
# ── Jira account lookup ───────────────────────────────────────────────────────
def resolve_account_id(email: str, jira_base_url: str,
jira_email: str, jira_token: str) -> str:
creds = base64.b64encode(f"{jira_email}:{jira_token}".encode()).decode()
headers = {"Authorization": f"Basic {creds}", "Accept": "application/json"}
url = f"{jira_base_url.rstrip('/')}/rest/api/3/user/search?query={email}&maxResults=5"
req = urllib.request.Request(url, headers=headers)
try:
with urllib.request.urlopen(req, timeout=10) as resp:
users = json.loads(resp.read())
except urllib.error.HTTPError as exc:
if exc.code == 401:
print("Error: Jira API returned 401. Check JIRA_EMAIL and JIRA_API_TOKEN.",
file=sys.stderr)
sys.exit(1)
raise RuntimeError(f"Jira API {exc.code}") from exc
match = next(
(u for u in users if u.get("emailAddress", "").lower() == email.lower()),
None,
)
if not match:
print(f"Error: Jira account not found for {email}.", file=sys.stderr)
sys.exit(1)
return match["accountId"]
# ── Tempo worklogs ────────────────────────────────────────────────────────────
_PROJECT_CODE_RE = re.compile(r"\b([A-Z][A-Z0-9]{1,9})-\d+\b")
def load_jira_ignore(path: Path) -> set[str]:
"""Load issue keys to ignore from a markdown table in jira_ignore.md."""
if not path.exists():
return set()
ignored: set[str] = set()
in_table = False
header_seen = False
for line in path.read_text(encoding="utf-8").splitlines():
line = line.strip()
if not (line.startswith("|") and line.endswith("|")):
in_table = False
header_seen = False
continue
cells = [c.strip() for c in line.strip("|").split("|")]
if not header_seen:
# First row is header or separator
if all(re.fullmatch(r":?-{3,}:?", c) for c in cells if c):
continue # separator row
header_seen = True
in_table = True
continue
if all(re.fullmatch(r":?-{3,}:?", c) for c in cells if c):
continue # separator row
if in_table and cells and cells[0]:
ignored.add(cells[0].strip().upper())
return ignored
def extract_project_code(issue_key: str) -> str | None:
"""Extract project code from a Jira issue key, e.g. 'IAA' from 'IAA-42'."""
m = _PROJECT_CODE_RE.match(issue_key.strip())
return m.group(1) if m else None
def fetch_worklogs(account_id: str, start: date, end: date,
token: str, base_url: str = "https://api.tempo.io/4") -> list[dict]:
"""Return all Tempo worklog entries for the account in [start, end]."""
results: list[dict] = []
headers = {"Authorization": f"Bearer {token}", "Accept": "application/json"}
next_url: str | None = (
f"{base_url}/worklogs/user/{account_id}"
f"?from={start}&to={end}&limit=1000"
)
while next_url:
req = urllib.request.Request(next_url, headers=headers)
try:
with urllib.request.urlopen(req, timeout=30) as resp:
data = json.loads(resp.read())
except urllib.error.HTTPError as exc:
body = exc.read().decode("utf-8", errors="replace")
if exc.code == 401:
print("Error: Tempo API returned 401. Check TEMPO_API_TOKEN.", file=sys.stderr)
sys.exit(1)
raise RuntimeError(f"Tempo API {exc.code}: {body[:200]}") from exc
results.extend(data.get("results", []))
next_url = data.get("metadata", {}).get("next")
return results
def resolve_issue_keys(worklogs: list[dict], jira_email: str,
jira_token: str) -> dict[int, str]:
"""
Resolve Jira issue IDs keys by calling each issue's self URL directly.
Uses the self URL already present in each Tempo worklog entry.
"""
creds = base64.b64encode(f"{jira_email}:{jira_token}".encode()).decode()
headers = {"Authorization": f"Basic {creds}", "Accept": "application/json"}
seen: dict[int, str] = {}
for wl in worklogs:
issue = wl.get("issue", {})
issue_id = issue.get("id")
if not issue_id or int(issue_id) in seen:
continue
self_url = issue.get("self", "")
if not self_url:
continue
# Swap api/2 → api/3 if needed; append fields filter
url = re.sub(r"/rest/api/\d+/", "/rest/api/3/", self_url) + "?fields=key"
req = urllib.request.Request(url, headers=headers)
try:
with urllib.request.urlopen(req, timeout=10) as resp:
data = json.loads(resp.read())
seen[int(issue_id)] = data.get("key", "")
except Exception as exc:
print(f"Warning: could not resolve issue {issue_id}: {exc}", file=sys.stderr)
return seen
# ── Main ──────────────────────────────────────────────────────────────────────
def run(args: argparse.Namespace) -> int:
skill_dir = Path(__file__).parent.parent
loaded = load_dotenv(Path.cwd() / ".env", skill_dir / ".env")
if loaded:
print(f"Loaded credentials from {loaded}")
tempo_token = os.environ.get("TEMPO_API_TOKEN", "")
jira_base_url = os.environ.get("JIRA_BASE_URL", "")
jira_email = os.environ.get("JIRA_EMAIL", "")
jira_token = os.environ.get("JIRA_API_TOKEN", "")
missing = [n for n, v in [
("TEMPO_API_TOKEN", tempo_token), ("JIRA_BASE_URL", jira_base_url),
("JIRA_EMAIL", jira_email), ("JIRA_API_TOKEN", jira_token),
] if not v]
if missing:
print(f"Error: missing credentials: {', '.join(missing)}", file=sys.stderr)
return 1
try:
start, end = parse_month(args.month)
except ValueError as exc:
print(f"Error: {exc}", file=sys.stderr)
return 1
# Load ignore list
ignore_path = Path.cwd() / "jira_ignore.md"
ignored_keys = load_jira_ignore(ignore_path)
if ignored_keys:
print(f"Ignoring issues: {', '.join(sorted(ignored_keys))}")
print(f"Month: {args.month} ({start} {end})")
print(f"User: {jira_email}")
# Resolve account ID
print("Resolving Jira account...")
account_id = resolve_account_id(jira_email, jira_base_url, jira_email, jira_token)
print(f" {jira_email}{account_id}")
# Working days (Polish calendar)
print("Fetching public holidays...")
holidays = get_public_holidays("PL", start.year)
wdays = working_days(start, end, holidays)
expected_hours = len(wdays) * 8.0
print(f" Working days: {len(wdays)} Expected hours: {expected_hours:.0f}h")
# Fetch Tempo worklogs
print("Fetching Tempo worklogs...")
worklogs = fetch_worklogs(account_id, start, end, tempo_token)
print(f" Worklogs fetched: {len(worklogs)}")
# Resolve issue IDs → keys via each issue's self URL (Tempo v4 omits the key)
unique_ids = {int(wl["issue"]["id"]) for wl in worklogs if wl.get("issue", {}).get("id")}
issue_key_map: dict[int, str] = {}
if unique_ids:
print(f" Resolving {len(unique_ids)} unique issue key(s) from Jira...")
issue_key_map = resolve_issue_keys(worklogs, jira_email, jira_token)
print(f" Resolved: {len(issue_key_map)}")
# Group hours by project code and by date, skipping ignored issues
project_hours: dict[str, float] = {}
hours_by_date: dict[str, float] = {}
unmatched_hours = 0.0
ignored_hours = 0.0
for wl in worklogs:
issue_id = int(wl.get("issue", {}).get("id", 0))
issue_key = issue_key_map.get(issue_id, "")
hours = wl.get("timeSpentSeconds", 0) / 3600
wl_date = wl.get("startDate", "")
if issue_key.upper() in ignored_keys:
ignored_hours += hours
# Still count toward date totals so missing-day detection stays accurate
if wl_date:
hours_by_date[wl_date] = hours_by_date.get(wl_date, 0.0) + hours
continue
code = extract_project_code(issue_key)
if code:
project_hours[code] = project_hours.get(code, 0.0) + hours
else:
unmatched_hours += hours
if wl_date:
hours_by_date[wl_date] = hours_by_date.get(wl_date, 0.0) + hours
# Sort projects by hours descending
projects = sorted(
[{"code": k, "hours": round(v, 2)} for k, v in project_hours.items()],
key=lambda x: x["hours"],
reverse=True,
)
total_logged = round(sum(project_hours.values()) + unmatched_hours + ignored_hours, 2)
total_billable = round(sum(project_hours.values()) + unmatched_hours, 2)
missing_hours = round(max(0.0, expected_hours - total_logged), 2)
# Find working days with no logged time
days_missing = [
d.isoformat() for d in wdays
if hours_by_date.get(d.isoformat(), 0.0) == 0.0
]
output = {
"month": args.month,
"user_email": jira_email,
"account_id": account_id,
"working_days": len(wdays),
"expected_hours": expected_hours,
"total_logged_hours": total_logged,
"total_billable_hours": total_billable,
"ignored_hours": round(ignored_hours, 2),
"ignored_keys": sorted(ignored_keys),
"missing_hours": missing_hours,
"unmatched_hours": round(unmatched_hours, 2),
"days_with_no_hours": days_missing,
"projects": projects,
}
output_dir = Path(args.output_dir)
output_dir.mkdir(parents=True, exist_ok=True)
out_path = output_dir / f"invoice_prep_{args.month}.json"
out_path.write_text(json.dumps(output, indent=2, ensure_ascii=False), encoding="utf-8")
print(f"Wrote {out_path}")
# Print quick summary
print(f"\nProjects:")
for p in projects:
print(f" {p['code']:12s} {p['hours']:.2f}h")
if unmatched_hours:
print(f" (no project) {unmatched_hours:.2f}h")
if ignored_hours:
print(f" (ignored) {ignored_hours:.2f}h ({', '.join(sorted(ignored_keys))})")
print(f"\nBillable total: {total_billable:.2f}h")
print(f"Total logged: {total_logged:.2f}h / {expected_hours:.0f}h expected")
if missing_hours > 0:
print(f"⚠ Missing: {missing_hours:.2f}h ({len(days_missing)} day(s) with no time logged)")
return 0
def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
description="Generate invoice line items from Tempo worklogs."
)
parser.add_argument("--month", required=True, help="Month to check: YYYY-MM")
parser.add_argument("--output-dir", required=True,
help="Directory for invoice_prep_YYYY-MM.json")
return parser
if __name__ == "__main__":
sys.exit(run(build_parser().parse_args()))

View file

@ -1,15 +0,0 @@
# Timesheet Checker — credentials template
# Copy this file to .env and fill in your values.
# .env is gitignored and never committed.
# Tempo API token — generate at: Tempo → Settings → API Integration
TEMPO_API_TOKEN=
# Your Atlassian / Jira Cloud URL
JIRA_BASE_URL=https://your-company.atlassian.net
# The email address you use to log into Jira
JIRA_EMAIL=your.email@company.com
# Jira personal API token — generate at: https://id.atlassian.com → Security → API tokens
JIRA_API_TOKEN=

View file

@ -1 +0,0 @@
.env

View file

@ -1,209 +0,0 @@
# Timesheet Checker
Checks whether employees have completed their timesheets for a given period by pulling Tempo worklogs. Compares against each employee's country calendar (public holidays fetched from date.nager.at) and a configurable minimum of hours per day. Produces a completion report ranked by missing hours.
---
## Prerequisites
The skill needs four credentials. The easiest setup is a `.env` file — copy the template and fill it in once:
macOS / Linux:
```bash
cp .claude/skills/timesheet-checker/.env.example .claude/skills/timesheet-checker/.env
```
Windows (PowerShell):
```powershell
Copy-Item .claude\skills\timesheet-checker\.env.example .claude\skills\timesheet-checker\.env
```
Or copy the file manually in Explorer / Finder — rename `.env.example` to `.env` in the skill folder.
Then open `.env` and fill in your values:
```
TEMPO_API_TOKEN=your_tempo_token
JIRA_BASE_URL=https://your-company.atlassian.net
JIRA_EMAIL=your.email@company.com
JIRA_API_TOKEN=your_jira_token
```
The script looks for `.env` in the **working directory first**, then the skill folder as a fallback. This means you can also place a `.env` in your project root if you prefer one file for all skills.
Environment variables set via `export` always take priority over `.env` values, so both approaches work side-by-side.
The `.env` file is gitignored — credentials are never committed.
| Variable | What it is | Where to get it |
|---|---|---|
| `TEMPO_API_TOKEN` | Tempo read token | Tempo → Settings → API Integration |
| `JIRA_BASE_URL` | Your Atlassian URL | The URL you use to open Jira |
| `JIRA_EMAIL` | Your Jira login email | Your Atlassian account email |
| `JIRA_API_TOKEN` | Jira personal API token | [id.atlassian.com](https://id.atlassian.com) → Security → API tokens |
If any credential is missing from both `.env` and the environment, the script stops with a clear error listing exactly what is needed.
---
## How to use in Claude Code
Claude Code runs as a CLI, desktop app, web app (claude.ai/code), or IDE extension. **Claude Code runs the full pipeline automatically — no manual script steps.**
**Install** — copy the skill folder into your project's Claude skills directory:
macOS / Linux:
```bash
cp -r src/timesheet-checker .claude/skills/
```
Windows (PowerShell):
```powershell
Copy-Item -Recurse -Path src\timesheet-checker -Destination .claude\skills\
```
**Use:**
Just describe what you want — Claude will pick up the skill automatically:
```
Check if timesheets are complete for last week
```
```
Timesheet completion for last month
```
```
Check logged hours for 2026-06-01 to 2026-06-30
```
Claude fetches Polish public holidays, pulls Tempo worklogs, computes completion per employee, generates the report, and shows you the results. Output files land in `./timesheet-output/`.
---
## How to use in the Claude.ai app
> Full automation (no manual script running) requires Claude Code. In the Claude.ai app, the Python steps must be run locally.
**Set up once (Project approach):**
1. Create a new Project in Claude.ai.
2. In **Project Instructions**, paste the full contents of `SKILL.md`.
3. Upload `references/employees.md` and `references/output-format.md` to the Project knowledge.
**Each time you run:**
1. Set all four credentials in your shell, then run the pre-processor locally:
```bash
python3 $SKILL_DIR/scripts/preprocess.py \
--period last-week \
--output-dir ./timesheet-output
```
The script will use `./employees.md` if it exists, otherwise `$SKILL_DIR/references/employees.md`.
2. Upload `timesheet_data.json` to the Project conversation.
3. Ask Claude to summarise and present the results.
4. Run the post-processor locally to generate the formatted report:
```bash
python3 $SKILL_DIR/scripts/postprocess.py \
--data ./timesheet-output/timesheet_data.json \
--output-dir ./timesheet-output
```
**Without Projects:** Paste the contents of `SKILL.md` at the start of a new conversation, attach the two reference files, then follow the same steps.
---
## What is included
| File | Description |
|---|---|
| `SKILL.md` | Orchestrating skill — instructions for Claude |
| `references/employees.md` | Editable list of employees and their Tempo account IDs |
| `references/output-format.md` | Structure of `timesheet_data.json` and the final report |
| `scripts/preprocess.py` | Fetches Tempo worklogs + Polish holidays, computes completion |
| `scripts/postprocess.py` | Formats `timesheet_data.json` into the markdown report |
| `sample/employees_sample.md` | Sample employees file with placeholder account IDs |
## Outputs (in `./timesheet-output/`)
| File | Description |
|---|---|
| `timesheet_data.json` | Computed stats per employee + period + summary |
| `timesheet_report.md` | Manager-readable report ranked by missing hours |
---
## Configuring employees
The script looks for the employees file in this order:
1. `employees.md` in the **current working directory** — place it here for a project-specific list
2. `references/employees.md` in the **skill folder** — the default shipped with the skill
To use a project-specific list, copy the sample and edit it:
```bash
cp .claude/skills/timesheet-checker/sample/employees_sample.md ./employees.md
```
The file has two tables:
**Defaults** — applies to every employee unless overridden:
| setting | value |
|---|---|
| `country` | ISO 3166-1 alpha-2 country code (e.g. `PL`, `GB`, `DE`). Used for public holiday lookup. |
| `expected_hours_per_day` | Minimum hours per working day |
**Employees** — one row per person. Leave any column blank to inherit the default.
| Column | What it controls |
|---|---|
| `name` | Display name in the report |
| `email` | Work email address (used to look up the Jira account ID automatically) |
| `country` | Override the default country for this employee |
| `expected_hours_per_day` | Override the default hours for this employee |
| `start_date` | First working day — leave blank if employed for the full period |
---
## Period formats
| Input | Meaning | End date |
|---|---|---|
| `last-week` | Previous MondayFriday | Last Friday |
| `last-month` | Previous calendar month | Last day of that month |
| `last-year` | Previous calendar year | 31 December |
| `yesterday` | Yesterday only | Yesterday |
| `current-week` | Monday of this week to yesterday | Yesterday |
| `current-month` | 1st of this month to yesterday | Yesterday |
| `current-year` | 1st January of this year to yesterday | Yesterday |
| `YYYY-MM-DD:YYYY-MM-DD` | Explicit date range | As specified |
For `current-*` periods the current in-progress day is always excluded — end date is yesterday. If today is the first day of the period (e.g. `current-week` on a Monday), the period contains no past working days and the report will show 0 expected hours.
In Claude Code, you can also use natural language ("last week", "this month", "yesterday", "June 2026") and Claude converts it automatically.
---
## How working days are calculated
The script fetches Polish public holidays from the [date.nager.at](https://date.nager.at) public API (no authentication needed). It then excludes weekends and those holidays from the period. This runs live each time — holiday data is never hardcoded.
---
## Integration
This skill connects to **Tempo Cloud** (`api.tempo.io/4`) by default. For self-hosted Tempo (Jira Data Center), pass the correct base URL:
```bash
python3 $SKILL_DIR/scripts/preprocess.py \
--tempo-base-url https://your-jira.company.com/rest/tempo-timesheets/4 \
...
```
Or in Claude Code:
```
Check timesheets for last week, Tempo is at https://jira.company.com/rest/tempo-timesheets/4
```
---
## Limits
- Requires `TEMPO_API_TOKEN` — the script stops immediately with a clear error if it is missing.
- Employees with 0 hours may be on leave; the tool does not distinguish between missing logs and approved absence. Review these manually.
- If the period end date is today or in the future, the period is still open and data will be partial.
- Polish holidays only. If your team uses a different country calendar, the `get_polish_holidays` function in `preprocess.py` needs updating.
- Tempo Cloud API v4 only. Tempo Server / Data Center may require a different base URL and authentication method.

View file

@ -1,115 +0,0 @@
---
name: timesheet-checker
description: >
Check timesheet completion for a given period by pulling Tempo worklogs for all reporters
found in Tempo (not just those in employees.md). Flags anyone missing from employees.md
and asks the user to add them. Requires TEMPO_API_TOKEN. Checks against the Polish working
calendar and 8h/day minimum. Produces a completion report ranked by missing hours.
TRIGGER this skill automatically whenever the user mentions checking, reviewing, or auditing
timesheets or timesheet completion — especially when they mention a period like last week,
last month, or a date range. Trigger phrases include: "check timesheets", "timesheet completion",
"review timesheets", "who hasn't logged hours", "timesheet audit", "check logged hours",
"hours reported", or any message asking about timesheet status or missing hours.
---
# Timesheet Checker
Use this skill when someone needs to know if employees have completed their timesheets for a given period. It pulls worklogs from Tempo, computes each person's completion against the Polish working calendar, and ranks employees by missing hours.
**In Claude Code, you run the full pipeline automatically.** The user provides a period — you handle everything else.
## Prerequisites
This skill requires four credentials. The easiest way to provide them is a `.env` file — copy `.env.example` in the skill folder and fill in the values:
```
cp {SKILL_DIR}/.env.example {SKILL_DIR}/.env
# then edit {SKILL_DIR}/.env
```
Or place the `.env` file in the working directory (where you run Claude Code) — that takes priority over the skill folder.
If the user prefers environment variables they can still use `export`, but the `.env` file removes the need to do that every session.
The four credentials:
| Variable | What it is | Where to get it |
|---|---|---|
| `TEMPO_API_TOKEN` | Tempo read token | Tempo → Settings → API Integration |
| `JIRA_BASE_URL` | Your Atlassian URL | The URL you use to open Jira |
| `JIRA_EMAIL` | Your Jira login email | Your Atlassian account email |
| `JIRA_API_TOKEN` | Jira personal API token | [id.atlassian.com](https://id.atlassian.com) → Security → API tokens |
`JIRA_BASE_URL`, `JIRA_EMAIL`, and `JIRA_API_TOKEN` are used to resolve employee email addresses to Jira account IDs before querying Tempo. `TEMPO_API_TOKEN` is used to pull the actual worklogs.
## Setup
Determine `SKILL_DIR` from the location of this file. Set the output directory to `./timesheet-output/` in the current working directory. Create it if it does not exist.
## Load First
Read these reference files:
1. The employees config — check `./employees.md` in the current working directory first; if not present, fall back to `{SKILL_DIR}/references/employees.md`. Tell the user which one is being used.
2. `{SKILL_DIR}/references/output-format.md` — the structure of `timesheet_data.json` and the final report.
## Extracting the period
The user may express the period in various ways. Convert to the CLI format before running:
| User says | CLI value | End date |
|---|---|---|
| "last week" | `last-week` | Last Friday |
| "last month" | `last-month` | Last day of previous month |
| "last year" | `last-year` | 31 December of previous year |
| "yesterday" | `yesterday` | Yesterday |
| "this week", "current week" | `current-week` | Yesterday |
| "this month", "current month" | `current-month` | Yesterday |
| "this year", "current year" | `current-year` | Yesterday |
| "June", "June 2026" | `2026-06-01:2026-06-30` | Explicit |
| Explicit dates | `YYYY-MM-DD:YYYY-MM-DD` | Explicit |
For `current-*` periods the end date is always yesterday — the current in-progress day is never included. If a user asks for "current week" on a Monday (the week only started today), there are no past working days yet; note this to the user.
If the user has not specified a period, ask: "Which period should I check? (e.g. last week, this month, yesterday, or a date range)"
## Workflow
**Step 1 - Fetch and compute.** Run this command with bash:
```bash
python3 {SKILL_DIR}/scripts/preprocess.py \
--period {period} \
--output-dir ./timesheet-output
```
This fetches Polish public holidays from date.nager.at, computes working days, pulls all Tempo worklogs for the period, and writes `timesheet_data.json` to `./timesheet-output/`.
**Step 2 - Interpret.** Read `./timesheet-output/timesheet_data.json`. Note:
- Any employee with `reported_hours: 0` — this may mean vacation, sick leave, or forgotten logging. Mention these explicitly.
- If the period end date is today or in the future, the period is still open — note this so the user knows the data is partial.
- The overall completion percentage and how many employees have gaps.
- **Unknown reporters:** If `unknown_reporters` is non-empty, show the user a table of those people (name, email, hours logged) and ask them to add the missing employees to `employees.md`. Then continue with Steps 34 using only the known employees, noting that the report excludes the unknown reporters until they are added.
**Step 3 - Generate report.** Run this command with bash:
```bash
python3 {SKILL_DIR}/scripts/postprocess.py \
--data ./timesheet-output/timesheet_data.json \
--output-dir ./timesheet-output
```
This writes `timesheet_report.md`.
**Step 4 - Present results.** Show the user:
- The summary block (expected / reported / overall completion %).
- The full employee table from the report.
- Any notable observations from Step 2.
- "Full report saved to `./timesheet-output/timesheet_report.md`."
## Guardrails
- Never invent or estimate hours. All data comes from Tempo.
- Working days are computed from the Polish public holidays API — do not hardcode holiday dates.
- If `TEMPO_API_TOKEN` is not set, stop and ask the user for it. Do not proceed without it.
- If any script step fails, show the error and stop. Do not proceed with partial output.

View file

@ -1,22 +0,0 @@
# Employees
## Defaults
These values apply to every employee unless overridden in the Employees table.
| setting | value |
| --- | --- |
| country | PL |
| expected_hours_per_day | 8 |
## Employees
Add a row for each employee. Use their work email address — the script looks up the Jira account ID automatically. Leave any column blank to inherit the default.
`country` is an ISO 3166-1 alpha-2 code (e.g. `PL`, `GB`, `DE`, `FR`). Public holidays for that country are fetched from date.nager.at. `start_date` is the employee's first working day — leave blank if they were employed for the full period.
| name | email | country | expected_hours_per_day | start_date |
| --- | --- | --- | --- | --- |
| Jan Kowalski | jan.kowalski@company.com | | | |
| Anna Nowak | anna.nowak@company.com | | | |
| Piotr Wiśniewski | piotr.wisniewski@company.com | | | 2026-05-15 |

View file

@ -1,49 +0,0 @@
# Output Format
## timesheet_data.json (written by preprocess.py)
Machine-readable intermediate file with all fetched and computed data.
```json
{
"period": {
"start": "2026-06-16",
"end": "2026-06-20",
"label": "last week",
"working_days": ["2026-06-16", "2026-06-17", "2026-06-18", "2026-06-19", "2026-06-20"],
"working_day_count": 5
},
"summary": {
"total_employees": 3,
"total_expected_hours": 120,
"total_reported_hours": 96.5,
"overall_completion_pct": 80.4,
"employees_fully_complete": 1,
"employees_with_gaps": 2
},
"employees": [
{
"name": "Jan Kowalski",
"account_id": "557058:aaaa...",
"expected_hours": 40,
"reported_hours": 32.5,
"missing_hours": 7.5,
"completion_pct": 81.3
}
]
}
```
Field notes:
- `employees` is sorted by `missing_hours` descending — the employee furthest behind appears first.
- `completion_pct` = `reported_hours / expected_hours × 100`. Can exceed 100 if overtime was logged.
- `missing_hours` = `max(0, expected_hours - reported_hours)`.
- Working days exclude weekends and Polish public holidays fetched from date.nager.at.
## timesheet_report.md (written by postprocess.py)
Manager-readable markdown report containing:
- Period details (dates, working day count)
- Summary block: expected total, reported total, overall completion %
- Employee table ranked by missing hours (most behind first)
- Completion badges: ✓ complete, ⚠ below 75%, ⚠⚠ below 50%

View file

@ -1,20 +0,0 @@
# Employees (sample)
## Defaults
| setting | value |
| --- | --- |
| country | PL |
| expected_hours_per_day | 8 |
## Employees
Replace email addresses with your team's real work emails.
| name | email | country | expected_hours_per_day | start_date |
| --- | --- | --- | --- | --- |
| Jan Kowalski | jan.kowalski@company.com | | | |
| Anna Nowak | anna.nowak@company.com | | | |
| Piotr Wiśniewski | piotr.wisniewski@company.com | | | 2026-05-15 |
| Marta Jabłońska | marta.jablonska@company.com | | | |
| Tom Smith | tom.smith@company.com | GB | | |

View file

@ -1,156 +0,0 @@
#!/usr/bin/env python3
"""Format timesheet completion data into a markdown report."""
from __future__ import annotations
import argparse
import json
import sys
from datetime import datetime, timezone
from pathlib import Path
from typing import Any
def fmt_h(hours: float) -> str:
"""Format hours with one decimal place, dropping .0 for whole numbers."""
if hours == int(hours):
return f"{int(hours)}h"
return f"{hours:.1f}h"
def completion_badge(pct: float) -> str:
if pct >= 100:
return ""
if pct >= 75:
return ""
if pct >= 50:
return ""
return "⚠⚠"
def build_report(data: dict[str, Any]) -> str:
period = data["period"]
summary = data["summary"]
employees = data["employees"]
generated_at = datetime.now(timezone.utc).strftime("%Y-%m-%d %H:%M UTC")
n = summary["total_employees"]
per_person = fmt_h(summary["total_expected_hours"] / n) if n else "0h"
wdays_by_country: dict[str, int] = period.get("working_days_by_country", {})
if len(wdays_by_country) == 1:
country, wdays = next(iter(wdays_by_country.items()))
wdays_line = f"Working days: {wdays} ({country} calendar, public holidays excluded)"
else:
parts = ", ".join(f"{c}: {d}d" for c, d in sorted(wdays_by_country.items()))
wdays_line = f"Working days by country: {parts} (public holidays excluded)"
lines = [
"# Timesheet Completion Report",
"",
f"Generated: {generated_at}",
f"Period: {period['label'].capitalize()} ({period['start']} to {period['end']})",
wdays_line,
"",
"## Summary",
"",
f"- Expected: **{fmt_h(summary['total_expected_hours'])}** ({n} employees × {per_person} avg)",
f"- Reported: **{fmt_h(summary['total_reported_hours'])}**",
f"- Overall completion: **{summary['overall_completion_pct']}%**",
f"- Fully complete: {summary['employees_fully_complete']} of {n} employees",
f"- With gaps: {summary['employees_with_gaps']} employees",
"",
"## Employees (ranked by missing hours)",
"",
"| Employee | Expected | Reported | Completion | Missing |",
"| --- | ---: | ---: | ---: | ---: |",
]
# Determine the most common working-day count to detect deviations
max_wdays = max(wdays_by_country.values()) if wdays_by_country else 0
for emp in employees:
badge = completion_badge(emp["completion_pct"])
missing = fmt_h(emp["missing_hours"]) if emp["missing_hours"] > 0 else ""
pct_str = f"{emp['completion_pct']}% {badge}".strip()
name = emp["name"]
notes = []
eff = emp.get("effective_working_days", max_wdays)
if emp.get("start_date") and eff < max_wdays:
notes.append(f"from {emp['start_date']}")
if emp.get("termination_date") and eff < max_wdays:
notes.append(f"until {emp['termination_date']}")
if len(wdays_by_country) > 1:
notes.append(emp.get("country", ""))
if notes:
name += f" _({', '.join(n for n in notes if n)})_"
lines.append(
f"| {name} | {fmt_h(emp['expected_hours'])} "
f"| {fmt_h(emp['reported_hours'])} | {pct_str} | {missing} |"
)
out_of_range = [
emp for emp in employees
if emp.get("hours_before_start", 0) > 0 or emp.get("hours_after_termination", 0) > 0
]
if out_of_range:
lines += [
"",
"## Hours logged outside employment period",
"",
"These hours are excluded from the completion calculation above.",
"",
"| Employee | Before start date | After termination date |",
"| --- | ---: | ---: |",
]
for emp in out_of_range:
before = fmt_h(emp["hours_before_start"]) if emp.get("hours_before_start", 0) > 0 else ""
after = fmt_h(emp["hours_after_termination"]) if emp.get("hours_after_termination", 0) > 0 else ""
lines.append(f"| {emp['name']} | {before} | {after} |")
countries_str = ", ".join(sorted(wdays_by_country.keys()))
lines += [
"",
"---",
"",
f"_Working days source: [date.nager.at](https://date.nager.at) ({countries_str})._",
"_Hours source: Tempo._",
]
return "\n".join(lines) + "\n"
def run(args: argparse.Namespace) -> int:
data_path = Path(args.data)
if not data_path.exists():
print(f"Error: {data_path} not found. Run preprocess.py first.", file=sys.stderr)
return 1
data = json.loads(data_path.read_text(encoding="utf-8"))
output_dir = Path(args.output_dir)
output_dir.mkdir(parents=True, exist_ok=True)
report = build_report(data)
report_path = output_dir / "timesheet_report.md"
report_path.write_text(report, encoding="utf-8")
print(f"Wrote {report_path}")
return 0
def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
description="Format timesheet completion data into a markdown report."
)
parser.add_argument(
"--data",
required=True,
help="Path to timesheet_data.json produced by preprocess.py.",
)
parser.add_argument(
"--output-dir",
required=True,
help="Directory for timesheet_report.md.",
)
return parser
if __name__ == "__main__":
sys.exit(run(build_parser().parse_args()))

View file

@ -1,649 +0,0 @@
#!/usr/bin/env python3
"""Fetch Tempo worklogs and compute timesheet completion for a given period."""
from __future__ import annotations
import argparse
import base64
import json
import os
import re
import sys
import urllib.error
import urllib.request
from datetime import date, timedelta
from pathlib import Path
from typing import Any
# ── .env loading ─────────────────────────────────────────────────────────────
def load_dotenv(*paths: Path) -> Path | None:
"""Load the first .env file found from the given paths.
Sets missing keys into os.environ (existing env vars take priority).
Returns the path that was loaded, or None if none were found.
"""
for path in paths:
if not path.exists():
continue
for raw in path.read_text(encoding="utf-8").splitlines():
line = raw.strip()
if not line or line.startswith("#") or "=" not in line:
continue
key, _, value = line.partition("=")
key = key.strip()
value = value.strip().strip('"').strip("'")
if key and key not in os.environ:
os.environ[key] = value
return path
return None
# ── Utilities ─────────────────────────────────────────────────────────────────
def clean_text(value: Any) -> str:
return re.sub(r"\s+", " ", str(value or "")).strip()
def normalize_key(value: Any) -> str:
return re.sub(r"[^a-z0-9]+", "_", clean_text(value).lower()).strip("_")
# ── Period parsing ────────────────────────────────────────────────────────────
def parse_period(period_str: str, today: date | None = None) -> tuple[date, date, str]:
"""Return (start, end, human_label) for the given period string.
Supported keywords:
last-week Previous MondayFriday
last-month Previous calendar month
last-year Previous calendar year
yesterday Yesterday only
current-week Monday of this week to yesterday
current-month 1st of this month to yesterday
current-year 1st January of this year to yesterday
YYYY-MM-DD:YYYY-MM-DD Explicit date range
For current-* periods the end date is always yesterday so the current
in-progress day is excluded. If yesterday falls before the period start
(e.g. calling current-week on a Monday), end < start and the period
will contain zero working days a warning is printed by run().
"""
if today is None:
today = date.today()
yesterday = today - timedelta(days=1)
s = period_str.strip().lower()
if s == "last-week":
last_monday = today - timedelta(days=today.weekday() + 7)
return last_monday, last_monday + timedelta(days=4), "last week"
if s == "last-month":
end = today.replace(day=1) - timedelta(days=1)
return end.replace(day=1), end, "last month"
if s == "last-year":
return date(today.year - 1, 1, 1), date(today.year - 1, 12, 31), "last year"
if s == "yesterday":
return yesterday, yesterday, "yesterday"
if s == "current-week":
start = today - timedelta(days=today.weekday())
return start, yesterday, "current week to date"
if s == "current-month":
return today.replace(day=1), yesterday, "current month to date"
if s == "current-year":
return date(today.year, 1, 1), yesterday, "current year to date"
if ":" in s:
parts = s.split(":", 1)
start = date.fromisoformat(parts[0].strip())
end = date.fromisoformat(parts[1].strip())
return start, end, f"{start} to {end}"
raise ValueError(
f"Unknown period '{period_str}'. "
f"Use last-week, last-month, last-year, yesterday, "
f"current-week, current-month, current-year, or YYYY-MM-DD:YYYY-MM-DD."
)
# ── Public holidays ───────────────────────────────────────────────────────────
def get_public_holidays(country: str, years: set[int]) -> set[date]:
"""Fetch public holidays from date.nager.at for the given country and years.
country is an ISO 3166-1 alpha-2 code (e.g. PL, GB, DE, FR).
Returns an empty set with a warning if the country is not supported.
"""
holidays: set[date] = set()
for year in sorted(years):
url = f"https://date.nager.at/api/v3/PublicHolidays/{year}/{country.upper()}"
req = urllib.request.Request(url, headers={"User-Agent": "timesheet-checker/1.0"})
try:
with urllib.request.urlopen(req, timeout=10) as resp:
items = json.loads(resp.read())
year_holidays = {date.fromisoformat(item["date"]) for item in items}
holidays |= year_holidays
print(f" {country.upper()} {year}: {len(year_holidays)} public holidays")
except urllib.error.HTTPError as exc:
if exc.code == 404:
print(
f"Warning: country '{country.upper()}' not supported by date.nager.at "
f"— no holidays will be excluded for {year}.",
file=sys.stderr,
)
else:
print(f"Warning: could not fetch holidays for {country.upper()} {year}: {exc}", file=sys.stderr)
except urllib.error.URLError as exc:
print(f"Warning: could not fetch holidays for {country.upper()} {year}: {exc}", file=sys.stderr)
return holidays
def get_working_days(start: date, end: date, holidays: set[date]) -> list[date]:
"""Return weekdays (MonFri) in [start, end] that are not public holidays."""
result = []
current = start
while current <= end:
if current.weekday() < 5 and current not in holidays:
result.append(current)
current += timedelta(days=1)
return result
# ── Employees config ──────────────────────────────────────────────────────────
def is_table_row(line: str) -> bool:
s = line.strip()
return s.startswith("|") and s.endswith("|")
def split_table_row(line: str) -> list[str]:
s = line.strip().lstrip("|").rstrip("|")
return [c.strip() for c in s.split("|")]
def is_separator_row(cells: list[str]) -> bool:
return all(re.fullmatch(r":?-{3,}:?", c.strip()) for c in cells)
def parse_markdown_tables(path: Path) -> list[dict[str, Any]]:
lines = path.read_text(encoding="utf-8").splitlines()
tables: list[dict[str, Any]] = []
i = 0
while i < len(lines):
if not is_table_row(lines[i]):
i += 1
continue
block: list[str] = []
while i < len(lines) and is_table_row(lines[i]):
block.append(lines[i])
i += 1
if len(block) < 3:
continue
headers = [normalize_key(c) for c in split_table_row(block[0])]
if not is_separator_row(split_table_row(block[1])):
continue
rows = []
for line in block[2:]:
vals = split_table_row(line)
if len(vals) != len(headers):
raise ValueError(f"Malformed table row in {path}: {line!r}")
rows.append(dict(zip(headers, vals)))
tables.append({"headers": headers, "rows": rows})
return tables
def find_table(tables: list[dict[str, Any]], required_headers: set[str]) -> list[dict[str, str]]:
for t in tables:
if required_headers.issubset(set(t["headers"])):
return t["rows"]
raise ValueError(f"No table with headers {sorted(required_headers)} found in employees file.")
def load_employees(path: Path) -> list[dict[str, Any]]:
tables = parse_markdown_tables(path)
# Load defaults (optional table with setting/value columns).
defaults: dict[str, Any] = {"country": "PL", "expected_hours_per_day": 8.0}
try:
for row in find_table(tables, {"setting", "value"}):
key = normalize_key(row["setting"])
val = clean_text(row["value"])
if not val:
continue
if key == "country":
defaults["country"] = val.upper()
elif key == "expected_hours_per_day":
try:
defaults["expected_hours_per_day"] = float(val)
except ValueError:
print(f"Warning: invalid default expected_hours_per_day '{val}' — using 8.", file=sys.stderr)
except ValueError:
pass # no defaults table; built-in defaults apply
rows = find_table(tables, {"name", "email"})
employees = []
for row in rows:
name = clean_text(row["name"])
email = clean_text(row["email"]).lower()
if not name or not email:
continue
raw_h = clean_text(row.get("expected_hours_per_day", ""))
try:
expected_h = float(raw_h) if raw_h else defaults["expected_hours_per_day"]
except ValueError:
expected_h = defaults["expected_hours_per_day"]
raw_country = clean_text(row.get("country", "")).upper()
country = raw_country if raw_country else defaults["country"]
raw_start = clean_text(row.get("start_date", ""))
start_date: date | None = None
if raw_start:
try:
start_date = date.fromisoformat(raw_start)
except ValueError:
print(
f"Warning: could not parse start_date '{raw_start}' for {name} — ignoring.",
file=sys.stderr,
)
raw_term = clean_text(row.get("termination_date", ""))
termination_date: date | None = None
if raw_term:
try:
termination_date = date.fromisoformat(raw_term)
except ValueError:
print(
f"Warning: could not parse termination_date '{raw_term}' for {name} — ignoring.",
file=sys.stderr,
)
employees.append({
"name": name,
"email": email,
"expected_hours_per_day": expected_h,
"country": country,
"start_date": start_date,
"termination_date": termination_date,
})
if not employees:
raise ValueError(f"No employees found in {path}. Check the table format.")
return employees
# ── Jira user lookup ──────────────────────────────────────────────────────────
def resolve_account_ids(
emails: list[str],
jira_base_url: str,
jira_email: str,
jira_token: str,
) -> dict[str, str]:
"""Return {email: account_id} by searching Jira's user API for each address."""
credentials = base64.b64encode(f"{jira_email}:{jira_token}".encode()).decode()
headers = {"Authorization": f"Basic {credentials}", "Accept": "application/json"}
resolved: dict[str, str] = {}
not_found: list[str] = []
for email in emails:
url = f"{jira_base_url.rstrip('/')}/rest/api/3/user/search?query={email}&maxResults=10"
req = urllib.request.Request(url, headers=headers)
try:
with urllib.request.urlopen(req, timeout=10) as resp:
users = json.loads(resp.read())
except urllib.error.HTTPError as exc:
if exc.code == 401:
print(
"Error: Jira API returned 401 Unauthorized.\n"
"Check JIRA_EMAIL and JIRA_API_TOKEN (generate at id.atlassian.com).",
file=sys.stderr,
)
sys.exit(1)
body = exc.read().decode("utf-8", errors="replace")
print(f"Error: Jira API returned {exc.code} for {email}: {body}", file=sys.stderr)
sys.exit(1)
match = next(
(u for u in users if u.get("emailAddress", "").lower() == email.lower()),
None,
)
if match:
resolved[email] = match["accountId"]
print(f" {email}{match['accountId']}")
else:
not_found.append(email)
if not_found:
print(
f"Error: Could not find Jira account for: {', '.join(not_found)}\n"
"Check the email addresses in employees.md or your JIRA_BASE_URL.",
file=sys.stderr,
)
sys.exit(1)
return resolved
# ── Tempo API ─────────────────────────────────────────────────────────────────
def fetch_tempo_worklogs(
start: date,
end: date,
token: str,
base_url: str,
) -> dict[str, dict[str, float]]:
"""Return {account_id: {date_str: hours}} for ALL reporters in the period.
Paginates automatically. Date strings are ISO format (YYYY-MM-DD).
"""
totals: dict[str, dict[str, float]] = {}
headers = {"Authorization": f"Bearer {token}", "Accept": "application/json"}
next_url: str | None = f"{base_url}/worklogs?from={start}&to={end}&limit=1000"
page = 0
while next_url:
page += 1
req = urllib.request.Request(next_url, headers=headers)
try:
with urllib.request.urlopen(req, timeout=30) as resp:
data = json.loads(resp.read())
except urllib.error.HTTPError as exc:
body = exc.read().decode("utf-8", errors="replace")
if exc.code == 401:
print(
"Error: Tempo API returned 401 Unauthorized. Check your TEMPO_API_TOKEN.",
file=sys.stderr,
)
else:
print(f"Error: Tempo API returned {exc.code}: {body}", file=sys.stderr)
sys.exit(1)
for worklog in data.get("results", []):
aid = worklog.get("author", {}).get("accountId", "")
date_str = worklog.get("startDate", "")
if aid and date_str:
if aid not in totals:
totals[aid] = {}
totals[aid][date_str] = (
totals[aid].get(date_str, 0.0) + worklog.get("timeSpentSeconds", 0) / 3600
)
next_url = data.get("metadata", {}).get("next")
print(f"Fetched worklogs (pages: {page})")
return totals
def resolve_account_info(
account_ids: list[str],
jira_base_url: str,
jira_email: str,
jira_token: str,
) -> dict[str, dict]:
"""Return {account_id: {display_name, email}} for the given Jira account IDs."""
credentials = base64.b64encode(f"{jira_email}:{jira_token}".encode()).decode()
headers = {"Authorization": f"Basic {credentials}", "Accept": "application/json"}
result: dict[str, dict] = {}
for aid in account_ids:
url = f"{jira_base_url.rstrip('/')}/rest/api/3/user?accountId={aid}"
req = urllib.request.Request(url, headers=headers)
try:
with urllib.request.urlopen(req, timeout=10) as resp:
user = json.loads(resp.read())
result[aid] = {
"display_name": user.get("displayName", ""),
"email": user.get("emailAddress", ""),
}
except (urllib.error.HTTPError, urllib.error.URLError):
result[aid] = {"display_name": f"Unknown ({aid})", "email": ""}
return result
# ── Stats computation ─────────────────────────────────────────────────────────
def compute_stats(
employees: list[dict[str, Any]],
period_start: date,
period_end: date,
holidays_by_country: dict[str, set[date]],
worklogs: dict[str, dict[str, float]],
) -> list[dict[str, Any]]:
results = []
for emp in employees:
holidays = holidays_by_country.get(emp["country"], set())
all_working_days = get_working_days(period_start, period_end, holidays)
emp_start: date | None = emp.get("start_date")
emp_term: date | None = emp.get("termination_date")
emp_working_days = all_working_days
if emp_start:
emp_working_days = [d for d in emp_working_days if d >= emp_start]
if emp_term:
emp_working_days = [d for d in emp_working_days if d <= emp_term]
expected_h = emp["expected_hours_per_day"] * len(emp_working_days)
# Partition logged hours into in-range vs out-of-range
date_logs = worklogs.get(emp["account_id"], {})
hours_in_range = 0.0
hours_before_start = 0.0
hours_after_termination = 0.0
for date_str, h in date_logs.items():
d = date.fromisoformat(date_str)
if emp_start and d < emp_start:
hours_before_start += h
elif emp_term and d > emp_term:
hours_after_termination += h
else:
hours_in_range += h
reported_h = round(hours_in_range, 2)
hours_before_start = round(hours_before_start, 2)
hours_after_termination = round(hours_after_termination, 2)
missing_h = round(max(0.0, expected_h - reported_h), 2)
pct = round((reported_h / expected_h * 100) if expected_h > 0 else 100.0, 1)
results.append({
"name": emp["name"],
"account_id": emp["account_id"],
"country": emp["country"],
"start_date": emp_start.isoformat() if emp_start else None,
"termination_date": emp_term.isoformat() if emp_term else None,
"effective_working_days": len(emp_working_days),
"expected_hours": expected_h,
"reported_hours": reported_h,
"hours_before_start": hours_before_start,
"hours_after_termination": hours_after_termination,
"missing_hours": missing_h,
"completion_pct": pct,
})
return sorted(results, key=lambda r: r["missing_hours"], reverse=True)
# ── Entry point ───────────────────────────────────────────────────────────────
def _resolve_employees_path(cli_value: str | None) -> Path | None:
"""Return the employees file to use, or None if not found.
Search order:
1. Explicit --employees argument (if given)
2. employees.md in the current working directory
3. references/employees.md in the skill folder (next to this script)
"""
if cli_value:
return Path(cli_value)
skill_dir = Path(__file__).parent.parent
for candidate in (
Path.cwd() / "employees.md",
skill_dir / "references" / "employees.md",
):
if candidate.exists():
return candidate
return None
def run(args: argparse.Namespace) -> int:
# Load .env from working directory first, then skill directory as fallback.
# Existing environment variables always take priority over .env values.
skill_dir = Path(__file__).parent.parent
loaded = load_dotenv(Path.cwd() / ".env", skill_dir / ".env")
if loaded:
print(f"Loaded credentials from {loaded}")
tempo_token = args.tempo_token or os.environ.get("TEMPO_API_TOKEN", "")
jira_base_url = args.jira_base_url or os.environ.get("JIRA_BASE_URL", "")
jira_email = args.jira_email or os.environ.get("JIRA_EMAIL", "")
jira_token = args.jira_token or os.environ.get("JIRA_API_TOKEN", "")
missing = [
name for name, val in [
("TEMPO_API_TOKEN", tempo_token),
("JIRA_BASE_URL", jira_base_url),
("JIRA_EMAIL", jira_email),
("JIRA_API_TOKEN", jira_token),
] if not val
]
if missing:
print(
f"Error: missing required credentials: {', '.join(missing)}\n"
"Set them as environment variables or pass via CLI flags.\n"
"See README.md for how to obtain each credential.",
file=sys.stderr,
)
return 1
start, end, label = parse_period(args.period)
print(f"Period: {label} ({start} to {end})")
employees_path = _resolve_employees_path(args.employees)
if employees_path is None:
skill_dir = Path(__file__).parent.parent
print(
"Error: No employees.md found.\n"
f"Looked in:\n"
f" {Path.cwd() / 'employees.md'}\n"
f" {skill_dir / 'references' / 'employees.md'}\n"
"Create one in your project directory or pass --employees PATH.",
file=sys.stderr,
)
return 1
employees = load_employees(employees_path)
print(f"Employees loaded: {len(employees)} (from {employees_path})")
countries = {e["country"] for e in employees}
years = {start.year, end.year}
print(f"Fetching public holidays from date.nager.at for: {', '.join(sorted(countries))} ...")
holidays_by_country = {c: get_public_holidays(c, years) for c in countries}
print("Resolving Jira account IDs from email addresses...")
email_to_id = resolve_account_ids(
[e["email"] for e in employees], jira_base_url, jira_email, jira_token
)
for emp in employees:
emp["account_id"] = email_to_id[emp["email"]]
known_ids = {e["account_id"] for e in employees}
print(f"Fetching Tempo worklogs from {args.tempo_base_url}...")
all_worklogs = fetch_tempo_worklogs(start, end, tempo_token, args.tempo_base_url)
# Detect reporters not in employees.md
unknown_ids = [aid for aid in all_worklogs if aid not in known_ids]
unknown_reporters: list[dict] = []
if unknown_ids:
print(f"Resolving {len(unknown_ids)} unknown reporter(s) from Jira...")
account_info = resolve_account_info(unknown_ids, jira_base_url, jira_email, jira_token)
for aid in unknown_ids:
info = account_info[aid]
total_h = round(sum(all_worklogs[aid].values()), 2)
unknown_reporters.append({
"account_id": aid,
"display_name": info["display_name"],
"email": info["email"],
"reported_hours": total_h,
})
print(f" Not in employees.md: {info['display_name']} ({info['email']}) — {total_h:.1f}h")
unknown_reporters.sort(key=lambda r: r["reported_hours"], reverse=True)
stats = compute_stats(employees, start, end, holidays_by_country, all_worklogs)
total_expected = sum(r["expected_hours"] for r in stats)
total_reported = round(sum(r["reported_hours"] for r in stats), 2)
overall_pct = round((total_reported / total_expected * 100) if total_expected > 0 else 100.0, 1)
# Annotate with email for readability
for stat in stats:
emp = next(e for e in employees if e["account_id"] == stat["account_id"])
stat["email"] = emp["email"]
working_days_by_country = {
c: len(get_working_days(start, end, holidays_by_country[c]))
for c in sorted(countries)
}
output = {
"period": {
"start": start.isoformat(),
"end": end.isoformat(),
"label": label,
"working_days_by_country": working_days_by_country,
},
"summary": {
"total_employees": len(stats),
"total_expected_hours": total_expected,
"total_reported_hours": total_reported,
"overall_completion_pct": overall_pct,
"employees_fully_complete": sum(1 for r in stats if r["completion_pct"] >= 100),
"employees_with_gaps": sum(1 for r in stats if r["completion_pct"] < 100),
},
"employees": stats,
"unknown_reporters": unknown_reporters,
}
output_dir = Path(args.output_dir)
output_dir.mkdir(parents=True, exist_ok=True)
out_path = output_dir / "timesheet_data.json"
out_path.write_text(json.dumps(output, indent=2), encoding="utf-8")
print(f"Wrote {out_path}")
return 0
def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
description="Fetch Tempo worklogs and compute timesheet completion."
)
parser.add_argument(
"--period",
required=True,
help="Period to check: 'last-week', 'last-month', or 'YYYY-MM-DD:YYYY-MM-DD'.",
)
parser.add_argument(
"--employees",
help=(
"Path to the employees markdown config file. "
"If omitted, looks for employees.md in the current directory, "
"then references/employees.md in the skill folder."
),
)
parser.add_argument(
"--output-dir",
required=True,
help="Directory for timesheet_data.json.",
)
parser.add_argument(
"--tempo-token",
help="Tempo API token. Defaults to TEMPO_API_TOKEN env var.",
)
parser.add_argument(
"--tempo-base-url",
default="https://api.tempo.io/4",
help="Tempo API base URL. Default: https://api.tempo.io/4 (Tempo Cloud).",
)
parser.add_argument(
"--jira-base-url",
help="Jira base URL, e.g. https://company.atlassian.net. Defaults to JIRA_BASE_URL env var.",
)
parser.add_argument(
"--jira-email",
help="Email address used to authenticate with Jira. Defaults to JIRA_EMAIL env var.",
)
parser.add_argument(
"--jira-token",
help="Jira personal API token. Defaults to JIRA_API_TOKEN env var.",
)
return parser
if __name__ == "__main__":
sys.exit(run(build_parser().parse_args()))

View file

@ -1,302 +0,0 @@
---
name: transcript-speaker-fill
description: Recover real speaker names in a Fireflies transcript that only has generic "Speaker N" labels, by fuzzy-matching timestamps against a second transcript of the same meeting that has real names but broken/unusable text (commonly caused by wrong-language detection). Use when the user has two Fireflies exports of one meeting - one with correct text but no speaker names, one with real names but garbled text - and wants the names filled into the good transcript. Trigger phrases: "fill in the speakers", "match speakers by timestamp", "recover speaker names", "one transcript is missing speakers and the other has the wrong language".
---
# Transcript speaker fill skill
## Purpose
Fireflies occasionally produces two broken outcomes for the same meeting:
one export has correct transcript text but generic `Speaker 1`, `Speaker 2`,
... labels (diarization worked, but names were never resolved or the
Fireflies bot lost the participant roster); another export - often from a
retry after fixing the language setting - has real speaker names but
garbage text (wrong language was detected, so the words are nonsense, but
the underlying speaker diarization and timestamps are still meaningful).
This skill cross-references the two: it can't read the garbled text, but it
can compare *when* each speaker was talking, and use that timing to guess
which generic label corresponds to which real name in the good transcript.
This is entirely mechanical - a bundled Python script does the parsing,
offset detection, and voting. Nothing here needs model judgment except
interpreting the final report and deciding whether the result is trustworthy
enough to use.
## Important: this is best-effort, not a solved match
Be direct with the user about this before and after running it:
- **Some speakers may be structurally unresolvable.** If the target
transcript has more distinct unnamed speakers than the broken transcript
has distinct real names, some target speakers simply aren't captured
with a real name *anywhere* in the broken file (they may have joined
late, or their voice wasn't separated out in that broken run). No amount
of tuning fixes this - the script detects and reports this gap explicitly,
but can still produce a confident-*looking* wrong answer for an affected
label, because it has no way to know a name is entirely absent from the
candidate pool.
- **Timestamps are start-of-utterance markers from two independent
diarization runs**, not a shared clock - they can disagree by several
seconds even for a genuine match, and the two files may have a constant
offset if the bots didn't start recording at exactly the same instant.
The script searches for that offset automatically; it can still get it
slightly wrong in a noisy recording.
- **Always spot-check the result** - especially any label resolved with
fewer than ~4-5 votes or under ~65% confidence, and *especially* any
label affected by the roster-gap warning. Read a couple of the actual
lines attributed to a resolved name and sanity-check against tone/content
(a name attributed to a monologue about hairdressers when the person is
known to talk mostly about delivery process, say, is a red flag).
## How to run this skill
This is a two-pass flow: a dry-run preview first, a write-to-disk second,
only after the matches have been confirmed with the user.
### Pass 1 - preview (dry run, default)
```bash
python3 "<skill-dir>/scripts/fill_speakers.py" "<file-1>" "<file-2>"
```
- `<file-1>` / `<file-2>` - the two transcripts, **in either order**. The
script auto-detects which one is the broken source (real speaker names,
garbled/wrong-language text) and which is the target (correct text,
generic `Speaker N` labels) by comparing how much of each file is still
labeled with generic `Speaker N` names - it does **not** rely on file
size or on which argument came first. It prints which role it assigned
to each file near the top of the output - check that this matches what
you'd expect from the filenames/content before trusting the rest of the
report. If the two files are too similar to tell apart (rare), it exits
with an error instead of guessing - open both and check by eye: the
broken one reads as nonsense/wrong language but has real names; the
target reads correctly but has `Speaker N` labels.
Resolve `<skill-dir>` to this skill's own directory.
With no `--apply` flag, the script **only prints the matching report** -
merges, roster-gap warning, offset, the per-label resolution table, and
any unresolved labels with text snippets. It does not touch disk yet.
Present this report to the user (see "After the preview" below) and get
their explicit confirmation - or their corrections via `--manual` - before
moving to pass 2.
### Pass 2 - apply (only after confirmation)
Once the user has confirmed the matches (and supplied any `--manual`
overrides for gaps or corrections), re-run the exact same command with
`--apply` added:
```bash
python3 "<skill-dir>/scripts/fill_speakers.py" "<file-1>" "<file-2>" --apply [--manual "Speaker N=Real Name" ...]
```
Only this run writes the output file. It never modifies either input
file - it writes a new file next to the target, named
`<target-stem>-speakers-filled.md` by default (override with `--output
<path>`), with every generic label it could confidently resolve (or that
was given via `--manual`) replaced by the real name throughout the
target - one resolution per label, applied consistently everywhere that
label appears, not line-by-line guessing.
Do not pass `--apply` on the first run, and do not treat pass 1's report
as final - it is a proposal for the user to react to, not a completed
action.
Useful tuning flags if the default report looks too conservative or too
noisy:
- `--tolerance <seconds>` (default 15) - how close two timestamps must be
to count as a candidate match.
- `--min-votes <n>` (default 2) / `--min-confidence <0-1>` (default 0.5) -
how many matches, and what vote share, a label needs before it gets
resolved instead of left as `Speaker N`. Raise both for a more
conservative (fewer, more trustworthy) result.
- `--offset-search <seconds>` (default 90) - how wide a window to search
for a constant clock offset between the two recordings.
- `--no-normalize` - by default, the broken file's own duplicate-diarization
name variants are merged before voting (see next section) - most
commonly a trailing digit Fireflies adds when it's unsure two segments
are the same voice cluster (e.g. `Robert Drazkowski` and
`Robert Drazkowski1` collapse into one candidate name, `Robert
Drazkowski`). Disable this only if that assumption is wrong for a given
pair of files - e.g. the trailing digit is genuinely disambiguating two
different real people who happen to share a name, which would be
unusual but isn't impossible.
- `--manual "Speaker N=Real Name"` (repeatable) - force a specific label to
a specific name directly, bypassing matching entirely. This is how you
apply names the user supplies for labels the automatic pass couldn't
resolve (see "After the preview" below) - it overrides any automatic
result for that label, confident or not, and works even for a label
that had zero automatic matches at all.
- **Partial names are matched against the broken file's roster
automatically.** If the user only gives a first name (or any partial
string) and it matches exactly one real name already appearing in the
broken file - e.g. `--manual "Speaker 3=Dawid"` when the roster
contains `Dawid Cieślicki` and no one else called Dawid - the script
expands it to the full roster name and prints a `NOTE (Speaker 3): ...`
line saying so. Relay that note to the user so they know which full
name actually got applied. If the partial name matches *more than one*
roster name, it's genuinely ambiguous - the script keeps the literal
string as given (does not guess) and prints a warning listing every
candidate it could have meant; relay that to the user and ask them to
supply the full name instead. If it matches *no one* in the roster, the
literal name is used as-is with no note - that's the expected, normal
case for a real participant who was never captured with a name
anywhere in the broken file either (the roster-gap scenario), not an
error.
- `--examples <n>` (default 5) - how many timestamped example lines to
print per label (spread across its timeline) for the user to spot-check
against the actual recording. Raise it if a label needs more coverage
before the user is comfortable confirming it.
- `--recording-url <url>` - a link to the target meeting's recording (e.g.
a Fireflies share URL). If given, it's printed once at the top and
repeated under every label next to its example lines, so a reviewer has
one click to the recording right where they need it. There's no verified
way to encode an exact-timestamp deep link for Fireflies (their own UI
has a "copy link to this moment" feature, but the query parameter it
produces isn't publicly documented), so this links to the recording
itself - the user still scrubs to each example's timestamp manually.
## Same speaker, two labels in the broken file
Fireflies sometimes emits two different name strings for what is actually
one person, when its diarization isn't confident two speech segments are
the same voice cluster - most visibly a trailing digit appended to an
otherwise-identical name (`Robert Drazkowski` / `Robert Drazkowski1`). Left
unhandled, this would split one real person's votes across two candidate
names and could prevent either from reaching the resolution threshold, or
worse, cause the script to treat them as two different (wrong) people.
The script merges these automatically before voting (`normalize_name`,
default on) and **prints exactly what it merged** near the top of its
output - always check that block. If it merged something that was actually
two different people, or missed a variant that isn't a bare trailing digit
(e.g. `Name (2)` or `Name_2` - the normalizer also handles these, but a
truly unusual format might slip through), re-run with `--no-normalize` and
handle that pair of names manually via `--manual` instead.
## After the preview (pass 1)
Report to the user, based on the script's own printed output, **before**
ever running pass 2:
- Which file it auto-detected as broken vs target - flag it if this looks
wrong given the filenames/content.
- Any merged duplicate-diarization labels (so they can sanity-check the
merge was correct, not two different people collapsed into one).
- The roster-gap warning, if any (how many target speakers can't
structurally be resolved).
- The offset it settled on.
- The per-label resolution table (name/UNRESOLVED, vote count, confidence)
- present this as proposed matches for the user to accept, not as a
done deal.
- **Any other candidates for a label, most to least probable.** Whenever a
label had votes for more than one real name, the script prints a second
line under it - `other candidates, most to least probable: ...` for a
resolved label, `all candidates, most to least probable: ...` for an
unresolved one - each with its own vote count and vote share. Always
relay this ranked list, not just the winning name, especially when the
top two candidates are close in vote share (e.g. 45% vs 40%): that's a
near-tie, not a confident resolution, and the user may recognize the
second-place name as the right one from the snippets.
- **The example lines printed under every label** - the script picks up to
`--examples` (default 5) lines per label, spread across that label's full
timeline rather than clustered at the start, and prints each with its
timestamp *in the target recording's own timeline* (not offset-adjusted -
these are the timestamps to scrub to in the actual recording/video, since
that's what the user has playback access to). This applies to every
label, not just unresolved ones - relay them for the resolved labels too,
and explicitly suggest the user jump to a couple of these timestamps in
the recording and confirm by ear who's actually speaking, especially for
anything under ~70% confidence. This is the concrete way to turn "the
vote count says X" into "I checked and it's actually X" - don't skip
offering it just because a label came back resolved. Raise `--examples`
if the user wants more per label to check. If you have a link to the
recording, pass it via `--recording-url` so it's printed alongside every
label's examples - don't make the user go find the meeting themselves.
- **When a `--manual` override is given, the automatic guess is still shown
underneath it, not discarded** - the header line says whether the manual
name *agrees with* the automatic guess, *overrides* it (naming what the
automatic pass would have picked instead, and at what confidence), or
fills a gap the automatic pass left unresolved. Relay this distinction -
an override that contradicts a high-confidence automatic guess is worth
flagging back to the user as a "you sure?" before applying, whereas one
that just fills an unresolved gap or agrees with the automatic guess
needs no extra scrutiny.
- The minority-vote lines flagged for manual review, if any.
- **If the roster-gap warning fired, look for suspicious patterns across
multiple labels** before taking individual resolutions at face value -
e.g. two different labels both resolving to the same real name, each at
middling confidence with the same runner-up(s), is a sign that one of
them is actually an unnamed real participant being misattributed, not
genuinely two clusters of the same person. Flag this pattern explicitly
to the user rather than reporting each label's line as independently
fine.
Then explicitly recommend spot-checking the lowest-confidence resolutions
against actual dialogue content - and, now that timestamps are available
for every label, against the actual recording audio - before accepting
them. Do not present the preview as a finished, verified transcript.
**Always relay the script's "Could not match" section and actually ask
the user about it** - don't just print it and move on. For each
unresolved label, show the line count and its example timestamps, then ask
something like: "I couldn't match Speaker 2, 4, and 9 - here's what each
said, with timestamps to check in the recording [examples]. Do you know
who any of these are?"
**Wait for explicit confirmation before running pass 2 (`--apply`).** The
user needs to either:
- confirm the proposed matches look right, and/or
- supply real names for any unresolved gaps, and/or
- correct any match that looks wrong (even a "resolved" one they don't
trust).
Fold whatever they give you into `--manual "Speaker N=Name"` flags (one
per label) on the pass-2 run - this overrides the automatic result for
that label, confident or not. Then re-run with `--apply` and show the
updated report. Don't hand-edit the output file directly, since a re-run
keeps the automatic resolutions and the merge/roster reporting consistent
with the final file.
If the user wants a stricter or looser automatic pass instead, re-run
pass 1 (still without `--apply`) with adjusted
`--tolerance`/`--min-votes`/`--min-confidence` and present the new
preview before applying anything.
## Edge cases
- **Neither file matches the expected `**Speaker** *[MM:SS]*: text` format**
(e.g. it's an `.mhtml` capture, not an extracted `.md`) - point the user
at the `extract-transcript` skill first to get a proper Markdown export.
- **A generic label appears in a burst with no broken-file activity nearby
at all** (e.g. everyone else was silent while this person spoke for a
while) - it will correctly come back `UNRESOLVED (no timestamp within
tolerance found at all)` rather than a forced guess.
- **The target file already has some real names mixed with `Speaker N`
labels** (partial resolution done elsewhere) - only the `Speaker N`
entries are touched; already-named lines are left exactly as they are.
This doesn't confuse auto-detection either, since it compares the
*share* of generic labels between the two files, not just their
presence.
- **Auto-detection can't tell the files apart** (exactly equal share of
generic labels in both, e.g. both 0% or both 100%) - the script exits
with an error rather than guessing, since swapping argument order
wouldn't change the outcome either. Inspect both files by eye instead -
it likely means one file isn't in the format expected, or this isn't
actually a broken/target pair.
- **Re-running after tuning flags or `--manual` overrides** - always safe
in either pass. Without `--apply` nothing is ever written, and with
`--apply` each run's output filename defaults to the same path
(overwritten on re-run, not accumulated).
- **`--manual` references a label that doesn't exist in the target** (typo,
or a label that already has a real name) - the script warns and ignores
it rather than silently doing nothing; check the warning if a manual
override doesn't seem to have taken effect.
- **A `--manual` partial name matches more than one roster name** (e.g. two
different real people in the broken file share a first name) - the
script refuses to guess, uses the literal string as given, and prints a
warning listing every candidate it could have meant. Relay this to the
user and get the full name before applying, rather than letting the
ambiguous literal string silently become the final label.

View file

@ -1,480 +0,0 @@
#!/usr/bin/env python3
"""Fill in real speaker names in a Fireflies-style transcript by fuzzy-matching
timestamps against a second transcript of the same meeting that has real
speaker names but broken/unusable text (e.g. wrong language was detected,
so the words are garbage but the diarization + speaker labels are fine).
Both inputs are expected in Fireflies' Markdown export shape:
**Speaker Name** *[MM:SS]*: text
The two file arguments can be given in either order - whichever one has a
higher share of generic "Speaker N" labels is auto-detected as the target
to fill in; the other is treated as the broken source of real names.
By default this is a dry run: it only prints the matching report. Nothing
is written until it is re-run with --apply, so the report can be reviewed
(and --manual overrides added for any gaps) before anything touches disk.
Deterministic, no model reasoning involved - see the accompanying SKILL.md
for when/how to invoke this.
"""
import argparse
import re
import sys
from pathlib import Path
from collections import defaultdict, Counter
LINE_RE = re.compile(
r'^\*\*(?P<speaker>[^*]+)\*\*\s*\*\[(?P<ts>\d{1,2}:\d{2}(?::\d{2})?)\]\*:\s*(?P<text>.*)$'
)
GENERIC_RE = re.compile(r'^Speaker\s+\d+$')
def parse_ts(ts):
parts = [int(p) for p in ts.split(':')]
if len(parts) == 2:
m, s = parts
return m * 60 + s
h, m, s = parts
return h * 3600 + m * 60 + s
def normalize_name(name):
"""Fireflies sometimes labels the same real person with two different
strings when it isn't sure two segments are the same voice cluster -
most commonly a trailing digit appended to an otherwise-identical name
(e.g. "Robert Drazkowski" and "Robert Drazkowski1" are the same person).
Strip that suffix so both collapse into one candidate name for voting.
Also strips a trailing " 2", "(2)", "_2" etc. in case Fireflies uses one
of those variants instead of a bare digit."""
cleaned = re.sub(r'[\s_]*\(?\d+\)?$', '', name).strip()
return cleaned if cleaned else name.strip()
def parse_manual_overrides(pairs):
"""Parse repeated --manual "Speaker N=Real Name" arguments into a dict."""
overrides = {}
for pair in pairs or []:
if '=' not in pair:
sys.exit(f"--manual expects 'Speaker N=Real Name', got: {pair!r}")
label, name = pair.split('=', 1)
overrides[label.strip()] = name.strip()
return overrides
def resolve_partial_name(name, roster):
"""If `name` is a partial reference (e.g. just a first name) to someone
who already appears in `roster` (the broken file's real-name roster),
suggest/expand to the matching full name instead of taking the partial
string literally. Returns (final_name, note_or_None).
Matching is deliberately conservative: an exact match short-circuits
immediately; otherwise a candidate qualifies only if `name` shares a
whole word with it (case-insensitive) or is a substring of it. If more
than one roster name qualifies, this is ambiguous - the literal name is
kept as given rather than guessing, with a warning listing the
candidates so the user can specify which one they meant. If none
qualify, the name is genuinely new (e.g. a real participant who was
never captured with a name anywhere in the broken file) and is used
as-is - that's expected, not an error."""
if name in roster:
return name, None
name_lower = name.strip().lower()
name_tokens = set(name_lower.split())
candidates = []
for full in roster:
full_lower = full.lower()
if name_lower == full_lower:
return full, f"'{name}' matches roster name '{full}' (case-insensitive) - using '{full}'."
full_tokens = set(full_lower.split())
if name_tokens & full_tokens or name_lower in full_lower:
candidates.append(full)
if len(candidates) == 1:
return candidates[0], f"'{name}' looks like a partial name - matched to the only candidate in the roster, '{candidates[0]}'. Using the full name."
if len(candidates) > 1:
return name, (
f"'{name}' is ambiguous - it could refer to any of: {', '.join(candidates)}. "
f"Used literally as given since I can't tell which one you meant - re-run with the full "
f"name to disambiguate if this isn't who you intended."
)
return name, None
def parse_file(path):
entries = []
for i, line in enumerate(Path(path).read_text(encoding='utf-8').splitlines()):
m = LINE_RE.match(line.strip())
if m:
entries.append({
'line_no': i,
'speaker': m.group('speaker').strip(),
'seconds': parse_ts(m.group('ts')),
'text': m.group('text'),
})
return entries
def classify(entries_a, path_a, entries_b, path_b):
"""Decide which of the two parsed transcripts is the 'target' (has
generic Speaker N labels needing real names filled in) and which is
the 'broken' source (has real names already, used only for its
timestamps). Whichever file has a higher share of generic-labeled
lines is the target - the broken source should have few or none,
since its diarization already resolved real names even though its
text is garbled. This replaces any assumption about file size or
argument order."""
def generic_fraction(entries):
if not entries:
return 0.0
generic = sum(1 for e in entries if GENERIC_RE.match(e['speaker']))
return generic / len(entries)
frac_a = generic_fraction(entries_a)
frac_b = generic_fraction(entries_b)
if frac_a == frac_b:
sys.exit(
f"Could not automatically tell which file needs speaker names filled in: "
f"both '{path_a}' and '{path_b}' have the same share of generic 'Speaker N' "
f"labels ({frac_a:.0%}). Check the files by eye - the target should read as "
f"real dialogue with generic labels, the broken source should have real names "
f"but garbled/wrong-language text."
)
if frac_a > frac_b:
return (entries_a, path_a), (entries_b, path_b)
return (entries_b, path_b), (entries_a, path_a)
def proximity_weight(delta, tolerance):
"""1.0 for an exact match, decaying linearly to just above 0 at the
tolerance boundary. A match a few seconds off is real signal; a match
12 seconds off inside a 15s tolerance is mostly noise - weight
accordingly rather than counting both as one equal 'vote'."""
return max(0.0, 1.0 - delta / (tolerance + 1))
def find_best_offset(anchor_entries, target_entries, tolerance, offset_range):
"""Search for a constant clock offset (seconds) between the two
recordings that maximizes total proximity-weighted overlap between
target and anchor timestamps. Handles the two Fireflies bots not
starting at exactly the same instant. Weighted (not a raw count of
"any match within tolerance") so a wide tolerance can't let a wrong
offset win just by picking up many loose, low-quality matches."""
anchor_times = [e['seconds'] for e in anchor_entries]
best_offset, best_score = 0, -1.0
for offset in range(-offset_range, offset_range + 1):
score = 0.0
for t in target_entries:
tt = t['seconds'] + offset
best_delta = min((abs(tt - at) for at in anchor_times), default=tolerance + 1)
score += proximity_weight(best_delta, tolerance)
if score > best_score or (score == best_score and abs(offset) < abs(best_offset)):
best_score, best_offset = score, offset
return best_offset, best_score
def nearest_match(seconds, anchor_entries, tolerance):
best, best_delta = None, tolerance + 1
for a in anchor_entries:
delta = abs(a['seconds'] - seconds)
if delta <= tolerance and delta < best_delta:
best_delta, best = delta, a
return best, best_delta
def sample_examples(entries, n):
"""Pick up to n example entries spread across the full span of entries
(not just the first n) so a spot-check sees variety across the
meeting's timeline rather than one early cluster."""
if not entries:
return []
if len(entries) <= n:
return entries
if n <= 1:
return [entries[0]]
idxs = sorted({round(i * (len(entries) - 1) / (n - 1)) for i in range(n)})
return [entries[i] for i in idxs]
def format_examples(entries, indent=' '):
lines = []
for e in entries:
mm, ss = divmod(e['seconds'], 60)
snippet = e['text'].strip()
if len(snippet) > 90:
snippet = snippet[:90].rstrip() + "..."
lines.append(f"{indent}[{mm:02d}:{ss:02d}] {snippet}")
return "\n".join(lines)
def main():
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument('file_a', help='One of the two transcripts (either order - the broken/target roles are auto-detected)')
ap.add_argument('file_b', help='The other transcript')
ap.add_argument('--apply', action='store_true',
help='Write the output file. Without this flag, only the matching report is '
'printed (dry run) so you can review and confirm before anything is written.')
ap.add_argument('--tolerance', type=int, default=15,
help='Max seconds between matched timestamps (default: 15)')
ap.add_argument('--offset-search', type=int, default=90,
help='Search +/- this many seconds for a global clock offset between the two recordings (default: 90)')
ap.add_argument('--min-votes', type=int, default=2,
help='Minimum matched votes required to resolve a label (default: 2)')
ap.add_argument('--min-confidence', type=float, default=0.5,
help='Minimum vote share (0-1) required to resolve a label (default: 0.5)')
ap.add_argument('--output', help='Output path (default: <target-stem>-speakers-filled.md next to target)')
ap.add_argument('--no-normalize', action='store_true',
help='Do not merge duplicate-diarization name variants (e.g. "Name" / "Name1") in the broken file before voting')
ap.add_argument('--manual', action='append', metavar='"Speaker N=Real Name"',
help='Force a specific label to a specific name, bypassing matching entirely. '
'Repeatable. Overrides any automatic result (confident or not) for that label.')
ap.add_argument('--examples', type=int, default=5,
help='Example lines (with timestamps) to print per label, spread across its full '
'timeline, so matches can be spot-checked against the actual recording (default: 5)')
ap.add_argument('--recording-url',
help="Link to the target meeting's recording (e.g. a Fireflies share URL). If given, "
"it's printed once at the top and repeated under every label so a reviewer can "
"jump straight to it and scrub to each example's timestamp. There is no verified "
"way to encode an exact-timestamp deep link for Fireflies (the parameter their "
"own UI uses for 'copy link to this moment' isn't publicly documented), so this "
"links to the recording itself, not a specific moment in it.")
args = ap.parse_args()
manual_overrides = parse_manual_overrides(args.manual)
entries_a = parse_file(args.file_a)
entries_b = parse_file(args.file_b)
if not entries_a:
sys.exit(f"No '**Speaker** *[MM:SS]*: text' lines found in {args.file_a} - check the format.")
if not entries_b:
sys.exit(f"No '**Speaker** *[MM:SS]*: text' lines found in {args.file_b} - check the format.")
(target_entries, target_path), (broken_entries, broken_path) = classify(
entries_a, args.file_a, entries_b, args.file_b
)
print(f"Auto-detected roles: '{broken_path}' has real names (broken source), "
f"'{target_path}' has generic labels to fill (target).")
print()
merged_variants = defaultdict(set)
if not args.no_normalize:
for b in broken_entries:
raw = b['speaker']
normalized = normalize_name(raw)
if normalized != raw:
merged_variants[normalized].add(raw)
b['speaker'] = normalized
if merged_variants:
print("Merged duplicate diarization labels in the broken file (treated as one person):")
for normalized, raws in sorted(merged_variants.items()):
variants = sorted(raws | {normalized})
print(f" {' / '.join(variants)} -> {normalized}")
print()
broken_roster = sorted(set(b['speaker'] for b in broken_entries))
target_generic_labels = sorted(set(t['speaker'] for t in target_entries if GENERIC_RE.match(t['speaker'])))
if len(target_generic_labels) > len(broken_roster):
gap = len(target_generic_labels) - len(broken_roster)
print(f"NOTE: the target has {len(target_generic_labels)} distinct unnamed speakers but the broken "
f"file only names {len(broken_roster)} real people ({', '.join(broken_roster)}). At least "
f"{gap} target speaker(s) are structurally impossible to resolve correctly - they (or their "
f"voice segments) simply aren't captured with a real name anywhere in the broken file, so "
f"the closest-timestamp match for them, if any, will be a coincidence, not a correspondence. "
f"Treat any resolution below with real skepticism, especially ones with few votes.")
print()
offset, offset_score = find_best_offset(broken_entries, target_entries, args.tolerance, args.offset_search)
weighted_votes = defaultdict(Counter) # label -> name -> summed proximity weight
raw_votes = defaultdict(Counter) # label -> name -> raw match count (for the min-votes gate)
minority_lines = []
for t in target_entries:
if not GENERIC_RE.match(t['speaker']):
continue
match, delta = nearest_match(t['seconds'] + offset, broken_entries, args.tolerance)
if match:
weighted_votes[t['speaker']][match['speaker']] += proximity_weight(delta, args.tolerance)
raw_votes[t['speaker']][match['speaker']] += 1
t['_matched_name'] = match['speaker']
t['_matched_delta'] = delta
resolution = {}
for label, counter in weighted_votes.items():
total_weight = sum(counter.values())
ranked = counter.most_common() # all candidates, highest weighted vote share first
name, top_weight = ranked[0]
top_count = raw_votes[label][name]
confidence = top_weight / total_weight if total_weight else 0.0
resolved = top_count >= args.min_votes and confidence >= args.min_confidence
candidates = [
{
'name': cand_name,
'raw_count': raw_votes[label][cand_name],
'share': weight / total_weight if total_weight else 0.0,
}
for cand_name, weight in ranked
]
resolution[label] = {
'name': name if resolved else None,
'top_count': top_count,
'total_votes': sum(raw_votes[label].values()),
'confidence': confidence,
'candidates': candidates,
}
if resolved:
for t in target_entries:
if t['speaker'] == label and t.get('_matched_name') and t['_matched_name'] != name:
minority_lines.append((t['line_no'], t['seconds'], label, t['_matched_name'], name))
unknown_manual_labels = [label for label in manual_overrides if label not in target_generic_labels]
if unknown_manual_labels:
print(f"WARNING: --manual referenced label(s) not found as a generic speaker in the target file: "
f"{', '.join(unknown_manual_labels)} - ignoring them. Known generic labels: "
f"{', '.join(target_generic_labels)}")
print()
# Resolve any partial names in --manual against the broken file's real-name
# roster (e.g. "Dawid" -> "Dawid Cieślicki" if that's the only roster match),
# rather than taking the literal string when a better match is available.
manual_final = {}
for label, raw_name in manual_overrides.items():
if label not in target_generic_labels:
continue
resolved_name, note = resolve_partial_name(raw_name, broken_roster)
manual_final[label] = resolved_name
if note:
print(f"NOTE ({label}): {note}")
if manual_final:
print()
# Automatic `resolution` is left untouched here (it stays the source of
# truth for the automatic guess/candidates, printed for every label
# below regardless of whether a manual override wins); `final_name` is
# what actually gets written to disk.
final_name = {}
for label in target_generic_labels:
info = resolution.get(label)
final_name[label] = info['name'] if info else None
final_name.update(manual_final)
lines = Path(target_path).read_text(encoding='utf-8').splitlines()
resolved_line_count = 0
generic_totals = Counter(t['speaker'] for t in target_entries if GENERIC_RE.match(t['speaker']))
for t in target_entries:
if not GENERIC_RE.match(t['speaker']):
continue
name = final_name.get(t['speaker'])
if name:
old = f"**{t['speaker']}**"
new = f"**{name}**"
lines[t['line_no']] = lines[t['line_no']].replace(old, new, 1)
resolved_line_count += 1
out_path = Path(args.output) if args.output else Path(target_path).with_name(
Path(target_path).stem + "-speakers-filled.md"
)
if args.apply:
out_path.write_text("\n".join(lines) + "\n", encoding='utf-8')
total_generic_lines = sum(generic_totals.values())
resolved_labels = sum(1 for name in final_name.values() if name)
print(f"Global offset applied: {offset:+d}s (best fit: {offset_score}/{len(target_entries)} target lines matched at that offset)")
if args.apply:
print(f"Output written to: {out_path}")
else:
print(f"DRY RUN - no file written. Would write to: {out_path}")
if args.recording_url:
print(f"Recording: {args.recording_url}")
print("(no verified way to deep-link an exact timestamp - open this and scrub to each example below)")
print()
print("Speaker label resolution:")
unresolved_labels = []
for label in sorted(generic_totals, key=lambda l: -generic_totals[l]):
info = resolution.get(label)
manual_name = manual_final.get(label)
n_lines = generic_totals[label]
label_entries = [t for t in target_entries if t['speaker'] == label]
examples = sample_examples(label_entries, args.examples)
# Header line: the FINAL decision for this label (manual wins if given).
if manual_name:
agreement = ""
if info and info.get('name') == manual_name:
agreement = " - agrees with automatic guess"
elif info and info.get('name'):
agreement = f" - OVERRIDES automatic guess of '{info['name']}' ({info['confidence']:.0%} confidence)"
elif info:
agreement = " - automatic pass left this unresolved"
print(f" {label:12s} -> {manual_name:25s} (manually provided{agreement}) - {n_lines} lines")
elif info is None:
print(f" {label:12s} -> UNRESOLVED (no timestamp within tolerance found at all) - {n_lines} lines")
unresolved_labels.append(label)
elif info['name']:
print(f" {label:12s} -> {info['name']:25s} ({info['top_count']}/{info['total_votes']} votes, "
f"{info['confidence']:.0%} confidence) - {n_lines} lines")
else:
print(f" {label:12s} -> UNRESOLVED (top guess {info['top_count']}/{info['total_votes']} votes, "
f"{info['confidence']:.0%} confidence, below threshold) - {n_lines} lines")
unresolved_labels.append(label)
# All candidates with their confidence, for every label regardless of
# whether the final answer came from automatic matching or --manual -
# so a manual override's plausibility can still be judged against
# what the timestamps alone suggested.
if info and info.get('candidates'):
winner = info.get('name')
others = [c for c in info['candidates'] if c['name'] != winner] if winner else info['candidates']
label_str = "other candidates" if winner else "all candidates"
if others:
ranked_str = ", ".join(
f"{c['name']} ({c['raw_count']}/{info['total_votes']} votes, {c['share']:.0%})"
for c in others
)
print(f" {label_str}, most to least probable: {ranked_str}")
if examples:
print(f" example lines (jump to these timestamps in the recording to verify):")
print(format_examples(examples))
if args.recording_url:
print(f" recording: {args.recording_url}")
print()
print(f"Resolved {resolved_labels}/{len(generic_totals)} distinct generic labels, "
f"covering {resolved_line_count}/{total_generic_lines} generic-labeled lines.")
if minority_lines:
print()
print(f"{len(minority_lines)} individual line(s) disagreed with their label's majority vote "
f"(kept the majority name, flagging for manual review):")
for line_no, seconds, label, minority_name, majority_name in minority_lines[:20]:
mm, ss = divmod(seconds, 60)
print(f" line {line_no + 1} [{mm:02d}:{ss:02d}] {label}: nearest match was "
f"'{minority_name}', used majority '{majority_name}' instead")
if len(minority_lines) > 20:
print(f" ... and {len(minority_lines) - 20} more")
if unresolved_labels:
print()
print("=" * 70)
print(f"Could not match {len(unresolved_labels)} speaker(s): {', '.join(unresolved_labels)} "
f"(see their example lines/timestamps above).")
print("=" * 70)
print("If you know who any of these are, provide their real names and re-run with, e.g.:")
example = unresolved_labels[0]
print(f' --manual "{example}=Real Name"' + (' --manual "..."' if len(unresolved_labels) > 1 else ''))
if not args.apply:
print()
print("-" * 70)
print("DRY RUN - nothing was written. Review the resolution table above (and any "
"unresolved gaps), add --manual \"Speaker N=Real Name\" for anything to correct "
"or fill in, then re-run with --apply to write the output file.")
if __name__ == '__main__':
main()

View file

@ -15,3 +15,7 @@
- Confirmed local `main` is clean against `origin/main`, fetched `codeberg`, and found divergent histories: local has commits Codeberg lacks, while Codeberg has commits local lacks.
- Planned a normal merge of `codeberg/main` into local `main` before pushing, avoiding any force push.
- During conflict review, user clarified that the token-shaped value in Codeberg's invoice checker `.env.example` is a test token and should remain.
## Dual-remote sync turn
- User asked to commit and push all changes to both `origin` and `codeberg`.
- Confirmed the only working-tree changes were deletions of the Codeberg-added skill files, including `.env.example` templates, and user explicitly approved including those deletions.