ckb/.agents/skills/ckb-index-external/SKILL.md
Michał Kopeć 55c8c352d1 Add connector-backed libs/ with self-contained external source indexing
libs/<name>/ now supports a second population mode alongside git-copy
clones: a user-authored source.yaml declares a live external source
(SharePoint, Google Drive, a plain URL, or another connector), and the
new ckb-index-external skill builds a self-contained generated index
for it (index.md/entities/graph/log.md), scoped entirely to that
connector and never blended into the main wiki/.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 15:56:37 +02:00

175 lines
8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
name: ckb-index-external
description: Walk every connector-backed libs/<name>/ (identified by a source.yaml declaring a live external source — SharePoint, Google Drive, a plain URL, or another connector) and build/refresh a self-contained generated index for it — index.md/entities/graph/log.md, mirroring wiki/'s own shape but scoped entirely to that one connector. Never touches wiki/, never touches source.yaml, never touches a git-copy lib. Use when the user says "index external sources", "index libs", "refresh the external index", or "scan the SharePoint/Drive folder". This is the on-demand workflow from CLAUDE.md/AGENTS.md §4, distinct from `ckb-ingest` (which processes raw/inbox/ into wiki/) and from a plain git-copy libs/<name>/ (a frozen clone, never touched by any skill).
---
# Index external sources skill
## Purpose
Give `libs/<name>/` a second population mode alongside the existing
git-copy one: a `libs/<name>/source.yaml` declares a *live* external
source — a SharePoint folder, a Google Drive folder, a plain URL, or
another connector — that the user doesn't want to fully mirror locally.
This skill walks that source and builds a small, self-contained index of
what's there: one entity page per document found, plus the same
`index.md`/`graph`/`log.md` shape `wiki/` already uses, rooted at
`libs/<name>/` instead.
This is deliberately **not** blended into the main `wiki/entities/` or
`wiki/graph/edges.json` — the generated index lives entirely inside its own
`libs/<name>/`, at the lowest cascade layer, the same way a git-cloned
KB's own files would. If something in it needs to override or correct what
the connector says, that's what writing the corrected version into `wiki/`
is for (cascade priority already covers that — `wiki/` always wins).
This skill only ever writes within a connector-backed `libs/<name>/` (never
`source.yaml` itself, never a git-copy lib, never anything under `wiki/`)
— and it never touches this repo's own git remote; pushing the resulting
changes is a separate, explicit step (see Step 6).
## Trigger phrases
Use this skill when the user says:
- "Index external sources"
- "Index libs"
- "Refresh the external index"
- "Scan the SharePoint/Drive folder" (or similarly names a specific
connector)
Do **not** use this skill for "Ingest" / "Sync the wiki" (that's
`ckb-ingest` — raw material into `wiki/`, unrelated to `libs/`) or for
"Lint" (that's `ckb-lint`, which now also health-checks a connector's own
generated index, but doesn't build it in the first place).
## How to run this skill
### Step 1 — Find connector-backed libs
Scan `libs/*/` for a `source.yaml`. A `libs/<name>/` with one is
connector-backed and in scope; a `libs/<name>/` without one is a git-copy
lib and stays completely untouched by this skill.
If no `libs/<name>/source.yaml` exists anywhere, tell the user there's
nothing to index yet and how to add one (a `source.yaml` with `connector`
and `location`), and stop — don't fabricate an index from nothing.
### Step 2 — Read the config and resolve the connector
Read `source.yaml`:
```yaml
connector: sharepoint # sharepoint | google_drive | web | custom — open string, not a fixed enum
location: "https://contoso.sharepoint.com/sites/Finance/Shared Documents/Reports"
description: "Finance team's shared reports folder" # optional
```
Resolve `connector` to whatever live tool is actually available this
session:
- `sharepoint` / a Microsoft-hosted location → search (`ToolSearch`) for a
connected Microsoft 365 MCP tool.
- `google_drive` → search for a connected Google Drive MCP tool.
- `web` / `custom` / anything else with a plain URL → `WebFetch`.
If the needed connector isn't authorized in this session (per the current
MCP auth state), **don't fail the whole run** — tell the user which
connector needs authorizing (via claude.ai connector settings, or `claude
mcp`/`/mcp` for non-claude.ai servers), skip that source, and continue with
any others.
### Step 3 — Enumerate and summarize documents
For each document found at the location:
- Read enough content to write a real 1-2 sentence `tldr` when the
connector allows fetching content.
- If only metadata is available (listing only, no read access), say so
honestly in the `tldr` (e.g. "Metadata only — read access not granted")
and set `confidence` low (e.g. 0.3) rather than fabricating a summary.
Create/update one entity page per document at
`libs/<name>/entities/<doc-slug>.md`:
```yaml
---
type: document
resource: <document's own URL/path>
tldr: ...
confidence: 0.30.9 # per the rule above
quality: ...
last_updated: YYYY-MM-DD
freshness_window_days: 30 # shorter than a typical wiki page — external sources change without notice
retention: medium
---
```
### Step 4 — Update the hub page and graph
Create/update `libs/<name>/index.md` — the root routing page for this
connector, mirroring `wiki/index.md`'s own role:
```yaml
---
type: external-source
resource: <location, from source.yaml>
tldr: One sentence — what this source is and how many documents were found
last_updated: YYYY-MM-DD
---
```
followed by a short routing table pointing to `entities/index.md` and
`graph/index.md`.
Record any real relationships between documents (e.g. folder hierarchy) as
edges in `libs/<name>/graph/edges.json`, using a `contains`/`part_of`-style
edge type — this is fine to be thin or empty for a flat source with no
useful structure beyond a document list.
### Step 5 — Update index/log, following the Recursive Index & Log Convention
Update `libs/<name>/entities/index.md` and `libs/<name>/graph/index.md`
(flat bullet lists, no frontmatter, same convention as `wiki/entities/
index.md`/`wiki/graph/index.md`). Log every created/updated page in
`libs/<name>/log.md`, same reverse-chronological format as Rule B
(`CLAUDE.md`/`AGENTS.md` §7) — this log is independent of `wiki/log.md`;
**nothing under `wiki/` is touched by this skill at all.**
### Step 6 — Remind to review and sync
This is always the last step, every time this skill makes any change at
all. Close with a short reminder — do not sync or push anything yourself
here:
> "External source indexing complete for `libs/<name>/` — see
> `libs/<name>/log.md` for the details. Please review the changes, and
> once you're happy with them, say 'sync changes' to commit and push them
> to `origin`."
If nothing was actually indexed (no connector-backed libs found, or every
source was skipped for lack of authorization), skip this reminder.
## Edge cases
- **No connector-backed `libs/<name>/` exists** — report that and explain
how to add one; don't touch anything.
- **Connector not authorized this session** — report which one and how to
authorize it (claude.ai connector settings / `claude mcp` / `/mcp`),
skip that source, keep going with the rest. Never ask the user for
tokens or callback URLs.
- **Re-running against an already-indexed source** — refresh existing
entity pages in place (update `tldr`/`last_updated`/`confidence`); never
duplicate a document's page.
- **A previously-indexed document is no longer found at the source** —
don't delete its page. Update its `tldr` with a note ("No longer found
at source as of YYYY-MM-DD") and set `retention: low`, so the existing
`ckb-lint` retention sweep archives it naturally on a later pass — no new
archival mechanism needed.
- **A source is very large** — cap what gets fully fetched/summarized in
one run and explicitly report what was skipped (e.g. "42 of ~300
documents summarized this pass — re-run to continue"). Never silently
truncate while implying full coverage.
- **`libs/<name>/` has both real content files and a `source.yaml`** —
ambiguous and out of scope for this skill to resolve; report it and stop
for that folder (this is the same case `ckb-lint`'s new external-source
config check flags).
- **`source.yaml` is missing `connector` or `location`** — report it as
malformed and skip that source; don't guess at intent.
---
*Licensed under the Apache License, Version 2.0 — see [LICENSE](../../../LICENSE)
at the repository root, or <http://www.apache.org/licenses/LICENSE-2.0>.*