Research
Duplicate Doc Clusters
Group exact and near-duplicate documents without losing version history.
Configure this skill
The page and the file follow your answers.
CRM
Email
Chat
Documents
Calendar
Project tools
6 of 6 systems connected
KeepDirectMissing
Your assistantSKILL.md
Download ↓
When to use it
Group exact and near-duplicate documents without losing version history.
What it covers
Inputs
Document set, comparison scope, cutoff.
Result
Duplicate clusters with IDs, duplicate type, differences and proposed canonical source.
What it uses
CalendarEmailDocumentsSlackProject toolsClaudeKeep memory
For developers
Retrieval instructions
Resolve people, accounts and projects by stable identifiers. Use only the sources this task needs. Cite the source and date for each finding. Keep source systems unchanged.
Data sources
Google Workspace, Slack, GitHub, Claude, Keep memory.
Procedure
- Resolve inputs for Duplicate Doc Clusters: topic, time_window, project_key, person. Default time window: canonical docs + last 30 days of related threads (override if caller provides time_window). Ask only for.
- Run keep_search with these skill-specific intents (keep meaning; adjust wording to corpus):
- Additional keep_search passes for residual signal: