Sirchmunk: Raw data to self-evolving intelligence, real-time.

Quick Start · Key Features · Web UI · How it Works · FAQ

🔍 Agentic Search • 🧠 Knowledge Clustering • 📊 Monte Carlo Evidence Sampling
⚡ Indexless Retrieval • 🔄 Self-Evolving Knowledge Base • 💬 Real-time Chat

🌰 Why “Sirchmunk”？

Intelligence pipelines built upon vector-based retrieval can be rigid and brittle. They rely on static vector embeddings that are expensive to compute, blind to real-time changes, and detached from the raw context. We introduce Sirchmunk to usher in a more agile paradigm, where data is no longer treated as a snapshot, and insights can evolve together with the data.

✨ Key Features

1. Embedding-Free: Data in its Purest Form

Sirchmunk works directly with raw data -- bypassing the heavy overhead of squeezing your rich files into fixed-dimensional vectors.

Instant Search: Eliminating complex pre-processing pipelines in hours long indexing; just drop your files and search immediately.
Full Fidelity: Zero information loss —- stay true to your data without vector approximation.

2. Self-Evolving: A Living Index

Data is a stream, not a snapshot. Sirchmunk is dynamic by design, while vector DB can become obsolete the moment your data changes.

Context-Aware: Evolves in real-time with your data context.
LLM-Powered Autonomy: Designed for Agents that perceive data as it lives, utilizing token-efficient reasoning that triggers LLM inference only when necessary to maximize intelligence while minimizing cost.

3. Intelligence at Scale: Real-Time & Massive

Sirchmunk bridges massive local repositories and the web with high-scale throughput and real-time awareness.
It serves as a unified intelligent hub for AI agents, delivering deep insights across vast datasets at the speed of thought.

Traditional RAG vs. Sirchmunk

Dimension	Traditional RAG	✨Sirchmunk
💰 Setup Cost	High Overhead (VectorDB, GraphDB, Complex Document Parser...)	✅ Zero Infrastructure Direct-to-data retrieval without vector silos
🕒 Data Freshness	Stale (Batch Re-indexing)	✅ Instant & Dynamic Self-evolving index that reflects live changes
📈 Scalability	Linear Cost Growth	✅ Extremely low RAM/CPU consumption Native Elastic Support, efficiently handles large-scale datasets
🎯 Accuracy	Approximate Vector Matches	✅ Deterministic & Contextual Hybrid logic ensuring semantic precision
⚙️ Workflow	Complex ETL Pipelines	✅ Drop-and-Search Zero-config integration for rapid deployment

Demonstration

Access files directly to start chatting

🎉 News

🎉🎉 Jan 22, 2026: Introducing Sirchmunk: Initial Release v0.0.1 Now Available!

🚀 Quick Start

Prerequisites

Python 3.10+
LLM API Key (OpenAI-compatible endpoint, local or remote)
Node.js 18+ (Optional, for web interface)

Installation

# Create virtual environment (recommended)
conda create -n sirchmunk python=3.13 -y && conda activate sirchmunk 

pip install sirchmunk

# Or via UV:
uv pip install sirchmunk

# Alternatively, install from source:
git clone https://github.com/modelscope/sirchmunk.git && cd sirchmunk
pip install -e .

Python SDK Usage

import asyncio

from sirchmunk import AgenticSearch
from sirchmunk.llm import OpenAIChat

llm = OpenAIChat(
        api_key="your-api-key",
        base_url="your-base-url",   # e.g., https://api.openai.com/v1
        model="your-model-name"     # e.g., gpt-4o
    )

async def main():
    
    agent_search = AgenticSearch(llm=llm)
    
    result: str = await agent_search.search(
        query="How does transformer attention work?",
        search_paths=["/path/to/documents"],
    )
    
    print(result)

asyncio.run(main())

⚠️ Notes:

Upon initialization, AgenticSearch automatically checks if ripgrep-all and ripgrep are installed. If they are missing, it will attempt to install them automatically. If the automatic installation fails, please install them manually.
- References: https://github.com/BurntSushi/ripgrep | https://github.com/phiresky/ripgrep-all
Replace "your-api-key", "your-base-url", "your-model-name" and /path/to/documents with your actual values.

🖥️ Web UI

The web UI is built for fast, transparent workflows: chat, knowledge analytics, and system monitoring in one place.

_{Home — Chat with streaming logs, file-based RAG, and session management.}

_{Monitor — System health, chat activity, knowledge analytics, and LLM usage.}

Installation

git clone https://github.com/modelscope/sirchmunk.git && cd sirchmunk

pip install ".[web]"

npm install --prefix web

Note: Node.js 18+ is required for the web interface.

Running the Web UI

# Start frontend and backend
python scripts/start_web.py 

# Stop frontend and backend
python scripts/stop_web.py

Access the web UI at (By default):

Backend APIs: http://localhost:8584/docs
Frontend: http://localhost:8585

Configuration:

Access Settings → Envrionment Variables to configure LLM API, and other parameters.

🏗️ How it Works

Sirchmunk Framework

Core Components

Component	Description
AgenticSearch	Search orchestrator with LLM-enhanced retrieval capabilities
KnowledgeBase	Transforms raw results into structured knowledge clusters with evidences
EvidenceProcessor	Evidence processing based on the MonteCarlo Importance Sampling
GrepRetriever	High-performance indexless file search with parallel processing
OpenAIChat	Unified LLM interface supporting streaming and usage tracking
MonitorTracker	Real-time system and application metrics collection

Data Storage

All persistent data is stored in the configured WORK_PATH (default: ~/.sirchmunk/):

{WORK_PATH}/
  ├── .cache/
    ├── history/              # Chat session history (DuckDB)
    │   └── chat_history.db
    ├── knowledge/            # Knowledge clusters (Parquet)
    │   └── knowledge_clusters.parquet
    └── settings/             # User settings (DuckDB)
        └── settings.db

❓ FAQ

How is this different from traditional RAG systems?

Sirchmunk takes an indexless approach:

No pre-indexing: Direct file search without vector database setup
Self-evolving: Knowledge clusters evolve based on search patterns
Multi-level retrieval: Adaptive keyword granularity for better recall
Evidence-based: Monte Carlo sampling for precise content extraction

What LLM providers are supported?

Any OpenAI-compatible API endpoint, including (but not limited too):

OpenAI (GPT-4, GPT-4o, GPT-3.5)
Local models served via Ollama, llama.cpp, vLLM, SGLang etc.
Claude via API proxy

How do I add documents to search?

Simply specify the path in your search query:

result = await search.search(
    query="Your question",
    search_paths=["/path/to/folder", "/path/to/file.pdf"]
)

No pre-processing or indexing required!

Where are knowledge clusters stored?

Knowledge clusters are persisted in Parquet format at:

{WORK_PATH}/.cache/knowledge/knowledge_clusters.parquet

You can query them using DuckDB or the KnowledgeManager API.

How do I monitor LLM token usage?

Web Dashboard: Visit the Monitor page for real-time statistics
API: GET /api/v1/monitor/llm returns usage metrics
Code: Access search.llm_usages after search completion

📋 Roadmap

Text-retrieval from raw files
Knowledge structuring & persistence
Real-time chat with RAG
Web UI support
Web search integration
Multi-modal support (images, videos)
Distributed search across nodes
Knowledge visualization and deep analytics
More file type support

🤝 Contributing

We welcome contributions !

📄 License

This project is licensed under the Apache License 2.0.

ModelScope · ⭐ Star us · 🐛 Report a bug · 💬 Discussions

✨ Sirchmunk: Raw data to self-evolving intelligence, real-time.

❤️ Thanks for Visiting ✨ Sirchmunk !

Name		Name	Last commit message	Last commit date
Latest commit History 118 Commits
.github/workflows		.github/workflows
assets		assets
requirements		requirements
scripts		scripts
src		src
web		web
.gitignore		.gitignore
.pre-commit-config.yaml		.pre-commit-config.yaml
LICENSE		LICENSE
README.md		README.md
README_zh.md		README_zh.md
pyproject.toml		pyproject.toml
requirements.txt		requirements.txt
setup.py		setup.py

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

Sirchmunk: Raw data to self-evolving intelligence, real-time.

🌰 Why “Sirchmunk”？

✨ Key Features

1. Embedding-Free: Data in its Purest Form

2. Self-Evolving: A Living Index

3. Intelligence at Scale: Real-Time & Massive

Traditional RAG vs. Sirchmunk

Demonstration

🎉 News

🚀 Quick Start

Prerequisites

Installation

Python SDK Usage

🖥️ Web UI

Installation

Running the Web UI

🏗️ How it Works

Sirchmunk Framework

Core Components

Data Storage

❓ FAQ

📋 Roadmap

🤝 Contributing

📄 License

About

Uh oh!

Releases 2

Packages

Uh oh!

Contributors 2

Uh oh!

Languages

License

modelscope/sirchmunk

Folders and files

Latest commit

History

Repository files navigation

Sirchmunk: Raw data to self-evolving intelligence, real-time.

🌰 Why “Sirchmunk”？

✨ Key Features

1. Embedding-Free: Data in its Purest Form

2. Self-Evolving: A Living Index

3. Intelligence at Scale: Real-Time & Massive

Traditional RAG vs. Sirchmunk

Demonstration

🎉 News

🚀 Quick Start

Prerequisites

Installation

Python SDK Usage

🖥️ Web UI

Installation

Running the Web UI

🏗️ How it Works

Sirchmunk Framework

Core Components

Data Storage

❓ FAQ

📋 Roadmap

🤝 Contributing

📄 License

About

Resources

License

Uh oh!

Stars

Watchers

Forks

Releases 2

Packages 0

Uh oh!

Contributors 2

Uh oh!

Languages

Packages