ContextCrate
ContextCrate turns website and Git sources into a durable pipeline of raw artifacts, normalized documents, chunks, and search-index records.
The default standalone profile runs in one JVM with file-backed H2, a database queue, filesystem artifacts, and Lucene. The distributed profile runs independently scalable roles with PostgreSQL, RabbitMQ, S3, and OpenSearch. Both modes share the same domain and work contracts.
Version 1 outcome
- Create a Website or Git source, then attach one or more ingestion jobs.
- Start an immutable ingestion run snapshot.
- Acquire eligible pages or repository files while enforcing connector safety policy.
- Store raw content outside the queue.
- Parse, normalize, extract links and metadata, and create stable chunks.
- Index document and chunk records.
Retrieval supports lexical BM25, semantic-vector, RRF-hybrid, and optional cross-encoder reranking. The default embedding provider runs a local multilingual ONNX model; an OpenAI-compatible embeddings endpoint can be selected instead. The remaining roadmap covers feedback, evaluation, and learning-to-rank rollout controls.