nexusai.stackwise.site · AI Product

Nexus AI Workspace – Agentic RAG

Nexus AI Workspace: agentic RAG with hybrid search — P95 retrieval <500ms*, 100% traceable citations*, 1,000+ pages/hour parse*

<500ms*
P95 retrieval
100%*
Traceable cites
1k+/hr*
Pages parsed
Agentic
RAG

Executive summary

  • •Problem: teams need trustworthy answers from their own PDFs — standard retrieve-once RAG fails when context is thin.
  • •Solution: Docling ingest → OpenSearch hybrid index → LangGraph agent (retrieve/rewrite/grade) → cited answers.
  • •Result (product-published*): P95 retrieval <500ms, traceable citations, 1,000+ pages parsed/hour.

The business challenge

Knowledge work depends on contracts, policies, and reports locked in complex PDFs. Naive RAG retrieves once and answers anyway. Buyers need page-level citations, confidence, and an agent that can rewrite/retrieve again when context is insufficient.

  • •Wrong answers from thin retrieval context
  • •Hours lost hunting clauses across PDFs
  • •Inability to audit which page an answer came from
  • •Wrapper tools that cannot parse complex tables/layouts

Goals & non-goals

Goals

  • •Parse complex PDFs (tables, multi-column) cleanly
  • •Hybrid retrieval (BM25 + dense vectors)
  • •Agentic loop: evaluate → retrieve/rewrite → grade → answer
  • •Every claim traceable to page + excerpt + confidence

Non-goals

  • •Training foundation models on customer corpora
  • •Answer-without-source modes as default

The solution

Nexus AI Workspace runs an agentic RAG pipeline: Docling parsing, dual indexing in OpenSearch (BM25 + embeddings), LangGraph reasoning agent, and grounded responses with page numbers, quotes, and confidence.

Ingest

Docling parses complex layouts into structured chunks

Index

Hybrid BM25 + sentence-transformer vectors

Reason

LangGraph agent can re-retrieve / rewrite / grade

Answer

Page citations + excerpts + confidence

Modular

Knowledge Retriever can run standalone

Technical architecture

LayerTechnologyWhyAlternatives
ParsingDoclingTables/headers/multi-column fidelityNaive text extract
SearchOpenSearch hybridKeyword precision + semantic recallVector-only OR keyword-only
AgentLangGraphMulti-step retrieval decisionsSingle-shot RAG
EmbeddingsSentence transformersDense retrieval qualityKeyword only
TrustCitations + confidenceAuditable answersChat without sources

Implementation path

  1. 1.PDF ingest + Docling chunking
  2. 2.Dual index build (BM25 + vectors)
  3. 3.LangGraph agent graph for retrieve/grade/answer
  4. 4.Citation packing with page anchors + confidence

Challenges & trade-offs

Complex PDFs

Why hard: Tables/multi-column break naive extractors

Solution: Docling advanced parsing

Trade-off: Heavier ingest CPU vs plain text

Thin context

Why hard: Retrieve-once answers confidently wrong

Solution: Agentic re-retrieve/rewrite/grade loop

Trade-off: Higher latency budget than single-shot

Results & metrics

Labels: WP = website-published · OPS = operational outcome · ARCH = design target

MetricBeforeAfterSourceLabel
P95 retrieval latency—< 500ms*stackwise.site product pageWP
Citation traceabilityUncited chat100% claims intended source-traceable*product pageWP
Parse throughput—1,000+ pages/hour*product pageWP
RAG typeStandard retrieve-onceAgentic LangGraph loopArchitectureOPS

Value by stakeholder

Technical evaluators

Named stack: Docling, OpenSearch, LangGraph

Risk/compliance

Page-level citations for audit

Ops/knowledge teams

Ask questions across corpora instead of hunting files

Procurement

Performance metrics published for comparison

Planning something similar?

Tell us the stuck flow. We can start with a scoped paid diagnosis — reproduce, investigate, and give written options.