IJCOPE Journal

UGC Logo DOI / ISO Logo

International Journal of Creative and Open Research in Engineering and Management

A Peer-Reviewed, Open-Access International Journal Supporting Multidisciplinary Research, Digital Publishing Standards, DOI Registration, and Academic Indexing.
Journal Information
ISSN: 3108-1754 (Online)
Crossref DOI: Available
ISO Certification: 9001:2015
Publication Fee: 599/- INR
Compliance: UGC Journal Norms
License: CC BY 4.0
Peer Review: Double Blind
Volume 02, Issue 8

Published on: August 2026

ADAPTIVE MULTIMODAL DOCUMENT INGESTION AND SELF-CORRECTING HYBRID RAG VIA LANGGRAPH MULTI-AGENT WORKFLOW

Yashas.B.S

Advanced Analyst III, EY GDS and Reva University

Article Status

Plagiarism Passed Peer Reviewed Open Access

Available Documents

Abstract

Document digitization and knowledge extraction remain challenging when dealing with heterogeneous PDF repositories comprising both machine-readable text and degraded, scanned visual artifacts. Traditional Optical Character Recognition (OCR) systems enforce rigid linear pipelines, while conventional Retrieval-Augmented Generation (RAG) models suffer from hallucination when context is sparse or noisy. In this paper, we propose DocuMind-AI, an end-to-end autonomous agentic architecture for document processing and intelligent question answering orchestrated via LangGraph. The framework introduces a dynamic routing agent that assesses character density metrics to intelligently dispatch inputs between fast native text extractors and multimodal vision large language models. Extracted text is normalized into structured Markdown through an automated cleansing agent and ingested into a dual-engine hybrid retrieval index combining BM25 lexical search with dense vector embeddings via Reciprocal Rank Fusion (RRF). Furthermore, a Corrective RAG (CRAG) self-reflection loop audits retrieved document chunks for semantic relevance, triggering automated query reformulation when retrieval confidence is low, and performs secondary hallucination auditing on synthesized answers. Empirical benchmarks demonstrate that our adaptive routing reduces multimodal API overhead by 68.4% on mixed corpora while achieving a 94.2% answer grounding accuracy, outperforming conventional single-engine RAG pipelines in both precision and computational efficiency.

 

Keywords— Agentic AI; Optical Character Recognition; Corrective RAG; LangGraph; Multimodal LLM; Hybrid Search.

How to Cite this Paper

Yashas.B.S, (2026). Adaptive Multimodal Document Ingestion and Self-Correcting Hybrid RAG via LangGraph Multi-Agent Workflow. International Journal of Creative and Open Research in Engineering and Management, <i>02</i>(8), 1-9. https://doi.org/10.55041/ijcope.v2i8.197

Yashas.B.S, . "Adaptive Multimodal Document Ingestion and Self-Correcting Hybrid RAG via LangGraph Multi-Agent Workflow." International Journal of Creative and Open Research in Engineering and Management, vol. 02, no. 8, 2026, pp. 1-9. doi:https://doi.org/10.55041/ijcope.v2i8.197.

Yashas.B.S, . "Adaptive Multimodal Document Ingestion and Self-Correcting Hybrid RAG via LangGraph Multi-Agent Workflow." International Journal of Creative and Open Research in Engineering and Management 02, no. 8 (2026): 1-9. https://doi.org/https://doi.org/10.55041/ijcope.v2i8.197.

Search & Index

References

[1] A. Vaswani et al., 'Attention is all you need,' in Advances in Neural Information Processing Systems (NeurIPS), vol. 30, pp. 5998–6008, 2017.

[2] H. Cui et al., 'Document AI: Benchmarks, models and applications,' ACM Computing Surveys, vol. 56, no. 4, pp. 1–38, 2024.

[3] P. Lewis et al., 'Retrieval-augmented generation for knowledge-intensive NLP tasks,' in Advances in Neural Information Processing Systems (NeurIPS), vol. 33, pp. 9459–9474, 2020.

[4] J. Gao et al., 'Retrieval-augmented generation for large language models: A survey,' arXiv preprint arXiv:2312.10997, 2023.

[5] R. Smith, 'An overview of the Tesseract OCR engine,' in Ninth International Conference on Document Analysis and Recognition (ICDAR), vol. 2, pp. 629–633, 2007.

[6] Y. Huang et al., 'LayoutLMv3: Pre-training for document AI with unified text and image masking,' in Proc. ACM Multimedia, pp. 4083–4091, 2022.

[7] Gemini Team, 'Gemini: A family of highly capable multimodal models,' arXiv preprint arXiv:2312.11805, 2023.

[8] G. Izacard and E. Grave, 'Leveraging passage retrieval with generative models for open domain question answering,' in Proc. EACL, pp. 874–880, 2021.

[9] V. Karpukhin et al., 'Dense passage retrieval for open-domain question answering,' in Proc. EMNLP, pp. 6769–6781, 2020.

[10] S. Robertson and H. Zaragoza, 'The probabilistic relevance framework: BM25 and beyond,' Foundations and Trends in Information Retrieval, vol. 3, no. 4, pp. 333–389, 2009.

Ethical Compliance & Review Process

  • All submissions are screened under plagiarism detection.
  • Review follows editorial policy.
  • Authors retain copyright.
  • Peer Review Type: Double-Blind Peer Review
  • Published on: Aug 24 2026
CCBYNC

This article is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License. You are free to share and adapt this work for non-commercial purposes with proper attribution.

View License
Scroll to Top