IJCOPE Journal

UGC Logo DOI / ISO Logo

International Journal of Creative and Open Research in Engineering and Management

A Peer-Reviewed, Open-Access International Journal Supporting Multidisciplinary Research, Digital Publishing Standards, DOI Registration, and Academic Indexing.
Journal Information
ISSN: 3108-1754 (Online)
Crossref DOI: Available
ISO Certification: 9001:2015
Publication Fee: 599/- INR
Compliance: UGC Journal Norms
License: CC BY 4.0
Peer Review: Double Blind
Volume 02, Issue 8

Published on: August 2026

A MODULAR MULTIMODAL FRAMEWORK FOR AUTOMATED MISINFORMATION INVESTIGATION: INTEGRATING OCR, NLP, VISION-LANGUAGE CAPTIONING, AND SPEECH ANALYSIS FOR EVIDENCE-BASED TRUST SCORING

Shivani Tiwari

Military College of Telecommunication Engineering, Mhow, India

Article Status

Plagiarism Passed Peer Reviewed Open Access

Available Documents

Abstract

The rapid spread of manipulated images, doctored videos, and misleading text across digital platforms has created a pressing need for automated tools that can assist human fact-checkers rather than replace them. This paper presents the design and implementation of a modular, multimodal misinformation investigation platform that combines optical character recognition (OCR), natural language processing (NLP), vision-language image captioning, and automatic speech recognition (ASR) within a single evidence-fusion pipeline. The system extracts on-screen text, spoken audio, visual content descriptions, and linguistic sentiment/clickbait indicators from a submitted image or video, and combines these independent signals into a composite trust score and human-readable explanation through a rule-based fusion engine. The platform is implemented as a FastAPI backend with a MySQL relational store, JSON Web Token (JWT) based authentication, and a React single-page-application frontend that supports investigation history, analytics, and report export. We describe the architecture, the responsibilities of each processing module, the design of the evidence-fusion scoring function, and the security and data-persistence layers. We further discuss the current limitations of the rule-based fusion approach, including its sensitivity to keyword-based clickbait detection and its reliance on unimodal confidence rather than learned cross-modal weighting, and we outline directions for extending the system toward a trained fusion model and stronger authentication practices.

Keywords— misinformation detection, multimodal fusion, optical character recognition, natural language processing, vision-language models, speech recognition, trust scoring, FastAPI, digital forensics.

How to Cite this Paper

Tiwari, S. (2026). A Modular Multimodal Framework for Automated Misinformation Investigation: Integrating OCR, NLP, Vision-Language Captioning, and Speech Analysis for Evidence-Based Trust Scoring. International Journal of Creative and Open Research in Engineering and Management, <i>02</i>(8), 1-9. https://doi.org/10.55041/ijcope.v2i8.147

Tiwari, Shivani. "A Modular Multimodal Framework for Automated Misinformation Investigation: Integrating OCR, NLP, Vision-Language Captioning, and Speech Analysis for Evidence-Based Trust Scoring." International Journal of Creative and Open Research in Engineering and Management, vol. 02, no. 8, 2026, pp. 1-9. doi:https://doi.org/10.55041/ijcope.v2i8.147.

Tiwari, Shivani. "A Modular Multimodal Framework for Automated Misinformation Investigation: Integrating OCR, NLP, Vision-Language Captioning, and Speech Analysis for Evidence-Based Trust Scoring." International Journal of Creative and Open Research in Engineering and Management 02, no. 8 (2026): 1-9. https://doi.org/https://doi.org/10.55041/ijcope.v2i8.147.

Search & Index

References


  • Vosoughi, D. Roy, and S. Aral, "The spread of true and false news online," Science, vol. 359, no. 6380, pp. 1146–1151, 2018.

  • Rashkin, E. Choi, J. Y. Jang, S. Volkova, and Y. Choi, "Truth of varying shades: Analyzing language in fake news and political fact-checking," in Proc. Conf. Empirical Methods in Natural Language Processing (EMNLP), 2017, pp. 2931–2937.

  • Li, D. Li, C. Xiong, and S. Hoi, "BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation," in Proc. Int. Conf. Machine Learning (ICML), 2022.

  • Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and

    1. Sutskever, "Robust speech recognition via large-scale weak supervision," arXiv preprint arXiv:2212.04356, 2022.



  • Honnibal and I. Montani, "spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing," 2017. [Online software].

  • Sanh, L. Debut, J. Chaumond, and T. Wolf, "DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter," arXiv preprint arXiv:1910.01108, 2019.

  • Shu, A. Sliva, S. Wang, J. Tang, and H. Liu, "Fake news detection on social media: A data mining perspective," ACM SIGKDD Explorations Newsletter, vol. 19, no. 1, pp. 22–36,2017.



  • Zhou and R. Zafarani, "A survey of fake news: Fundamental theories, detection methods, and opportunities," ACM Computing Surveys, vol. 53, no. 5, pp. 1–40, 2020.

  • JaidedAI, "EasyOCR: Ready-to-use OCR with 80+ supported languages," GitHub repository, 2020. [Online software].

  • Ramalho et al., "FastAPI: A modern, fast web framework for building APIs with Python," 2018–present. [Online software].

Ethical Compliance & Review Process

  • All submissions are screened under plagiarism detection.
  • Review follows editorial policy.
  • Authors retain copyright.
  • Peer Review Type: Double-Blind Peer Review
  • Published on: Aug 19 2026
CCBYNC

This article is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License. You are free to share and adapt this work for non-commercial purposes with proper attribution.

View License
Scroll to Top