Multimodal AI

Slide Narrator — Multimodal AI Presentation Intelligence

A multimodal pipeline that understands presentation slides using OCR, table extraction, image captioning and grounded LLM summarization.

Role

Data Scientist / AI Engineer

Context

Selected AI Project

Period

2025–2026

The challenge

Presentation slides mix text, charts, tables and images, so plain text extraction loses structure and often produces weak summaries or narration.

What I built

Designed a slide-aware multimodal pipeline that extracts text with OCR, preserves tables, captions visual elements and stores metadata for each slide element. The LLM then generates grounded summaries and narration with review checkpoints and style adaptation for different audiences.

Architecture

  • OCR and layout-aware text extraction
  • Table extraction and structure preservation
  • Image captioning for visual context
  • Slide-level chunking and element metadata
  • Grounded summarization with human-in-the-loop review

Outcomes

  • Better preservation of slide structure and context
  • More grounded narration across text, tables and images
  • Audience-aware explanation and summarization workflows

Technology

PythonOCRMultimodal LLMsImage CaptioningRAGStructured Metadata
LET'S BUILD
CONNECTION CHANNEL OPEN
FINAL SYSTEM / PORTFOLIO END
NEXT PROJECT / NEXT SYSTEM / NEXT IDEA

LET'S BUILD

SOMETHING

INTELLIGENT.

Interested in building intelligent products, production AI systems, agentic workflows, machine learning platforms, or something that does not exist yet? Start the conversation.

CREATED BYPathan Afnan Khan✦
2026 © ALL RIGHTS RESERVEDPORTFOLIO / SYSTEM COMPLETE
Built with♡& intelligence
AGENTIC AI✦
GENERATIVE AI✦
MACHINE LEARNING✦
DATA ENGINEERING✦
CLOUD AI✦
MLOPS✦
INTELLIGENT SYSTEMS✦
AGENTIC AI✦
GENERATIVE AI✦
MACHINE LEARNING✦
DATA ENGINEERING✦
CLOUD AI✦
MLOPS✦
INTELLIGENT SYSTEMS✦