The challenge
Presentation slides mix text, charts, tables and images, so plain text extraction loses structure and often produces weak summaries or narration.
Multimodal AI
A multimodal pipeline that understands presentation slides using OCR, table extraction, image captioning and grounded LLM summarization.
Role
Data Scientist / AI Engineer
Context
Selected AI Project
Period
2025–2026
Presentation slides mix text, charts, tables and images, so plain text extraction loses structure and often produces weak summaries or narration.
Designed a slide-aware multimodal pipeline that extracts text with OCR, preserves tables, captions visual elements and stores metadata for each slide element. The LLM then generates grounded summaries and narration with review checkpoints and style adaptation for different audiences.