StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description

Published in European Conference on Computer Vision (ECCV), 2026

Recommended citation: Hahm, S. H., Dinh, M. T., & Jin, S. (2026). "StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description." European Conference on Computer Vision (ECCV). https://arxiv.org/abs/2607.11798

Long-form audio description needs to preserve more than the visible action in an isolated clip: viewers should be able to follow characters, events, relationships, and story context across scenes. StoryTeller is a training-free framework that maintains a verified narrative memory while processing long-form video chronologically, producing descriptions that are more coherent and grounded.

The accompanying StoryAD-QA benchmark evaluates whether generated descriptions preserve the information needed to answer grounded questions about a story.

Read the paper on arXiv

StoryAD-QA release on GitHub