Live Caption Delivery System Using Follower Script Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech follower systems are ineffective in live theatre settings due to poor audio quality, variability in speech styles and speeds, background noise, and non-verbal performance elements, making it difficult to automatically cue captions accurately.
Innovation Solution
A caption delivery system that tracks live performances against a pre-defined script using a follower script and caption script, with waypoints and metadata to assist the speech follower in synchronizing captions, including 'soft' metadata to manage non-dialogue events and 'hard' metadata for cue handling, ensuring accurate caption timing and display.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If speech recognition systems are used to automatically follow speech in live theatre, then automation of caption delivery is improved, but accuracy deteriorates due to poor audio quality, background noise, and variability in speech styles
Solution Approach 1:
The patent introduces an intermediary system comprising a follower script and caption script that mediates between the live performance and the caption delivery. The follower script acts as a reference framework that translates the complex live performance into a structured sequence of expected events, allowing the system to compare actual performance against this pre-defined structure and maintain accurate caption timing despite audio challenges.
Solution Approach 2:
The system performs preliminary action by pre-defining the follower script and caption script before the live performance occurs. The follower script is created by analyzing the production script and identifying key performance events, allowing the system to be prepared with an expected timeline and structure that can guide accurate caption delivery during the actual performance, overcoming the unpredictability of live speech.
2Measurement precision
If manual caption triggering is used in live theatre, then accuracy of caption timing is improved, but productivity and accessibility deteriorate due to manual operation requirements
Solution Approach 1:
The system enables self-service by allowing the live performance itself to trigger the caption delivery automatically. Through the speech follower component that monitors actual performance events against the follower script, the system serves itself by using the performance data to automatically determine when captions should be displayed, eliminating the need for manual intervention while maintaining timing accuracy.
Solution Approach 2:
The patent implements feedback by continuously comparing the actual live performance against the pre-defined follower script and using this comparison to adjust caption delivery timing. The speech follower component receives feedback from the actual performance events and uses this information to synchronize captions accurately, creating a closed-loop system that maintains precision without manual control.
3Measurement precision
If speech follower systems are trained to standard speech styles, then processing accuracy is improved, but adaptability to live theatre performance variability deteriorates
Solution Approach 1:
The patent applies segmentation by dividing the performance tracking into discrete, identifiable events based on the follower script. Instead of attempting to process continuous speech variations, the system segments the performance into key events (dialogue, music, sound effects) that can be individually tracked and matched to corresponding captions, allowing accurate processing without requiring the system to adapt to every speech style variation.
Solution Approach 2:
The system uses parameter changes by adjusting the reference framework from standard speech patterns to a performance-specific follower script that is customized for each production. This allows the system to maintain high processing accuracy by comparing actual performance against the specific expected parameters of that particular production, rather than attempting to accommodate all possible speech variations with a single standardized model.
Data Source
AI summary
A system and method of delivering an information output to a viewer of a live performance. The information output can be displayed text or an audio description at predefined times in the live performance relative to stage events. A follower script with entries organised along a timeline, and metadata at time points between at least some of the entries is generated. The metadata is associated with stage events in the live performance. The system uses speech recognition to track spoken dialogue against the entries in the follower script, and the stage events, to aid in following the live performance.


