Live Event Captioning via Multi-Source Audio Fingerprinting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In live events, such as sports or concerts, audience members seated far from the action often struggle to understand conversations between performers due to inaudible vocal deliverances, impacting their experience.

Innovation Solution

A system comprising a server and display device that captures and processes audio segments from multiple performers, extracts common caption information, and displays it in real-time, allowing viewers to follow conversations between participants.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If display panels are used to show subtitles or closed captions over videos in live events, then audience members can read text transcriptions of dialogue, but they still cannot understand conversations between performers who are far away

Engineering Contradiction:
Improveinformation about performer conversationsVSAvoidcomprehensibility of vocal deliverance
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent introduces an intermediary system consisting of audio capture devices, processors, and display devices that mediate between the performers' actual conversations and the audience's understanding. The system captures audio segments from multiple performers, processes them to extract verbatim text, and displays this text as captions, serving as a bridge that translates inaudible speech into comprehensible text form.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical/acoustic limitation of human hearing with an electronic information processing system. Instead of relying on the physical propagation of sound waves that diminish with distance, the system uses audio capture devices, digital signal processing, and text display to convey conversation content, substituting the natural acoustic mechanism with an engineered information transmission system.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If audio capture devices are placed close to each performer to capture their vocal deliverance, then clear audio of individual performers can be obtained, but the system complexity and number of devices required increases significantly

Engineering Contradiction:
Improveaudio capture accuracyVSAvoidnumber of audio capture devices
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent makes the audio capture devices universal by enabling them to serve multiple functions: capturing audio from multiple different performers, identifying which performer is speaking, extracting verbatim text from each performer's speech, and displaying the appropriate captions. This multi-functionality reduces the need for dedicated capture devices for each performer.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system employs feedback mechanisms where the processor analyzes audio segments, identifies which performer is speaking based on the captured audio, and uses this identification to retrieve or generate the corresponding verbatim text from the appropriate audio source. This feedback loop enables the system to dynamically switch between capturing and processing audio from different performers using the same hardware infrastructure.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11211073B2Display control of different verbatim text of vocal deliverance of performer-of-interest in a live event
Publication Date: 2021.12.28 SONY GROUP CORP
  • US11211073B2 patent drawing
  • US11211073B2 patent drawing
  • US11211073B2 patent drawing

AI summary

A display system includes a display device and a server. The server receives a plurality of audio segments from a plurality of audio-capture devices. The server receives a user-input that corresponds to a selection of a first user interface (UI) element that represents a first performer-of-interest or a first audio-capture device attached to the first performer-of-interest. The server detects a second performer-of-interest associated with a second audio-capture device within a threshold range of the first audio-capture device. The server extracts a first audio segment of a first vocal deliverance of the first performer-of-interest and a second audio segment of a second vocal deliverance of the second performer-of-interest. The server deduces new caption information from a first verbatim text that is common between the first audio segment and the second audio segment and controls display of the new caption information on the display device.