Closed Captioning Model Training via Cross-Location Data Merging

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Customizing automatic closed captioning systems for broadcast news can be expensive due to the need for manual transcription of recent broadcast data and continuous updates with new content, especially when prior captioning services are not available at local stations, and existing solutions do not effectively leverage parallel data from other locations.

Innovation Solution

An automatic closed captioning system that uses a base model to request and process relevant closed caption data from multiple data collection locations, computes confidence scores, and selects data subsets for training, allowing for seamless deployment and continuous updates by sharing common news stories across locations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual transcription of local broadcast data is performed for customization, then captioning accuracy is improved, but costs and time consumption increase

Engineering Contradiction:
Improvecaptioning accuracyVSAvoidcost
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent merges training data from multiple broadcast locations into a unified training corpus. Instead of manually transcribing data at each individual location, the system combines available caption data from various stations, allowing all locations to benefit from a shared, larger training dataset that improves overall captioning accuracy without proportionally increasing costs at each site.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal training approach where a single consolidated training corpus serves multiple broadcast locations simultaneously. The customized ASR model trained on this shared data can be deployed across different stations, making the customization process universally applicable and cost-effective rather than requiring separate manual transcription projects for each location.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If manual transcription of local broadcast data is performed for customization, then captioning accuracy is improved, but time consumption increases

Engineering Contradiction:
Improvecaptioning accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent merges training data from multiple broadcast locations into a unified training corpus. Instead of manually transcribing data at each individual location, the system combines available caption data from various stations, allowing all locations to benefit from a shared, larger training dataset that improves overall captioning accuracy without proportionally increasing time consumption at each site.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent performs preliminary consolidation of training data from multiple sources before deployment. By pre-processing and combining caption data from various broadcast locations into a unified corpus in advance, the system eliminates the need for time-consuming manual transcription at each individual location when the customization is actually needed.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If separate data collection is performed at each broadcast location, then local customization accuracy is improved, but device complexity and processing requirements increase

Engineering Contradiction:
Improvelocal customization accuracyVSAvoiddata collection and processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges training data from multiple broadcast locations into a unified training corpus. Instead of manually transcribing data at each individual location, the system combines available caption data from various stations, allowing all locations to benefit from a shared, larger training dataset that improves overall captioning accuracy without proportionally increasing time consumption at each site.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal training approach where a single consolidated training corpus serves multiple broadcast locations simultaneously. The customized ASR model trained on this shared data can be deployed across different stations, making the customization process universally applicable and cost-effective rather than requiring separate manual transcription projects for each location.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11250872B2Using closed captions as parallel training data for customization of closed captioning systems
Publication Date: 2022.02.15 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11250872B2 patent drawing
  • US11250872B2 patent drawing
  • US11250872B2 patent drawing

AI summary

Method, apparatus, and computer program product are provided for customizing an automatic closed captioning system. In some embodiments, at a data use (DU) location, an automatic closed captioning system that includes a base model is provided, search criteria are defined to request from one or more data collection (DC) locations, a search request based on the search criteria is sent to the one or more DC locations, relevant closed caption data from the one or more DC locations are received responsive to the search request, the received relevant closed caption data are processed by computing a confidence score for each of a plurality of data sub-sets of the received relevant closed caption data and selecting one or more of the data sub-sets based on the confidence scores, and the automatic closed captioning system is customized by using the selected one or more data sub-sets to train the base model.