Live Broadcast Script Synchronization via Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Social networking systems lack the ability to provide users with a seamless live-broadcasting experience that allows for real-time interaction and media streaming, particularly in video or audio formats, while enabling users to deliver prepared scripts during live broadcasts.

Innovation Solution

The social networking system implements a live-broadcast service that allows users to broadcast media streams in real-time, providing a script-composing functionality for users to prepare and deliver scripted content during live sessions, with features like script selection, display, and synchronization with vocal expressions, enabling users to broadcast using a single device.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If users broadcast media streams in real-time with script-composing functionality, then user engagement and sense of presence are improved, but device complexity and operation difficulty increase

Engineering Contradiction:
Improveuser engagementVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent combines multiple functions (media streaming, script composition, speech recognition, and script display) into a single integrated live-broadcast service. The script-composing functionality is merged with the broadcasting interface, allowing users to prepare and deliver scripted content through one unified system rather than separate tools, thereby improving engagement without proportionally increasing complexity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The live-broadcast service is designed as a multi-functional platform that handles media capture, script composition, speech-to-text conversion, and real-time script display within a single service. This universal approach allows the same system to perform diverse functions (broadcasting, scripting, recognition, display) that would traditionally require multiple separate applications or devices.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If speech recognition is used to identify words in vocal expressions, then script synchronization is improved, but processing time and system complexity increase

Engineering Contradiction:
Improvescript synchronizationVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-processing and storing transcribed script content before the live broadcast. The speech recognition engine converts vocal expressions to text in advance or in real-time with optimized processing, so that when it's time to synchronize and display the script, the text is already prepared and ready for immediate presentation, reducing latency during the actual broadcast.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If a single device is used for broadcasting with script display, then ease of operation is improved, but display area and information visibility are limited

Engineering Contradiction:
Improvebroadcasting simplicityVSAvoiddisplay area
Core Design Contradiction:
Ease of operationVSArea of stationary object

Solution Approach 1:

The script display is segmented into manageable portions rather than showing the entire script at once. The interface divides the script content into sections that can be displayed sequentially or in focused views, allowing the limited screen real estate to effectively present relevant information without overwhelming the user or requiring a larger display area.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10645460B2Real-time script for live broadcast
Publication Date: 2020.05.05 META PLATFORMS INC
  • US10645460B2 patent drawing
  • US10645460B2 patent drawing
  • US10645460B2 patent drawing

AI summary

In one embodiment, a method includes retrieving, from one or more data stores, a script including multiple text strings, where the script is associated with a user of a social-networking system. The method also includes capturing an incoming media stream including audio data corresponding to vocal expression by the user, where the media stream is transmitted to the social-networking system for broadcast and identifying, using a speech recognition process, one or more words in the vocal expression corresponding to a text string of the script. The method also includes providing the corresponding text string for display in conjunction with a subsequent text string of the script.