Video Conference Recording With Speaker-Linked Transcripts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video conferencing systems require manual preparation of verbatim transcripts, which is time-consuming and often fail to accurately identify speakers, making it difficult to determine who made certain speeches during subsequent viewing.

Innovation Solution

A video conferencing system and method that automatically converts voices of different speakers into text content using voice processing algorithms, associates the text with the corresponding speaker, and displays it in an ordered manner, along with a timeline for recording.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual transcription is used to prepare verbatim transcripts, then the accuracy of transcript content can be controlled, but the time consumption increases significantly

Engineering Contradiction:
Improvetranscript accuracyVSAvoidtranscription time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical transcription process with an automated voice processing algorithm that converts speech to text automatically. The system processes audio signals from video conferences and generates transcripts without human intervention, thereby eliminating the time-consuming manual typing process while maintaining acceptable accuracy through algorithmic processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-service by automatically processing the audio signals from video conferences and generating transcripts autonomously. The voice processing algorithm analyzes the audio input and produces text content without requiring external manual input, enabling the system to serve itself in the transcription task

Inventive Principle:
Principle #25Self-service

2Reliability

If manual transcription is used, then the transcript content can be reviewed and corrected, but the ability to identify speakers is lost

Engineering Contradiction:
Improvetranscript review capabilityVSAvoidspeaker identification
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent segments the transcript output by associating each text content segment with specific speaker identification information. The system divides the audio signal into segments and processes them individually, maintaining the ability to identify which speaker said what while generating the full transcript, thus preserving speaker identification information that would otherwise be lost in manual transcription

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces speaker identification as an intermediary element between the audio signal and the text transcript. By processing the audio signal through voice processing algorithms that recognize and identify speakers, the system creates a layered structure where speaker information mediates between the raw audio and the final text output, ensuring that both transcript content and speaker attribution are maintained

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If automated voice processing is used to convert audio to text, then the time consumption is reduced, but the complexity of the processing system increases

Engineering Contradiction:
Improvetranscription speedVSAvoidprocessing system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements a universal voice processing algorithm that can handle multiple functions within a single system: audio signal processing, speaker identification, and text generation. By making the processing system multi-functional, the patent reduces the need for separate specialized components for each function, thereby managing system complexity while maintaining high productivity in the transcription process

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12500994B2Method for recording video conference and video conferencing system
Publication Date: 2025.12.16 MERRY ELECTRONICS (SHENZHEN) CO LTD
  • US12500994B2 patent drawing
  • US12500994B2 patent drawing
  • US12500994B2 patent drawing

AI summary

A method for recording a video conference and a video conferencing system are provided. The method includes: providing a user interface to a display device, in which the user interface includes a first area, a second area, and a timeline; in response to obtaining an image corresponding to each of multiple participants from a video signal through a person recognition algorithm, displaying the image of each participant in the first area; in response to converting an audio segment of one of the participants obtained from an audio signal into text content through a voice processing algorithm, associating the text content with the corresponding one of the participants, and based on an order of speaking, displaying the text content in the second area; and adjusting a time length of the timeline according to a recording time of the video conference.