Video Conference Recording With Speaker-Linked Transcripts
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video conferencing systems require manual preparation of verbatim transcripts, which is time-consuming and often fail to accurately identify speakers, making it difficult to determine who made certain speeches during subsequent viewing.
Innovation Solution
A video conferencing system and method that automatically converts voices of different speakers into text content using voice processing algorithms, associates the text with the corresponding speaker, and displays it in an ordered manner, along with a timeline for recording.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual transcription is used to prepare verbatim transcripts, then the accuracy of transcript content can be controlled, but the time consumption increases significantly
Solution Approach 1:
The patent replaces the manual mechanical transcription process with an automated voice processing algorithm that converts speech to text automatically. The system processes audio signals from video conferences and generates transcripts without human intervention, thereby eliminating the time-consuming manual typing process while maintaining acceptable accuracy through algorithmic processing
Solution Approach 2:
The system performs self-service by automatically processing the audio signals from video conferences and generating transcripts autonomously. The voice processing algorithm analyzes the audio input and produces text content without requiring external manual input, enabling the system to serve itself in the transcription task
2Reliability
If manual transcription is used, then the transcript content can be reviewed and corrected, but the ability to identify speakers is lost
Solution Approach 1:
The patent segments the transcript output by associating each text content segment with specific speaker identification information. The system divides the audio signal into segments and processes them individually, maintaining the ability to identify which speaker said what while generating the full transcript, thus preserving speaker identification information that would otherwise be lost in manual transcription
Solution Approach 2:
The system introduces speaker identification as an intermediary element between the audio signal and the text transcript. By processing the audio signal through voice processing algorithms that recognize and identify speakers, the system creates a layered structure where speaker information mediates between the raw audio and the final text output, ensuring that both transcript content and speaker attribution are maintained
3Productivity
If automated voice processing is used to convert audio to text, then the time consumption is reduced, but the complexity of the processing system increases
Solution Approach 1:
The patent implements a universal voice processing algorithm that can handle multiple functions within a single system: audio signal processing, speaker identification, and text generation. By making the processing system multi-functional, the patent reduces the need for separate specialized components for each function, thereby managing system complexity while maintaining high productivity in the transcription process
Data Source
AI summary
A method for recording a video conference and a video conferencing system are provided. The method includes: providing a user interface to a display device, in which the user interface includes a first area, a second area, and a timeline; in response to obtaining an image corresponding to each of multiple participants from a video signal through a person recognition algorithm, displaying the image of each participant in the first area; in response to converting an audio segment of one of the participants obtained from an audio signal into text content through a voice processing algorithm, associating the text content with the corresponding one of the participants, and based on an order of speaking, displaying the text content in the second area; and adjusting a time length of the timeline according to a recording time of the video conference.


