Voice Recognition Model for Multi-Speaker Meeting Transcription
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for recording meeting contents require a separate stenographer, leading to additional costs and varying quality due to the stenographer's competency, while also being inefficient in processing multiple speakers.
Innovation Solution
A computing device that receives voice data from multiple user terminals, generates integrated voice data, and converts it into a conversation record using a voice recognition model, eliminating the need for a stenographer and improving efficiency in handling multiple speakers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a separate stenographer is used to write meeting minutes, then the conversation record can be obtained, but additional costs occur and the quality varies depending on stenographer competency
Solution Approach 1:
The system uses automatic voice recognition technology to transcribe and process meeting content without requiring a separate stenographer. The voice recognition model autonomously converts speech to text, performs speaker separation, and generates structured meeting minutes, enabling the system to serve itself rather than relying on external human resources.
Solution Approach 2:
The patent replaces the mechanical human process of stenography with an automated electronic system. The voice recognition model, combined with speaker separation algorithms and natural language processing, substitutes the manual typing and listening tasks previously performed by human stenographers, thereby eliminating the need for additional personnel while improving consistency and reliability.
2Reliability
If a separate stenographer is required to write minutes during the meeting, then conversation records can be obtained, but work efficiency is reduced due to additional personnel requirements
Solution Approach 1:
The system autonomously processes meeting content through automated voice recognition and transcription. The voice recognition model continuously transcribes speech, performs real-time speaker separation, and generates structured minutes without requiring human intervention, thereby maintaining conversation record availability while eliminating the productivity loss associated with employing additional stenographers.
Solution Approach 2:
The automated system operates continuously throughout the meeting, processing voice data in real-time without interruption. The voice recognition model maintains continuous transcription and speaker separation operations, ensuring that conversation records are generated throughout the entire meeting duration without the start-stop nature that might occur with human stenographers, thereby improving overall work efficiency.
3Loss of information
If voice data from multiple users is processed, then comprehensive conversation records can be generated, but the complexity of separating and processing multiple speakers increases
Solution Approach 1:
The system applies speaker separation technology that divides the mixed voice data from multiple users into distinct speaker channels. The voice recognition model segments the audio stream by identifying and isolating individual speaker voices, assigning each to a specific user, and processing them separately. This segmentation approach maintains complete conversation content while managing the complexity of multiple speakers through systematic division and individual processing of each speaker's contribution.
Data Source
AI summary
Disclosed is a computer program executable by one or more processors and stored in a computer-readable storage medium, the computer program causing the one or more processors to perform one or more operations below, the operations including: an operation of receiving first voice data from a first user terminal and receiving second voice data from a second user terminal; an operation of generating integrated voice data based on the first voice data and the second voice data; and an operation of generating the integrated voice data as a conversation record by using a voice recognition model.


