Token-Block Voice Reconstruction for Chronological Speaker Flow

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice recognition technologies struggle to accurately reconstruct speaker-specific voice recognition data in a conversation format, especially when multiple speakers overlap, leading to inaccurate identification and disrupted conversation flow.

Innovation Solution

A method and apparatus that divide speaker-specific voice recognition data into blocks using predefined criteria, arrange them chronologically, merge continuous utterances of the same speaker, and reconstruct the data in a conversation format to maintain the natural flow of the conversation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speaker-specific voice recognition data is acquired for each speaker in a voice conversation, then speaker identification accuracy is improved, but the complexity of reconstructing the data in conversation format increases

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoiddata reconstruction complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments speaker-specific voice recognition data into individual speaker streams, processes each speaker's data separately through blocking and merging operations, and then reconstructs the conversation format by combining the processed segments. This segmentation approach maintains speaker identification accuracy while managing reconstruction complexity through systematic processing of divided data units.

Inventive Principle:
Principle #1Segmentation

2Speed

If voice recognition is performed in real-time with partial results generated every predetermined time, then response speed is improved, but the accuracy of conversation reconstruction deteriorates due to incomplete speech recognition

Engineering Contradiction:
Improveresponse speedVSAvoidconversation reconstruction accuracy
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The patent performs preliminary blocking of speaker-specific voice recognition data into token-based units before final conversation reconstruction. This preliminary organization of data into manageable blocks allows for more accurate merging and reconstruction operations, improving conversation reconstruction accuracy while maintaining real-time processing capabilities through efficient pre-structuring of the data.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If multiple speakers speak simultaneously causing voice overlap, then conversation naturalness is improved, but voice recognition accuracy deteriorates

Engineering Contradiction:
Improveconversation naturalnessVSAvoidvoice recognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent segments voice recognition data by speaker identity, creating separate data streams for each speaker even when they speak simultaneously. This speaker-specific segmentation allows the system to process and reconstruct overlapping speech more accurately by maintaining distinct speaker boundaries, thereby improving voice recognition accuracy while preserving the naturalness of simultaneous conversation.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12431150B2Method and apparatus for reconstructing voice conversation
Publication Date: 2025.09.30 LLSOLLU CO LTD
  • US12431150B2 patent drawing
  • US12431150B2 patent drawing
  • US12431150B2 patent drawing

AI summary

A voice conversation reconstruction method performed by a voice conversation reconstruction apparatus is disclosed. The method includes acquiring speaker-specific voice recognition data about voice conversation, dividing the speaker-specific voice recognition data into a plurality of blocks using a boundary between tokens according to a predefined division criterion, arranging the plurality of blocks in chronological order irrespective of a speaker, merging blocks from continuous utterance of the same speaker among the arranged plurality of blocks, and reconstructing the plurality of blocks subjected to the merging in a conversation format in chronological order and based on a speaker.