Speaker Change Point Detection Through Continuous Integrate-and-Fire Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speaker change point detection methods in voice interactions are inaccurate and complex, failing to detect rapid speaker changes and optimize overall performance due to their pipelined structure and inability to simulate human brain processing.

Innovation Solution

A method utilizing a continuous integrate-and-fire (CIF) mechanism to integrate and fire speaker characterization vectors frame by frame, encoding acoustic features with a bi-directional LSTM network to determine speaker change points accurately.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional pipelined speaker change point detection methods are used, then the detection process is simpler to implement, but the detection accuracy is low and rapid speaker changes cannot be detected

Engineering Contradiction:
Improvespeaker change point detection accuracyVSAvoiddetection mechanism complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical pipelined processing with a neural network-based continuous integrate-and-fire mechanism that simulates human brain processing. This substitution enables the system to detect rapid speaker changes by continuously integrating acoustic features frame-by-frame, achieving higher detection accuracy without being constrained by fixed pipeline stages

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces dynamic processing by continuously updating speaker characterization vectors frame-by-frame through the integrate-and-fire mechanism. This dynamic approach allows the system to adapt to rapid speaker changes in real-time, contrasting with static pipelined methods that process audio in fixed stages

Inventive Principle:
Principle #15Dynamics

2Productivity

If pipelined detection methods are used, then the system structure is simpler, but the overall performance cannot be optimized due to sequential processing limitations

Engineering Contradiction:
Improvedetection performanceVSAvoidprocessing mechanism complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements continuous processing by continuously integrating acoustic features and updating speaker characterization vectors frame-by-frame without sequential pipeline interruptions. This continuous action allows the system to maintain optimal detection performance throughout the entire audio stream, unlike pipelined methods that process segments sequentially

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The patent performs preliminary encoding of acoustic features into speaker characterization vectors using a neural network before the integrate-and-fire detection process. This preliminary action prepares the data in an optimized format that enables more efficient and accurate subsequent detection operations

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If traditional methods are used, then the implementation is more straightforward, but rapid speaker changes cannot be detected accurately

Engineering Contradiction:
Improverapid speaker change detection accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The continuous integrate-and-fire mechanism processes audio frames without interruption, enabling immediate detection of rapid speaker changes as they occur. This eliminates the time loss associated with pipelined methods that must complete fixed processing stages before detecting changes

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS12374339B2Method, apparatus, device, and storage medium for speaker change point detection
Publication Date: 2025.07.29 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • US12374339B2 patent drawing
  • US12374339B2 patent drawing
  • US12374339B2 patent drawing

AI summary

A method, apparatus, device, and storage medium for speaker change point detection, the method including: acquiring target voice data to be detected; and extracting an acoustic feature characterizing acoustic information of the target voice data from the target voice data; encoding the acoustic feature to obtain speaker characterization vectors of the target voice data; integrating and firing the speaker characterization vectors of the target voice data based on a continuous integrate-and-fire CIF mechanism, to obtain a sequence of speaker characterizations in the target voice data; and determining the speaker change points, according to the sequence of the speaker characterizations bounded by the speaker change points in the target voice data. This method can effectively improve the accuracy of the detection result of a speaker change point in target voice data with a type of interaction.