Zero-configuration adaptive speaker recognition method and system

Through the zero-configuration adaptive speaker recognition method, voiceprint embedding vectors are extracted from the audio stream in real time for online clustering to generate an identity pool, which solves the problems of high deployment cost and poor real-time performance in traditional technologies, and realizes flexible real-time recognition and efficient adaptation in complex environments.

CN120708626AActive Publication Date: 2025-09-26北京文聿科技有限公司

Patent Information

Application Number
CN202511068328.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-09-26
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

Traditional speaker recognition technology requires pre-registration of voiceprint templates, resulting in high deployment costs and poor flexibility. It is unable to achieve real-time response and adapt to dynamic changes in speakers, and cannot meet the real-time recognition needs of scenarios such as multi-person meetings.

Method used

It adopts a zero-configuration adaptive speaker recognition method, detects voice activity by receiving audio streams, divides them into single-person voice segments, extracts voiceprint embedding vectors for online clustering, generates a speaker identity pool, calculates multi-dimensional fusion similarity, dynamically updates the identity pool, integrates voiceprint, time series, emotion and text features, and supports real-time recognition and adaptation to speaker changes.

Benefits of technology

It realizes real-time speaker recognition without pre-registering voiceprint templates, reduces the difficulty of system deployment, adapts to temporary users and dynamic scenarios, improves the robustness and flexibility of recognition, supports multi-dimensional information fusion, and adapts to complex environments and emotional changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708626A_ABST
    Figure CN120708626A_ABST
Patent Text Reader

Abstract

The invention discloses a zero-configuration adaptive speaker recognition method and system, and relates to the technical field of voice signal processing. The method comprises the following steps: receiving an audio stream, carrying out voice activity detection, obtaining a single-person voice segment, extracting a voiceprint embedding vector of the single-person voice segment, and carrying out online clustering to generate a speaker identity pool; and the identity pool is updated by calculating the multi-dimensional fusion similarity between the remaining single-person voice segments and the voice model, and temporary identity tags corresponding to the single-person voice segments are output. According to the speaker recognition method provided by the invention, a voiceprint template does not need to be registered in advance, and the flexibility and real-time performance of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech signal processing technology, and in particular to a zero-configuration adaptive speaker recognition method and system. Background Art

[0002] In scenarios requiring speech transcription, such as multi-person meetings, interviews, and customer service quality control, accurately distinguishing who said what and when is a crucial core requirement. Traditional speaker recognition technology typically relies on pre-registered voiceprint templates to achieve this goal. This requires collecting user voice data during system initialization, resulting in high deployment costs and limited flexibility.

[0003] Currently, traditional speaker recognition solutions are open-source solutions based on clustering algorithms. These require obtaining complete, long-length recordings of conversations, then globally analyzing and clustering all the voice segments within the recording to ultimately identify the speakers. Their main drawbacks include an inability to meet real-time requirements. For scenarios requiring real-time generation of subtitles or meeting minutes with speaker identification, this "post-analysis" approach is useless, requiring users to wait for a lengthy processing time after the meeting ends before seeing the results.

[0004] In summary, existing technologies are unable to provide a speaker recognition solution that can respond in real time and automatically adapt to dynamic changes in speakers, which to a large extent hinders the application of speaker recognition technology in a wider range of real-world scenarios. Summary of the Invention

[0005] The purpose of this application is to provide a zero-configuration adaptive speaker recognition method and system, aiming to solve the problems of low processing efficiency, high training cost, and inability to achieve real-time speaker recognition in traditional methods.

[0006] According to a first aspect of the embodiments of the present application, a zero-configuration adaptive speaker recognition method is provided, comprising the following steps:

[0007] Receive audio stream;

[0008] Perform voice activity detection on the audio stream and split the audio stream into multiple single-person voice segments;

[0009] Extracting a preset number of single-person voice segments as a first sample set, and the remaining single-person voice segments as a second sample set;

[0010] Obtaining a voiceprint embedding vector corresponding to each single-speaker speech segment in the first sample set, performing online clustering on the voiceprint embedding vectors, and generating a speaker identity pool, the speaker identity pool including multiple temporary identity tags and speech models corresponding to the temporary identity tags;

[0011] Calculate the multi-dimensional fusion similarity between each single-person speech segment and the speech model in the second sample set;

[0012] The speaker identity pool is updated based on the multi-dimensional fusion similarity, and the temporary identity label corresponding to the single-person speech segment is output.

[0013] In other embodiments of the present application, calculating the multi-dimensional fusion similarity between each single-person voice segment in the second sample set and the voice model includes:

[0014] Obtain multi-dimensional information of a single person's speech clip, including voiceprint embedding vector, temporal continuity vector, emotional state vector, and text content coherence vector;

[0015] Determine the weight of each dimension of multi-dimensional information by integrating the decision function;

[0016] According to the weight of each dimension of information, the multi-dimensional fusion similarity between the single-person voice clip and each voice model is calculated.

[0017] Furthermore, the fusion decision function adopts at least one of a weighted average, a Bayesian network or a neural network model, and the weights are adaptively optimized according to the recognition results of the single-person speech segment.

[0018] Furthermore, the voiceprint embedding vectors are clustered online to generate a speaker identity pool, including:

[0019] Using an unsupervised clustering algorithm, the number of clusters is determined, and the voiceprint embedding vectors are divided into clusters;

[0020] A speech model is generated based on the centroid of the voiceprint embedding vector in the cluster, and a temporary identity label is assigned to each cluster.

[0021] Furthermore, the method for determining the number of clusters includes at least one of the following:

[0022] The number of clusters is determined by clustering effectiveness index;

[0023] Adaptively determine the number of clusters through density analysis;

[0024] Dynamically and incrementally modify the number of preset clusters.

[0025] In other embodiments of the present application, updating the speaker identity pool according to the multi-dimensional fusion similarity includes:

[0026] When the multi-dimensional fusion similarity is greater than or equal to a preset threshold, a temporary identity tag corresponding to the single-person voice segment is determined, and the voice model corresponding to the temporary identity tag is updated at a preset learning rate based on the voiceprint embedding vector corresponding to the single-person voice segment;

[0027] When the multi-dimensional fusion similarity is less than the preset threshold, a new temporary identity tag and a new speech model are created and added to the speaker identity pool.

[0028] Furthermore, the method further includes the following steps:

[0029] Provide a user interface that allows an operator to perform at least one of the following actions:

[0030] Rename or bind a temporary identity tag to another identity;

[0031] Corrected the temporary identity label corresponding to the single-person voice clip;

[0032] Merge multiple temporary identity tags;

[0033] Forces the update of the speech model in the speaker identity pool to generate an optimized model.

[0034] Furthermore, a preset number of single-person voice segments are extracted as the first sample set, and the remaining single-person voice segments are extracted as the second sample set, including:

[0035] When the cumulative number of extracted single-person voice segments reaches a preset number, the extracted single-person voice segments are used as the first sample set, and the subsequently extracted single-person voice segments are used as the second sample set;

[0036] Alternatively, a preset number of single-person voice segments are randomly extracted as the first sample set, and the remaining single-person voice segments are used as the second sample set.

[0037] Furthermore, a deep neural network model is used to obtain the voiceprint embedding vector corresponding to each single-person speech segment in the first sample set, and the model parameters of the deep neural network model are dynamically adjusted during the recognition process.

[0038] A second aspect of the embodiments of the present application provides a zero-configuration adaptive speaker recognition method system, including:

[0039] Audio acquisition module, used to receive audio stream;

[0040] A voice activity detection module is used to perform voice activity detection on the audio stream and split the audio stream into multiple single-person voice segments;

[0041] A sample set division module is used to extract a preset number of single-person voice segments as a first sample set and the remaining single-person voice segments as a second sample set;

[0042] A voiceprint embedding extraction module is used to obtain a voiceprint embedding vector corresponding to each single-person speech segment in the first sample set;

[0043] An online clustering module, used to perform online clustering on the voiceprint embedding vectors to generate a speaker identity pool, which includes multiple temporary identity tags and speech models corresponding to the temporary identity tags;

[0044] A multi-dimensional similarity calculation module, used to calculate the multi-dimensional fusion similarity between each single-person voice segment and the voice model in the second sample set;

[0045] The identity pool update module is used to update the speaker identity pool based on the multi-dimensional fusion similarity and output the temporary identity label corresponding to the single-person voice segment.

[0046] Compared with the prior art, the beneficial effects of the embodiments of the present application are as follows: the present application does not require users to pre-register voiceprint templates during the system initialization phase. The system automatically extracts voiceprint embedding vectors from real-time voice streams, and dynamically generates temporary identity tags and voice models through online clustering algorithms to build an identity pool. The present application does not require pre-registration, which reduces the difficulty of system deployment and the user's operating burden. It has low resource consumption and can be used at any time. It can be better applied to temporary users or dynamic scenarios, thereby expanding the flexibility of the application scenarios of speaker recognition. At the same time, the present application uses continuous human-computer collaborative training, and not only relies on voiceprint embedding vectors for speaker recognition, but also integrates multi-dimensional contextual information such as temporal continuity, emotional state, and text coherence to calculate comprehensive similarity, significantly improving recognition robustness in complex environments, and having stronger adaptability to short voice clips and emotional change scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 A flowchart of a zero-configuration adaptive speaker recognition method provided by this application;

[0048] Figure 2 A schematic diagram of the process of generating a speaker identity pool provided for this application;

[0049] Figure 3 A flowchart of the method for calculating the multi-dimensional fusion similarity between each single-person voice segment and the voice model in the second sample set provided by this application;

[0050] Figure 4 A flowchart of an embodiment of a zero-configuration adaptive speaker recognition method provided by the present application;

[0051] Figure 5 This is a structural diagram of a zero-configuration adaptive speaker recognition system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0052] In order to make the technical problems, technical solutions and beneficial effects to be solved by this application more clearly understood, this application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0053] For speaker recognition technology in multi-person conversation scenarios, the current mainstream technical solutions have the following three technical bottlenecks:

[0054] Existing commercial systems generally rely on voiceprint pre-registration. This mechanism requires pre-collecting specific / free speech samples of the target speaker and extracting voiceprint features to establish a voiceprint database tied to the user's identity. First, the implementation requires an additional user registration process, which reduces system availability. Second, in open-ended conversation scenarios (such as when a temporary participant intervenes or an unknown customer calls), the system completely loses its recognition function due to the lack of pre-stored voiceprint data.

[0055] Some open-source solutions use offline clustering algorithms, requiring complete conversation recordings and then performing global cluster analysis of speech segments to distinguish speakers. This process is not real-time and cannot meet time-sensitive requirements such as generating real-time captions for conferences. Users must endure the processing wait time after the recording is completed.

[0056] Some specific recognition algorithms have technical limitations related to a predefined number of speakers, forcing all speech clips to be classified as belonging to a fixed number of speakers. This solution cannot automatically adjust the recognition strategy when the actual number of speakers changes dynamically (for example, when attendees increase or decrease during a meeting), resulting in the incorrect classification of new speakers.

[0057] Existing technical solutions have difficult-to-overcome technical barriers in terms of real-time performance, adaptability, and scenario adaptability, which seriously restricts the promotion and application of this technology in real scenarios.

[0058] Figure 1 A flowchart of a zero-configuration adaptive speaker recognition method provided in the first aspect of the present application is shown, and is described in detail as follows:

[0059] S1. Receive audio stream.

[0060] It should be noted that the method provided in this application can receive a continuous audio input stream in real time, and can also receive a recorded audio stream.

[0061] S2. Perform voice activity detection on the audio stream and divide the audio stream into multiple single-person voice segments.

[0062] For example, the real-time continuous audio stream is preprocessed, with the stream divided into frames according to fixed time windows. Frames can overlap to maintain continuity. Each frame is then high-pass filtered to compensate for high-frequency attenuation in the speech signal and enhance subsequent feature extraction.

[0063] Secondly, the short-term energy of each frame is calculated to reflect the intensity change of the speech signal. The number of times the signal crosses the zero level in each frame is counted to distinguish between unvoiced and voiced sounds.

[0064] Optionally, advanced acoustic features such as Mel-frequency cepstral coefficients, spectral entropy, and formant frequency can be further extracted to improve the accuracy of voice activity detection.

[0065] When the frame energy exceeds a preset first energy threshold, it is determined to be the start of a speech segment;

[0066] When the frame energy is lower than a preset second energy threshold, it is determined that the speech segment ends;

[0067] The first energy threshold and the second energy threshold are dynamically adjusted according to the ambient noise level. For example, the first energy threshold is lowered in a quiet environment to avoid missing weak speech detection; the second energy threshold is increased in a noisy environment to reduce false detection of non-speech signals.

[0068] Starting from the start moment of the speech segment, speech frames are continuously recorded until the end moment of the speech segment is detected, generating an independent single-person speech segment.

[0069] Remove short non-speech frames (such as silent gaps) at the front and back ends to ensure that the segment contains only valid speech content; merge adjacent short speech segments to avoid segment splitting due to short pauses.

[0070] As a preferred technical solution, multi-channel audio streams can be processed in parallel to support real-time segmentation in multi-person conversation scenarios, and a streaming computing framework that processes while receiving can be adopted to reduce latency and ensure real-time performance.

[0071] As an optional technical solution, a lightweight neural network can be trained to input multiple frames of acoustic features and output the probability of judging "speech fragments / non-speech fragments", further improving detection accuracy in complex scenarios. It is also possible to use the judgment results of historical frames and the characteristics of the current frame to optimize the stability of speech and non-speech fragments through a state machine (such as a hidden Markov model).

[0072] S3. Extract a preset number of single-person voice segments as a first sample set, and the remaining single-person voice segments as a second sample set.

[0073] The segmented single-person voice segments are distinguished. Specifically for real-time voice and recorded voice, this application provides two different solutions for sample distinction.

[0074] If the acquired audio stream is a real-time audio stream, extract the number of single-person voice segments in sequence according to the time sequence in which the single-person voice segments are acquired. When the number of extracted single-person voice segments reaches a preset number, use the extracted single-person voice segments as the first sample set, and use the subsequently extracted single-person voice segments as the second sample set;

[0075] If the acquired audio stream is a recorded audio stream, all acquired single-person voice segments may be randomly extracted, and a preset number of single-person voice segments may be extracted as the first sample set, and the remaining unextracted single-person voice segments may be extracted as the second sample set.

[0076] S4. Obtain a voiceprint embedding vector corresponding to each single-person speech segment in the first sample set, perform online clustering on the voiceprint embedding vectors, and generate a speaker identity pool. The speaker identity pool includes multiple temporary identity tags and speech models corresponding to the temporary identity tags.

[0077] It should be noted that the speech model is a deep learning model that includes at least the centroid of the speaker's voiceprint feature vector cluster. For details, see Figure 2 , obtaining the voiceprint embedding vector corresponding to each single-person speech segment in the first sample set, performing online clustering on the voiceprint embedding vectors, and generating a speaker identity pool, including the following steps:

[0078] S401: Perform voiceprint feature extraction on the segmented speech segments to generate a fixed-dimensional voiceprint embedding vector.

[0079] Exemplarily, the voiceprint feature extraction method in the present application can use a pre-trained neural network model to extract high-dimensional voiceprint features, and the vector represents the acoustic characteristics of the speaker; or, extract Mel-frequency cepstral coefficients, spectral centroid, resonance peaks and other features and convert them into fixed-length vectors.

[0080] The voiceprint embedding vector is normalized so that the vectors of speech with different lengths have comparable modulus lengths.

[0081] S402: Determine the number of clusters by using an unsupervised clustering algorithm, and divide the voiceprint embedding vectors into the clusters.

[0082] Specifically, the method for determining the number of clusters includes at least one of the following methods:

[0083] The number of clusters is determined by clustering effectiveness indicators, and the number of clusters can be determined by, for example, the elbow rule, silhouette coefficient, and gap statistics.

[0084] The number of clusters is adaptively determined by density analysis, and the number of clusters can be determined by density peak clustering or density-based clustering.

[0085] The number of preset clusters can be dynamically incrementally modified by setting the initial number of clusters and then dynamically adjusting the initial number of clusters according to new data to determine the number of clusters.

[0086] S403: Generate a speech model according to the centroid of the voiceprint embedding vector in the cluster, and assign a temporary identity label to each cluster.

[0087] For each cluster, its centroid is calculated as the voiceprint feature representative of the cluster.

[0088] Centroid calculation formula:

[0089] Among them, centroid i is the centroid vector of the i-th cluster;

[0090] n i is the number of voiceprint vectors contained in the i-th cluster;

[0091] vector j is the j-th voiceprint embedding vector in the i-th cluster.

[0092] If there are 10 voiceprint vectors in a cluster (each dimension is 512), the centroid is the dimension-wise average of these 10 vectors.

[0093] The centroid vector of each cluster is used as the speech model;

[0094] It should be noted that the model can be supplemented with other meta-information (such as the interval distribution of the speech segments corresponding to the model, the emotion vector representation corresponding to the model, and the text content coherence score, etc.).

[0095] Generate a unique temporary identity tag for each cluster for subsequent identification and tracking.

[0096] S5. Calculate the multi-dimensional fusion similarity between each single-person voice segment in the second sample set and the voice model.

[0097] In the zero-configuration adaptive speaker recognition system, in order to improve the accuracy and robustness of recognition, a single voiceprint feature is often not enough to cope with complex scenarios (such as noise, short speech, emotional changes). Therefore, this application can also comprehensively represent the voice segment by extracting multi-dimensional information (voiceprint embedding vector, temporal continuity vector, emotional state vector, text content coherence vector).

[0098] For details, see Figure 3 , including the following steps:

[0099] S501: Acquire multi-dimensional information of a single-person voice segment.

[0100] Specifically, the multi-dimensional information includes voiceprint embedding vector, temporal continuity vector, emotional state vector and text content coherence vector.

[0101] After performing pre-processing such as framing, pre-emphasis, and windowing on the single-person voice clips in the second sample set;

[0102] Calculate Mel-frequency cepstral coefficients or spectral features;

[0103] A deep neural network model is used to map spectral features into high-dimensional embedding vectors.

[0104] Each single-person speech segment generates a fixed-length voiceprint embedding vector to represent the identity characteristics of the speaker of the current single-person speech segment.

[0105] The temporal continuity vector is scored by counting the time intervals between the current speech segment and the previous n recognized speech segments, and calculating the mean, standard deviation, maximum interval and other indicators between the speech segments to form a temporal continuity feature;

[0106] This allows for better identification of the same speaker in a continuous conversation or differentiation of switching between different speakers.

[0107] Use speech emotion recognition models to analyze the acoustic characteristics (such as pitch, energy, speaking rate) and spectral characteristics of speech;

[0108] categorize emotions into discrete categories or emotion intensity scores;

[0109] The emotion category or emotion intensity score is mapped to an emotion state vector. If an emotion change is detected (such as from calm to excited), the weight of the voiceprint similarity calculation may be adjusted.

[0110] Use an automatic speech recognition model to convert speech clips into text, and extract the semantic embedding vector of the text through natural language processing technology; calculate the semantic similarity between the current text and historical recognition text, or analyze the topic consistency; and combine indicators such as semantic similarity and topic consistency to form a text content coherence vector.

[0111] When the voiceprint features are interfered with by noise, the same speaker can be confirmed through the text content, or the specific language habits of different speakers can be distinguished.

[0112] S502: Determine the weight of each dimension of information in the multi-dimensional information through a fusion decision function.

[0113] Concatenate or weight the four dimensional vectors into a comprehensive feature vector;

[0114] Dynamically adjust the weight of each dimension according to the environment (such as noise level, speech length);

[0115] Specifically, the general form of the fusion decision function can be expressed as:

[0116] W=f(E,S,T,C)

[0117] Among them, W is the dimension weight vector, E, S, T, and C are the voiceprint embedding vector, temporal continuity vector, emotional state vector, and text content coherence vector respectively; f() is the fusion function, which can be designed as a linear weighted, nonlinear model (such as a neural network) or rule engine according to specific needs.

[0118] Specifically, a fixed weight template can be set based on prior knowledge. For example:

[0119] High-noise scenario: W = [0.4, 0.2, 0.1, 0.3], reducing the voiceprint weight and increasing the text weight.

[0120] Short speech scene: W = [0.3, 0.4, 0.1, 0.2] to enhance temporal continuity.

[0121] According to the recognition accuracy, misjudgment rate and other indicators, the fixed weight template is updated by the gradient descent method:

[0122]

[0123] Among them, η is the learning rate, Loss is the recognition error function, W t is the weight of the current iteration, Y is the true value, is the predicted value.

[0124] Alternatively, the weights can be regarded as hyperparameters, and the optimal weight combination can be searched through a Bayesian optimization algorithm (such as Gaussian process regression); or, a multi-input neural network can be constructed, with features of each dimension input into independent branches, and finally the weights are calculated through a fusion layer (such as weighted summation, attention mechanism).

[0125] S503: Calculate the multi-dimensional fusion similarity between the single-person voice segment and each voice model based on the weight of each dimension of information.

[0126] Specifically, the similarity calculation of single-dimensional information is performed, and the similarity score is determined based on the weight vector of each dimension.

[0127] S6. Update the speaker identity pool based on the multi-dimensional fusion similarity and output the temporary identity label corresponding to the single-person voice segment.

[0128] Specifically, if the similarity score obtained by the highest similarity model is greater than or equal to a preset threshold, the information of the voice segment (such as voiceprint vector, time series features, etc.) is merged into the corresponding model;

[0129] Update model parameters based on the information of the speech segment, such as weighted average voiceprint embedding vector, cumulative time interval statistics, updated emotion / text feature distribution, etc.

[0130] Directly returns the matching existing identity tag.

[0131] If the similarity score obtained by the highest similarity model is less than the preset threshold, it is determined to be a new speaker, and a new temporary identity tag is generated. The multi-dimensional information of the current voice segment is used as the voice model parameters of the new speaker, and the voice model of the new speaker is inserted at the end of the speaker identity pool and marked as "temporarily added state". If the same temporary tag is matched multiple times in subsequent recognition, the "temporarily added state" mark is removed.

[0132] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0133] In some embodiments, the present application may further include the following steps:

[0134] Provide a user interface that allows an operator to perform at least one of the following actions:

[0135] Rename or bind a temporary identity tag to another identity;

[0136] Corrected the temporary identity label corresponding to the single-person voice clip;

[0137] Merge multiple temporary identity tags;

[0138] Forces the update of the speech model in the speaker identity pool to generate an optimized model.

[0139] Specifically, the present application provides an optional manual supervision and identity management interface that runs in parallel with the real-time recognition process. Operators (such as conference moderators or recorders) can use this interface to intervene in and manage the system's automatic recognition results.

[0140] Operators can perform management operations including at least one or more of the following:

[0141] Identity renaming, for example: changing the anonymous label automatically generated by the system (such as speaker_A) to a real identity with clear semantics (such as Zhang San).

[0142] Attribution correction: For example, when the system incorrectly attributes a voice clip to Speaker_A, the operator can manually correct it to the correct Speaker_B.

[0143] Identity merging. For example, when the system mistakenly identifies the same person as Speaker_C and Speaker_D, the operator can merge the two identities into one.

[0144] Each supervision operation by the user will be regarded as a supervision signal with the highest priority by the system, and will be used to force the update of the speech model in the "speaker identity pool". Specifically:

[0145] For the attribution correction operation, the system will embed the voiceprint of the speech clip into a vector, subtract it from the model of the original wrong identity (if it has been updated), and update it to the model of the correct identity with a high weight to achieve instant error correction of the model.

[0146] For the identity merging operation, the system will fuse all the voiceprint information under the two identities to generate a more robust and comprehensive new voice model.

[0147] Through this closed-loop optimization of human-machine collaboration, the system's long-term recognition accuracy and robust performance are continuously and rapidly improved.

[0148] refer to Figure 4 , showing the complete process of a voiceprint recognition system:

[0149] The system first processes the continuous audio stream, identifies valid speech segments through voice activity detection (VAD), and extracts voiceprint features to form an embedding vector.

[0150] The extracted voiceprint embedding vectors are clustered to initially form M identity models representing different speakers.

[0151] For a new voiceprint embedding vector, the system calculates its similarity with the existing identity model and decides whether to assign it to an existing identity or create a new identity model.

[0152] Finally, the voice models in the identity pool are managed through manual supervision to ensure the accuracy and reliability of the system and achieve closed-loop optimization.

[0153] See also Figure 5 , shows a zero-configuration adaptive speaker recognition method system, including:

[0154] Audio acquisition module, used to receive audio stream;

[0155] A voice activity detection module is used to perform voice activity detection on the audio stream and split the audio stream into multiple single-person voice segments;

[0156] A sample set division module is used to extract a preset number of single-person voice segments as a first sample set and the remaining single-person voice segments as a second sample set;

[0157] A voiceprint embedding extraction module is used to obtain a voiceprint embedding vector corresponding to each single-person speech segment in the first sample set;

[0158] An online clustering module, used to perform online clustering on the voiceprint embedding vectors to generate a speaker identity pool, which includes multiple temporary identity tags and speech models corresponding to the temporary identity tags;

[0159] A multi-dimensional similarity calculation module, used to calculate the multi-dimensional fusion similarity between each single-person voice segment and the voice model in the second sample set;

[0160] The identity pool update module is used to update the speaker identity pool based on the multi-dimensional fusion similarity and output the temporary identity label corresponding to the single-person voice segment.

[0161] The user interface is used to allow the user to perform at least one of the following operations:

[0162] Rename or bind a temporary identity tag to another identity;

[0163] Fixed temporary identity tags for single-player voice clips;

[0164] Merge multiple temporary identity tags;

[0165] Forces an update of the speech model in the speaker identity pool to produce an optimized model.

[0166] Based on the above embodiments, it can be seen that this application can achieve efficient multi-person conversation speaker recognition in a zero-configuration, adaptive manner, without the need for pre-registered voiceprints, automatically build an initial identity pool, and achieve rapid deployment and startup. At the same time, this application adopts an incremental update mechanism that can adapt to dynamic changes in speakers (such as joining / exiting midway) in real time, and integrate voiceprints, timing, emotions, and text features to automatically adjust the speech model and recognition threshold to further improve recognition accuracy. It can also achieve closed-loop optimization through a manual supervision interface, thereby continuously improving long-term recognition accuracy.

[0167] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A zero-configuration adaptive speaker recognition method, characterized in that: The method comprises the following steps: Receive audio stream; Performing voice activity detection on the audio stream, and dividing the audio stream into multiple single-person voice segments; Extracting a preset number of the single-person voice segments as a first sample set, and the remaining single-person voice segments as a second sample set; Obtaining a voiceprint embedding vector corresponding to each of the single-speaker speech segments in the first sample set, performing online clustering on the voiceprint embedding vectors to generate a speaker identity pool, the speaker identity pool including a plurality of temporary identity tags and speech models corresponding to the temporary identity tags; Calculating the multi-dimensional fusion similarity between each of the single-person voice segments in the second sample set and the voice model; The speaker identity pool is updated according to the multi-dimensional fusion similarity, and a temporary identity tag corresponding to the single-speaker voice segment is output.

2. The method according to claim 1, wherein The calculating of the multi-dimensional fusion similarity between each of the single-speaker voice segments in the second sample set and the voice model includes: Acquiring multi-dimensional information of the single-person voice segment, the multi-dimensional information including a voiceprint embedding vector, a temporal continuity vector, an emotional state vector, and a text content coherence vector; Determining the weight of each dimension of information in the multi-dimensional information by a fusion decision function; According to the weight of each dimension of information, the multi-dimensional fusion similarity between the single-person voice segment and each of the voice models is calculated.

3. The method according to claim 2, wherein The fusion decision function adopts at least one of a weighted average, a Bayesian network or a neural network model, and the weight is adaptively optimized according to the recognition result of the single-person voice segment.

4. The method according to claim 1, wherein The online clustering of the voiceprint embedding vectors to generate a speaker identity pool includes: Determine the number of clusters by using an unsupervised clustering algorithm, and divide the voiceprint embedding vectors into the clusters; A speech model is generated according to the centroid of the voiceprint embedding vectors in the clusters, and a temporary identity label is assigned to each of the clusters.

5. The method according to claim 4, wherein The method for determining the number of clusters includes at least one of the following: Determining the number of clusters by clustering effectiveness index; Adaptively determining the number of clusters through density analysis; Dynamically and incrementally modify the number of preset clusters.

6. The method according to claim 1, wherein The updating of the speaker identity pool according to the multi-dimensional fusion similarity includes: When the multi-dimensional fusion similarity is greater than or equal to a preset threshold, determining a temporary identity tag corresponding to the single-person voice segment, and updating the voice model corresponding to the temporary identity tag at a preset learning rate based on the voiceprint embedding vector corresponding to the single-person voice segment; When the multi-dimensional fusion similarity is less than a preset threshold, a new temporary identity tag and a new speech model are created and added to the speaker identity pool.

7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises the following steps: Provide a user interface that allows an operator to perform at least one of the following actions: Renaming or binding the temporary identity tag to an identity; Correcting the temporary identity tag corresponding to the single-person voice segment; Merge multiple temporary identity tags; The speech model in the speaker identity pool is forcibly updated to generate an optimized model.

8. The method according to claim 7, wherein The extracting of a preset number of the single-person voice segments as a first sample set and the remaining single-person voice segments as a second sample set includes: When the cumulative number of the extracted single-person voice segments reaches a preset number, the extracted single-person voice segments are used as the first sample set, and the subsequently extracted single-person voice segments are used as the second sample set; Alternatively, a preset number of the single-person voice segments are randomly extracted as the first sample set, and the remaining single-person voice segments are used as the second sample set.

9. The method according to claim 8, wherein A deep neural network model is used to obtain the voiceprint embedding vector corresponding to each of the single-person voice segments in the first sample set, and the model parameters of the deep neural network model are dynamically adjusted during the recognition process.

10. A zero-configuration adaptive speaker recognition method system, comprising: Audio acquisition module, used to receive audio stream; a voice activity detection module, configured to perform voice activity detection on the audio stream and segment the audio stream into a plurality of single-person voice segments; a sample set division module, configured to extract a preset number of the single-person voice segments as a first sample set, and the remaining single-person voice segments as a second sample set; A voiceprint embedding and extraction module, configured to obtain a voiceprint embedding vector corresponding to each of the single-person voice segments in the first sample set; An online clustering module, configured to perform online clustering on the voiceprint embedding vectors to generate a speaker identity pool, wherein the speaker identity pool includes a plurality of temporary identity tags and speech models corresponding to the temporary identity tags; A multi-dimensional similarity calculation module, configured to calculate a multi-dimensional fusion similarity between each of the single-speaker voice segments in the second sample set and the voice model; An identity pool updating module, configured to update the speaker identity pool according to the multi-dimensional fusion similarity and output a temporary identity tag corresponding to the single-speaker voice segment; The user interface is used to allow the user to perform at least one of the following operations: Rename or bind a temporary identity tag to another identity; Correcting the temporary identity tag of the single-person voice segment; Merge multiple temporary identity tags; The speech models in the speaker identity pool are forced to be updated to generate an optimized model.

Citation Information

Patent Citations

  • Speaker recognition method and device based on clustering, equipment and storage medium

    CN113851136A

  • Speech recognition method and device, computer equipment and storage medium

    CN119559934A

  • Streaming speaker log method and system

    CN119673173A

  • Speaker recognition method based on bipartite graph matching and electronic equipment

    CN119993166A

  • Voiceprint recognition method, model construction method and system for authority verification

    CN120236590A

Cited By

  • Speech recognition method and system

    CN121306142A

  • Online speaker affiliation method and system based on voiceprint recognition

    CN121641032A

  • Marketing scene speaker distinguishing method and device based on multi-modal fusion algorithm

    CN121768403A