A speaker voice recognition method

Through multi-task model and feature weighting fusion technology, the accuracy problem of target speaker recognition in noisy environments is solved, and high-precision speaker recognition in complex scenarios is achieved.

CN119993154BActive Publication Date: 2025-07-08SHENZHEN HUOLI TIAN HUI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510480115.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-08
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

现有技术在嘈杂环境下目标说话人识别困难,尤其是周围有其他人声噪音时,难以准确判断目标声音。

Method used

A multi-task model is used to identify speakers, and a feature-weighted fusion is performed through speech activity detection, proximity judgment and target speaker judgment head, dynamic weight adjustment is performed in combination with attention mechanism, and speech features are extracted using convolutional neural network and Transformer structure, and the model is optimized through multi-task loss function.

Benefits of technology

The accuracy and robustness of the target speaker's judgment are improved, especially in long-term calls and complex scenarios to maintain high recognition accuracy, and dynamically adjust feature weights to adapt to the changes in the speaker's feature.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993154B_ABST
    Figure CN119993154B_ABST
Patent Text Reader

Abstract

The present application relates to a speaker voice recognition method, which relates to the technical field of telephone voice signal processing. The method includes: obtaining the call audio to be recognized in real time, and segmenting the call audio to be recognized into a plurality of audio blocks; for each of the audio blocks, extracting the voice features of the audio block and inputting them into a pre-trained speaker voice recognition model; outputting, by the speaker voice recognition model, the attribution probability that the audio block belongs to the target speaker; and if the attribution probability that the audio block belongs to the target speaker is greater than a preset speaker attribution probability threshold, determining that the audio block belongs to the target speaker. By using the present application, the recognition of the target speaker in a complex scenario can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of telephone voice signal processing, and particularly to a method for speaker voice recognition. Background Art

[0002] Currently, with the development of technologies such as speech recognition, intelligent assistants, and remote communication, the demand for target speaker recognition is increasing continuously. Existing technologies usually rely on voice activity detection (VAD) to determine whether an audio block has sound, and combine speaker verification (SV) for identity recognition. However, when the surrounding noise is noisy, the judgment is difficult. Existing human voice models have shown relatively good performance in filtering environmental noise, but if there are other human voice noises around, it will interfere with the judgment of the target voice. Summary of the Invention

[0003] Based on this, in view of the above technical problems, it is necessary to provide a method for speaker voice recognition, and the method includes:

[0004] Obtain the call audio to be recognized in real time, and segment the call audio to be recognized into several audio blocks;

[0005] For each of the audio blocks, extract the voice features of the audio block and input them into a pre-trained speaker voice recognition model;

[0006] Output the attribution probability that the audio block belongs to the target speaker through the speaker voice recognition model;

[0007] If the attribution probability that the audio block belongs to the target speaker is greater than a preset speaker attribution probability threshold, it is determined that the audio block belongs to the target speaker.

[0008] As an optional implementation manner, the training process of the speaker voice recognition model includes:

[0009] Obtain historical call audio, and segment the historical call audio into several audio blocks;

[0010] Perform training annotation on each of the audio blocks;

[0011] For each of the audio blocks, extract the voice features of the audio block;

[0012] Construct the speaker voice recognition model, and the speaker voice recognition model includes an encoder module and a multi-task output module;

[0013] Jointly optimize the speaker voice recognition model using a multi-task loss function.

[0014] As an optional implementation manner, the performing training annotation on each of the audio blocks includes:

[0015] Perform vocal state annotation on each of the audio chunks, where the vocal state includes a vocal state and a silent state;

[0016] Based on a pre-trained NEAR pre-model, perform proximal annotation on each of the audio chunks according to the sound energy, reverberation degree, and high-frequency retention of the audio chunk;

[0017] Perform target speaker annotation on each of the audio chunks.

[0018] As an optional implementation, the speech features include acoustic features, distance-related features, and high-level speech features.

[0019] As an optional implementation, the encoder module includes a convolutional neural network layer and a Transformer structure; where

[0020] The convolutional neural network layer is used to refine and fuse the speech features, capture time-frequency patterns, vocalization features, and reverberation information;

[0021] The Transformer structure is used to maintain context information using the hidden states of previous audio chunks.

[0022] As an optional implementation, the multi-task output module includes a voice activity detection head, a proximal speech judgment head, a speaker embedding extraction head, and a target speaker judgment head; where

[0023] The voice activity detection head is used to output the voice activity probability of the audio chunk;

[0024] The proximal speech judgment head is used to output the proximal probability of the audio chunk;

[0025] The speaker embedding extraction head is used to output the speaker embedding vector of the audio chunk;

[0026] The target speaker judgment head is used to comprehensively consider the voice activity probability, the proximal probability, the speaker embedding vector, and a preset target speaker reference vector, and output the attribution probability that the audio chunk belongs to the target speaker.

[0027] As an optional implementation, the preset target speaker reference vector is the global embedding mean of the speaker embedding vectors in the proximal audio chunks identified during the call.

[0028] As an optional implementation, the target speaker judgment head is further used to:

[0029] Calculate the similarity score between the speaker embedding vector of the audio chunk to be recognized and the target speaker reference vector;

[0030] Based on the multi-dimensional feature weighted fusion mechanism, the speech activity probability, the proximal probability, and the similarity score are weighted and calculated to determine the attribution probability that the audio block belongs to the target speaker.

[0031] As an alternative implementation, the method further includes:

[0032] An attention mechanism is introduced into the target speaker judgment head. By calculating the attention weight between the speaker embedding vector of the audio block and the target speaker reference vector, feature weighted fusion is performed on the target judgment to improve the accuracy of audio block attribution judgment.

[0033] As an alternative implementation, the loss function includes a speech activity detection loss function, a proximal judgment loss function, an embedding extraction loss function, and a target speaker judgment loss function; wherein,

[0034] The speech activity detection loss function is used to distinguish whether the audio block is vocal or silent;

[0035] The proximal judgment loss function is used to distinguish whether the audio block is proximal or distal;

[0036] The embedding extraction loss function is used to optimize the embedding discrimination;

[0037] The target speaker judgment loss function maximizes the attribution probability of the audio block of the target speaker and at the same time minimizes the misjudgment probability of the audio blocks of other speakers.

[0038] In a second aspect, a computer device is provided, including a memory and a processor. A computer program that can run on the processor is stored on the memory. When the processor executes the computer program, the method steps described in any item of the first aspect are implemented.

[0039] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method steps described in any item of the first aspect are implemented.

[0040] The present application provides a speaker voice recognition method, the method comprising: acquiring in real time a call audio to be recognized, and segmenting the call audio to be recognized into a plurality of audio chunks; for each of the audio chunks, extracting the voice features of the audio chunk and inputting same into a pre-trained speaker voice recognition model; outputting, by the speaker voice recognition model, the attribution probability that the audio chunk belongs to a target speaker; and if the attribution probability that the audio chunk belongs to the target speaker is greater than a preset speaker attribution probability threshold, determining that the audio chunk belongs to the target speaker. The technical solutions provided by the embodiments of the present application at least bring the following beneficial effects: In the traditional technology, the determination of the target speaker usually relies on a single feature and cannot make full use of multi-dimensional information, resulting in low recognition accuracy. By introducing a multi-task model, the present application performs feature weighted fusion on P vad (voice activity), P near (proximal determination), and S (embedding similarity), realizes the adaptive adjustment of different features at different call stages, and improves the accuracy of target speaker determination. The traditional method uses a fixed reference vector for matching and cannot perform adaptive adjustment according to the feature changes of the target speaker during the call, resulting in a decrease in recognition accuracy during long calls. By calculating the reference vector C ref of the target speaker and dynamically updating C ref during the call, and continuously optimizing C ref according to the embedding mean of proximal audio chunks with high confidence, the present application ensures that the model maintains a high determination accuracy during long calls. The present application introduces an attention mechanism, dynamically adjusts the weights of P vad , P near , and S according to the feature changes of the audio chunk, makes the feature weighting more flexible, and thus enhances the target speaker determination ability in complex scenarios.

[0041] It should be understood that the above general description and the following detailed description are only exemplary and explanatory and should not limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.

[0043] Figure 1 is a flowchart of a speaker voice recognition method provided by an embodiment of the present application;

[0044] Figure 2 is a flowchart of a training method of a speaker voice recognition model provided by an embodiment of the present application;

[0045] Figure 3 This is a schematic structural diagram of a computer device provided by an embodiment of the present application. Specific embodiments

[0046] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0047] Next, a speaker voice recognition method provided by an embodiment of the present application will be described in detail in conjunction with specific embodiments. Figure 1 This is a flowchart of a speaker voice recognition method provided by an embodiment of the present application, as Figure 1 shown, and the specific steps are as follows:

[0048] Step S101: Real-time obtain the call audio to be recognized, and segment the call audio to be recognized into a plurality of audio blocks.

[0049] In implementation, the computer can obtain the call audio to be recognized in real time during the user's call. The call audio to be recognized is input in the form of continuous time-series data. For more efficient processing, the call audio to be recognized can be segmented according to a fixed length (such as 0.5 seconds to 2 seconds). The segmented audio blocks can retain the time-series context information to ensure that key voice features are not lost during feature extraction.

[0050] Step S102: For each audio block, extract the voice features of the audio block and input them into a pre-trained speaker voice recognition model.

[0051] In implementation, the computer can extract the key features of the audio block, including low-level acoustic features, high-level voice features, distance-related features, etc., for the speaker voice recognition model to make subsequent judgments.

[0052] Step S103: Output the belonging probability that the audio block belongs to the target speaker through the speaker voice recognition model.

[0053] In implementation, the computer can output the belonging probability that the audio block belongs to the target speaker through the speaker voice recognition model.

[0054] Step S104: If the belonging probability that the audio block belongs to the target speaker is greater than a preset speaker belonging probability threshold, it is determined that the audio block belongs to the target speaker.

[0055] In implementation, the computer can determine according to the belonging probability (P targetPerform threshold judgment to determine whether the current audio block belongs to the target speaker. For example, the preset speaker attribution probability threshold (such as 0.75) can be adjusted according to the ROC curve of the model or its performance on historical data. When P target is higher than the speaker attribution probability threshold, it is determined that the current audio block belongs to the target speaker.

[0056] As an alternative implementation, Figure 2 is a flowchart of a training method for a speaker voice recognition model provided by an embodiment of the present application. As Figure 2 shown, the specific steps are as follows:

[0057] Step S201: Obtain historical call audio and split the historical call audio into several audio blocks.

[0058] In implementation, a technician can decompose the historical call audio into multiple audio blocks (chunks) of a fixed length by a computer to provide data for subsequent feature extraction and model training. The splitting method of the audio blocks can adopt strategies such as fixed-length splitting and overlapping window sliding splitting. The splitting length can be set between 0.5 seconds and 2 seconds to ensure that each block contains sufficient speech features without losing time information.

[0059] Step S202: Perform training annotation on each audio block.

[0060] In implementation, a training label can be added to each audio block.

[0061] As an alternative implementation, the specific method for performing training annotation on each audio block in step S202 is:

[0062] Perform voice state annotation on each audio block. The voice state includes a voiced state and an unvoiced state. Based on a pre-trained NEAR pre-model, proximal annotation is performed on each audio block according to the sound energy, reverberation degree, and high-frequency retention of the audio block. Perform target speaker annotation on each audio block.

[0063] In implementation, a VAD (Voice Activity Detection) model can be used to determine whether there is human voice in each audio block, providing labels for subsequent multi-task model training. Specifically, the time-domain and frequency-domain features of the audio block (such as energy, zero-crossing rate, spectral features, etc.) can be used to determine whether the audio block contains human voice. Determine whether the audio block is from the proximal end of the device, as the NEAR annotation (for example: 1 represents the proximal end, 0 represents the distal end), which is used to improve the robustness of target speaker recognition. Through a pre-trained NEAR model, the proximal or distal end can be determined based on the following features: the proximal block has higher sound energy, the distal block is accompanied by stronger reverberation, the high-frequency part of the proximal block is better preserved, while the high-frequency information of the distal block decays more. The speakers in the audio block can be clustered or authenticated through speaker clustering (SV clustering) or a speaker verification model (SV model). For example, unsupervised clustering techniques such as DBSCAN and K-means can be used, or a speaker verification (SpeakerVerification) model can be used for target speaker annotation. The similarity between the target speaker reference vector C ref and the embedding vector of the current block can be compared to determine whether the audio block belongs to the target speaker. The NEAR pre-model is an auxiliary model trained before the main speaker speech recognition model, which is used to determine whether the audio block is proximal speech. This model is trained using supervised learning. The input is the energy distribution, high-frequency energy ratio, reverberation characteristics, etc. of the audio block, and the output is the probability value that the current audio block belongs to proximal (device voice) or distal (the other party or background) speech. In the training data, the proximal / distal labels can be obtained through device-side recording and scene annotation. The NEAR pre-model can adopt a lightweight CNN network structure to ensure real-time performance. After training, the model weights are frozen and used as an auxiliary annotator in the subsequent main model training. In the training stage of the main speaker speech recognition model, the NEAR pre-model is used to perform prior annotation on each audio block, generating the target labels for the proximal speech judgment task in the training set. At the same time, in the model inference stage, the function of the NEAR pre-model is integrated into the "proximal speech judgment head" of the main model to achieve end-to-end processing. The use of this pre-model ensures high confidence in the annotation of proximal speech data, providing a basis for the accurate construction of the target speaker reference vector C ref , thereby improving the judgment accuracy and stability of the entire recognition system.

[0064] Step S203: For each audio block, extract the speech features of this audio block.

[0065] In implementation, the computer can extract multi-dimensional features from each audio block, such as acoustic features (e.g., MFCC, Fbank), distance judgment features (e.g., high-frequency energy ratio, reverberation features), and high-level features (e.g., high-level features extracted by pre-trained models such as Wav2Vec, HuBERT). Specifically, time-frequency features can be extracted through a convolutional neural network (CNN) or a Transformer model to generate embedding vectors, and then the speech features will be passed as input to the speaker speech recognition model for joint learning.

[0066] Step S204: Construct a speaker speech recognition model, which includes an encoder module and a multi-task output module.

[0067] In implementation, a speaker speech recognition model is constructed to simultaneously complete speech activity detection, proximal judgment, speaker feature extraction, and target speaker judgment. The speaker speech recognition model includes an encoder module and a multi-task output module.

[0068] As an optional implementation manner, in step S204, the encoder module includes a convolutional neural network layer and a Transformer structure; where

[0069] The convolutional neural network layer is used to refine and fuse speech features, capture time-frequency patterns, vocal features, and reverberation information;

[0070] The Transformer structure is used to maintain context information using the hidden states of previous audio blocks.

[0071] In implementation, the audio block can be feature-extracted through a convolutional neural network (CNN) structure to capture local time-frequency information, vocal patterns, and reverberation characteristics, generating a more discriminative feature representation. The convolutional kernel slides in the frequency domain and the time domain, capable of capturing local time-frequency patterns in speech features. Through multiple layers of convolution and pooling operations, CNN extracts multi-dimensional features, including time-frequency features: key information for identifying the speech spectrum; vocal features: differentiating the pronunciation patterns of different speakers; reverberation information: identifying the reverberation features generated when speaking proximally or distally. The multi-channel features are processed through pooling and activation functions and then feature-fused to form a time-frequency representation of the audio block. The context information of previous audio blocks is maintained through the Transformer structure, capturing cross-block temporal dependencies and providing a global context for target speaker judgment. The Transformer captures the dependencies between speech blocks in the time dimension by encoding the hidden states of previous blocks. At the same time, state information is shared between each block, thereby achieving a comprehensive understanding of the call context during target speaker judgment.

[0072] As an alternative implementation, in step S204, the multi-task output module includes a voice activity detection head, a near-end voice judgment head, a speaker embedding extraction head, and a target speaker judgment head; where,

[0073] The voice activity detection head is used to output the voice activity probability of the audio block;

[0074] The near-end voice judgment head is used to output the near-end probability of the audio block;

[0075] The speaker embedding extraction head is used to output the speaker embedding vector of the audio block;

[0076] The target speaker judgment head is used to comprehensively consider the voice activity probability, the near-end probability, the speaker embedding vector, and a preset target speaker reference vector, and output the belonging probability that the audio block belongs to the target speaker.

[0077] In implementation, the voice activity detection head (VAD head) can be used to determine whether the current audio block is a "voiced block" or a "silent block", and output the voice activity probability P that the block is voiced vad . Specifically, a CNN or RNN model can be used to determine the voice activity state of the audio block based on the time-domain and frequency-domain features of the audio block (such as energy, zero-crossing rate, MFCC, etc.). P is output through a binary classification model (voiced / silent) vad , representing the probability that the audio block is a "voiced block". The output of the voice activity detection head is a probability value within the range of [0, 1]. For example, when P vad > 0.5, it can be determined as a "voiced block", otherwise it is a "silent block". The near-end voice judgment head (NEAR head) can be used to determine whether the current audio block is from the "near end" or "far end" of the device, and output the probability P that the block is a near-end block near . Specifically, a CNN model can be used to perform binary classification (near end / far end) on the audio block by combining features such as sound energy, high-frequency retention, and reverberation degree, and output P near . For example, a block with a high-frequency energy ratio > 0.6 and a reverberation fraction < 0.3 is more likely to be a near-end block. For example, the near-end voice judgment head outputs a probability value within the range of [0, 1]. When P near > 0.5, it is considered that the block is from the near end. When P nearLess than 0.5 comes from the far end. The speaker embedding vector e in the current audio block can be extracted through the Speaker Verification (SV) head for speaker identity discrimination. Specifically, a CNN + Transformer model can be used for feature extraction to convert the time-frequency features into a speaker embedding vector of a fixed dimension (usually 256 or 512 dimensions). A pre-trained Speaker Verification (SV) model is used for speaker embedding extraction to ensure that the embedding vector has a high degree of discrimination. e is used as an important feature input for target speaker judgment and compared with the target speaker reference vector (C ref ) for similarity comparison. Through the Target Speaker Judgment (TARGET) head, comprehensive judgment is made by combining multi-dimensional information, and the probability P that the current audio block belongs to the target speaker is output target .

[0078] As an alternative implementation, the preset target speaker reference vector is the global embedding mean of the speaker embedding vectors in the identified near-end audio blocks during the call

[0079] In implementation, the embedding vectors of multiple near-end audio blocks during the call can be extracted and aggregated to generate a target speaker reference vector C ref , as the standard reference for subsequent target speaker judgment. Among them, during the call, the Near-End Voice Judgment (NEAR) head is used to judge each audio block as near end / far end. Only the audio blocks judged as "near end" (P near > 0.5) will be used to generate the target speaker reference vector. The speaker embedding vector e is extracted from the near-end audio block through the Speaker Verification (SV) head. Each near-end audio block can generate an embedding vector, and these embedding vectors together reflect the characteristic information of the target speaker. By performing weighted average or simple average on all the identified near-end audio block embedding vectors, a global embedding mean is obtained as the target speaker reference vector C ref . As more near-end audio blocks are identified, the model can continuously introduce new embedding vectors and use a sliding window or time-weighted method to dynamically update C ref ; this update mechanism can ensure that C ref is adaptively adjusted when the characteristics of the target speaker change, thereby improving the robustness of recognition. For example: in the first few seconds of the call, the model extracts embedding vectors from high-confidence near-end audio blocks, and these embedding vectors form the initial C ref after mean calculation and are used as the reference vector for subsequent blocks. Whenever a newly identified near-end audio block is determined to belong to the target speaker, its embedding vector is incorporated into C ref calculation, and a sliding window mean or time-weighted average method can be used to calculate C refPerform dynamic updates to adapt to changes in the characteristics of the target speaker.

[0080] As an alternative implementation, mean filtering or median filtering can also be used to denoise the embedding vectors, excluding abnormal embeddings or noise interference.

[0081] As an alternative implementation, the target speaker judgment head is also used for:

[0082] Calculate the similarity score between the speaker embedding vector of the audio block to be recognized and the target speaker reference vector;

[0083] Based on the multi-dimensional feature weighted fusion mechanism, perform weighted calculations on the voice activity probability, proximal probability, and similarity score to determine the attribution probability that the audio block belongs to the target speaker.

[0084] In implementation, the matching degree between the current audio block and the target speaker characteristics can be quantified by calculating the similarity score between the speaker embedding vector e of the current audio block and the target speaker reference vector C ref . e is a feature vector with a fixed dimension (usually 256-dimensional or 512-dimensional), representing the speaker characteristics of the current block. The similarity score S can be calculated using cosine similarity (CosineSimilarity) or dot product similarity. For example, the value range of S is [-1, 1]. The higher the similarity, the higher the matching degree between the current block and the target speaker characteristics. Through the multi-dimensional feature weighted fusion mechanism, feature information from different sources is integrated to improve the accuracy of target speaker judgment, and the attribution probability P that the audio block belongs to the target speaker is output target . The target speaker judgment head can perform weighted summation of P vad , P near and S to generate the final P target . Among them, the weights can be automatically optimized according to the performance during model training, or can be set by experience. For example, when the call background noise is large, the weight of P near can be increased.

[0085] As an alternative implementation, an attention mechanism is introduced into the target speaker judgment head. By calculating the attention weight between the speaker embedding vector of the audio block and the target speaker reference vector, feature weighted fusion of the target judgment is performed to improve the accuracy of audio block attribution judgment.

[0086] In implementation, the attention mechanism (AttentionMechanism) can be used to calculate the speaker embedding vector e of the current audio block and the target speaker reference vector C refThe weight relationship between them is used to dynamically adjust the importance of features, thereby optimizing the judgment result of the target speaker. Among them, the attention mechanism is a weighted feature aggregation mechanism. By calculating the correlation (similarity) between different features, higher weights are assigned to important features, thereby enhancing the model's attention to key information. The weights of P vad , P near and S are dynamically adjusted through the attention mechanism to make feature fusion more targeted, thereby improving the accuracy of the target speaker attribution probability P target output by the target judgment head. Through the attention mechanism, according to the matching degree of e(chunk) and C ref , the weights of different features are adjusted. When S is higher, the attention mechanism can assign it a higher weight, making it have a greater impact on the judgment of the target speaker; in a scene with high noise, the weight of P near is relatively increased to reduce the interference of environmental noise. For example: at the beginning of a call, the reference vector C ref of the target speaker is not yet fully stable, and the weights of P vad and P near are relatively high. In the middle of the call, the reference vector C ref of the target speaker gradually converges, and the matching degree of S increases. The attention mechanism automatically increases the weight of S, making the judgment more dependent on the matching degree of the speaker embedding features. At the end of the call, as the background noise increases, the reliability of P near and P vad relatively increases.

[0087] Step S205, use a multi-task loss function to jointly optimize the speaker speech recognition model.

[0088] In implementation, the model can be jointly optimized through a multi-task loss function to make it have good generalization ability in multi-tasks. Specifically, in the actual training process, the model can adopt the method of joint multi-task learning for end-to-end training. All sub-tasks (including voice activity detection, proximal judgment, speaker embedding extraction, target speaker judgment) share the preposed encoder module and output prediction results through different task heads. The multi-task loss function can be composed of the following parts:

[0089] Voice activity detection loss function: Use cross-entropy loss to optimize the model's judgment of whether an audio chunk is a voiced chunk;

[0090] Proximal judgment loss function: Use cross-entropy loss to judge whether an audio chunk comes from the proximal end;

[0091] Embedding extraction loss function: Use AM-Softmax (Additive Margin Softmax) or Triplet Loss to improve the discriminability of the embedding vector;

[0092] Target speaker judgment loss function: The cross-entropy loss is adopted to optimize the judgment result of whether an audio block belongs to the target speaker;

[0093] The final total loss function L of the model total can be expressed as the weighted sum of multiple sub-task loss functions:

[0094] L total =αL vad +βL near +γL sv +σL target ;

[0095] Among them, α, β, γ, and σ are the weight coefficients of each sub-task loss, which can be tuned by cross-validation or dynamically adjusted using an indefinite proportion learning strategy (such as GradNorm).

[0096] During the model training process, all audio blocks are simultaneously annotated and predicted for four tasks. Through this multi-task joint optimization strategy, not only the performance of a single task is improved, but also the information sharing and generalization ability between tasks are enhanced, effectively improving the recognition accuracy and robustness of the target speaker.

[0097] As an alternative implementation, the loss function in step S205 includes a voice activity detection loss function, a proximal judgment loss function, an embedding extraction loss function, and a target speaker judgment loss function; among them,

[0098] The voice activity detection loss function is used to distinguish between voiced and unvoiced audio blocks;

[0099] The proximal judgment loss function is used to distinguish between proximal and distal audio blocks;

[0100] The embedding extraction loss function is used to optimize the embedding discrimination;

[0101] The target speaker judgment loss function maximizes the attribution probability of the audio blocks of the target speaker and minimizes the misjudgment probability of the audio blocks of other speakers at the same time.

[0102] In implementation, the voice activity detection loss function (L vad ) can be used to optimize the model to distinguish between the "voiced state" and "unvoiced state" of audio blocks, improving the accuracy of voice activity detection. Specifically, the voice activity detection head judges whether the audio block contains voice activity and outputs P vad . The goal is to label voiced blocks as 1 and unvoiced blocks as 0, forming a binary classification task. Maximize the probability of P vad in voiced blocks and minimize the misjudgment probability of unvoiced blocks at the same time to improve the accuracy of voice activity detection. Through the proximal judgment loss function (L near), optimize the model to distinguish between the "proximal block" and "distal block" of the audio block, and improve the accuracy of proximal judgment. The proximal speech judgment head outputs P near , to judge whether the audio block is "proximal" or "distal" from the device. The goal is to correctly distinguish the proximal block (labeled as 1) and the distal block (labeled as 0). By embedding and extracting the loss function (L sv ), optimize the discriminability of the speaker embedding vector, and improve the ability to separate the features of the target speaker from other speakers. Extract the embedding vector e of the audio block through the speaker embedding extraction head, which is used to compare with the target speaker reference vector C ref . The speaker embedding extraction task can adopt the AM-Softmax (Additive Margin Softmax) loss function, where the scaling factor is used to control the distribution range, and the angular margin is used to increase the discriminability between the target and non-target speakers. Through the target speaker judgment loss function (L target ), optimize the model's judgment on the attribution of the target speaker block, maximize the correct attribution probability of the target block, and minimize the misjudgment probability of other speaker blocks at the same time. The target speaker judgment head outputs the probability P target that the current audio block belongs to the target speaker through multi-dimensional feature weighted fusion. The target speaker judgment belongs to a binary classification task, and the cross-entropy loss function can be used for optimization. After the true label of the target speaker (1 represents the target block, 0 represents the non-target block) is optimized, the attribution probability of the target speaker block will be further improved, while reducing the interference of other speaker blocks and enhancing the overall robustness of the system.

[0103] An embodiment of the present application provides a speaker speech recognition method, the method includes: obtaining the call audio to be recognized in real time, and splitting the call audio to be recognized into several audio blocks; for each audio block, extracting the speech features of the audio block and inputting them into a pre-trained speaker speech recognition model; outputting the attribution probability that the audio block belongs to the target speaker through the speaker speech recognition model; if the attribution probability that the audio block belongs to the target speaker is greater than the preset speaker attribution probability threshold, it is determined that the audio block belongs to the target speaker. The technical solution provided by the embodiment of the present application at least brings the following beneficial effects: In the traditional technology, the judgment of the target speaker usually relies on a single feature and cannot make full use of multi-dimensional information, resulting in low recognition accuracy. In this application, by introducing a multi-task model, P vad (voice activity), P near (proximal judgment), and S (embedding similarity) are weighted and fused for features, realizing the adaptive adjustment of different features at different call stages, and improving the accuracy of target speaker judgment. The traditional method uses a fixed reference vector for matching and cannot adaptively adjust according to the feature changes of the target speaker during the call, resulting in a decrease in recognition accuracy during long calls. In this application, by calculating the reference vector C of the target speakerref , and during the call, ref Dynamically update and continuously optimize C according to the embedding mean of high-confidence near-end audio blocks ref , thereby ensuring that the model maintains a high judgment accuracy during long calls. This application introduces an attention mechanism to monitor P according to the feature changes of the audio block. vad , P near Dynamic weight adjustment is performed on S to make feature weighting more flexible, thereby enhancing the target speaker judgment ability in complex scenarios.

[0104] It should be understood that although Figure 1 and Figure 2 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 and Figure 2 At least part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.

[0105] It can be understood that the same / similar parts between the various embodiments of the above method in this specification can refer to each other, and each embodiment focuses on the differences from other embodiments. For related points, please refer to the description of other method embodiments.

[0106] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program that can be run on the processor. When the processor executes the computer program, the method steps of dynamically identifying flight changes are implemented. Figure 3 A schematic diagram of the structure of a computer device provided in an embodiment of the present application is shown in FIG. Figure 3As shown in the figure, the computer device may include a processor 301, a system bus 302, a non-volatile storage medium 303, an internal memory 304, a network interface 305, a display screen 306, and an input device 307. Among them, the non-volatile storage medium 303 stores an operating system 3031 and a computer program 3032. The processor 301 is used to execute the computer program 3032 to implement the method steps of speaker speech recognition. The system bus 302 is used to connect the processor 301, the non-volatile storage medium 303, the internal memory 304, the network interface 305, the display screen 306, and the input device 307 to ensure efficient communication between the components. The internal memory 304 is used to temporarily store the programs and data being run, helping the processor 301 quickly access the required information, thereby improving the overall system performance. The network interface 305 (such as a network card) enables the computer device to connect to a local area network or the Internet to achieve data transmission and remote communication. The display screen 306 is used to present the information of speaker speech recognition processed by the computer device to the user in the form of graphics or text. The input device 307 (such as a keyboard, mouse, touch screen, etc.) is used to allow the user to input voice data and commands to the computer device to achieve interactive operations with the computer device.

[0107] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application may include non-volatile and / or volatile memories. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0108] It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.

[0109] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0110] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the corresponding part of the method embodiment for the relevant content.

[0111] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0112] The above-described embodiments only represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application patent should be subject to the appended claims.

Claims

1. A speaker voice recognition method, characterized in that, The method includes: Obtaining the call audio to be recognized in real time, and splitting the call audio to be recognized into a plurality of audio chunks; For each of the audio chunks, extracting the speech features of the audio chunk and inputting them into a pre-trained speaker speech recognition model; Outputting, by the speaker speech recognition model, the belonging probability that the audio chunk belongs to the target speaker; If the belonging probability that the audio chunk belongs to the target speaker is greater than a preset speaker belonging probability threshold, determining that the audio chunk belongs to the target speaker; The training process of the speaker speech recognition model includes: Obtaining historical call audio, and splitting the historical call audio into a plurality of audio chunks; Performing training annotation on each of the audio chunks; For each of the audio chunks, extracting the speech features of the audio chunk; Constructing the speaker speech recognition model, where the speaker speech recognition model includes an encoder module and a multi-task output module; Jointly optimizing the speaker speech recognition model using a multi-task loss function; The performing training annotation on each of the audio chunks includes: Performing vocal state annotation on each of the audio chunks, where the vocal state includes a vocal state and a silent state; Based on a pre-trained NEAR pre-model, performing proximal annotation on each of the audio chunks according to the sound energy, reverberation degree, and high-frequency retention condition of the audio chunk; Performing target speaker annotation on each of the audio chunks.

2. The method according to claim 1, wherein The speech features include acoustic features, distance-related features, and high-level speech features.

3. The method according to claim 1, wherein The encoder module includes a convolutional neural network layer and a Transformer structure; where The convolutional neural network layer is used to refine and fuse the speech features, capture time-frequency patterns, vocalization features, and reverberation information; The Transformer structure is used to maintain context information using the hidden states of previous audio chunks.

4. The method according to claim 1, wherein The multi-task output module includes a voice activity detection head, a proximal voice judgment head, a speaker embedding extraction head, and a target speaker judgment head; where The voice activity detection head is used to output the voice activity probability of the audio chunk; The proximal voice judgment head is used to output the proximal probability of the audio chunk; The speaker embedding extraction head is used to output the speaker embedding vector of the audio chunk; The target speaker judgment head is used to comprehensively output the belonging probability that the audio chunk belongs to the target speaker based on the voice activity probability, the proximal probability, the speaker embedding vector, and a preset target speaker reference vector; 5. The method according to claim 4, characterized in that, The preset target speaker reference vector is the global embedding mean of the speaker embedding vectors in the proximal audio chunks that have been recognized during the call.

6. The method according to claim 4, characterized in that, The target speaker judgment head is further used to: Calculate the similarity score between the speaker embedding vector of the audio chunk to be recognized and the target speaker reference vector; Based on a multi-dimensional feature weighted fusion mechanism, perform weighted calculation on the voice activity probability, the proximal probability, and the similarity score to determine the belonging probability that the audio chunk belongs to the target speaker.

7. The method according to claim 4, characterized in that The method further includes: An attention mechanism is introduced into the target speaker judgment head. By calculating the attention weights between the speaker embedding vector of the audio block and the target speaker reference vector, feature weighted fusion of the target judgment is performed to improve the accuracy of audio block attribution judgment.

8. The method according to claim 1, wherein The loss function includes a voice activity detection loss function, a proximal judgment loss function, an embedding extraction loss function, and a target speaker judgment loss function; among them, the voice activity detection loss function is used to distinguish whether the audio block is vocal or silent; the proximal judgment loss function is used to distinguish whether the audio block is proximal or distal; the embedding extraction loss function is used to optimize the embedding discrimination; the target speaker judgment loss function maximizes the attribution probability of the audio block of the target speaker and minimizes the misjudgment probability of the audio blocks of other speakers at the same time.

Citation Information

Patent Citations

  • Junk call identification method and device, computer equipment and storage medium

    CN119629636A