Speaker speech recognition method

By acquiring and processing call audio in real time, using a pre-trained speaker speech recognition model and multi-task output module, combining attention mechanism and dynamically updated target speaker reference vectors, the problem of target speaker judgment under noise and interference in the prior art is solved, and higher recognition accuracy and robustness are achieved.

CN119993154AActive Publication Date: 2025-05-13SHENZHEN HUOLI TIAN HUI TECH CO LTD

Patent Information

Application Number
CN202510480115.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-13
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify when judging the target speaker, especially in the case of noisy noise or other vocal interference.

Method used

By obtaining the call audio in real time, dividing it into audio blocks, voice features are extracted and input into the pre-trained speaker speech recognition model, and outputting the audio block belongs to the target speaker's attribution probability. The model includes an encoder module and a multi-task output module, which uses a multi-task loss function for joint optimization, and introduces an attention mechanism and dynamically updated target speaker reference vector.

Benefits of technology

The accuracy of the target speaker's judgment is improved, especially in long-term calls and complex scenarios, and the robustness and recognition accuracy of the model are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993154A_ABST
    Figure CN119993154A_ABST
Patent Text Reader

Abstract

The invention relates to a speaker voice recognition method, and relates to the technical field of telephone voice signal processing. The method comprises the following steps: acquiring a to-be-identified call audio in real time, and segmenting the to-be-identified call audio into a plurality of audio blocks; for each audio block, extracting voice features of the audio block, and inputting the voice features into a pre-trained speaker voice recognition model; outputting the affiliation probability that the audio block belongs to the target speaker through the speaker speech recognition model; and if the affiliation probability of the audio block belonging to the target speaker is greater than a preset speaker affiliation probability threshold, judging that the audio block belongs to the target speaker. According to the invention, target speaker identification in a complex scene can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of telephone voice signal processing, and in particular to a method for speaker voice recognition. Background Art

[0002] At present, with the development of technologies such as speech recognition, intelligent assistants, and remote communications, the demand for target speaker identification is increasing. Existing technologies usually rely on voice activity detection (VAD) to determine whether an audio block is audible, and combine it with speaker verification (SV) for identity recognition. However, when there is noisy surroundings, it is more difficult to judge. Existing human voice models have already performed well in filtering environmental noise, but if there is other human voice noise around, it will interfere with the judgment of the target voice. Summary of the invention

[0003] Based on this, it is necessary to provide a speaker speech recognition method for the above technical problems, the method comprising: Acquire the call audio to be identified in real time, and divide the call audio to be identified into a plurality of audio blocks; For each of the audio blocks, extract the speech features of the audio block and input them into a pre-trained speaker speech recognition model; Outputting the probability that the audio block belongs to the target speaker through the speaker speech recognition model; If the attribution probability of the audio block belonging to the target speaker is greater than a preset speaker attribution probability threshold, it is determined that the audio block belongs to the target speaker.

[0004] As an optional implementation, the training process of the speaker speech recognition model includes: Obtaining historical call audio, and dividing the historical call audio into a plurality of audio blocks; Performing training annotation on each of the audio blocks; For each of the audio blocks, extracting speech features of the audio block; Constructing the speaker speech recognition model, wherein the speaker speech recognition model includes an encoder module and a multi-task output module; The speaker speech recognition model is jointly optimized using a multi-task loss function.

[0005] As an optional implementation manner, the performing training annotation on each of the audio blocks includes: Marking the vocal state of each of the audio blocks, where the vocal state includes a voiced state and a silent state; Based on the pre-trained NEAR pre-model, each of the audio blocks is proximally labeled according to the sound energy, reverberation level, and high-frequency retention of the audio block; A target speaker is labeled for each of the audio blocks.

[0006] As an optional implementation, the speech features include acoustic features, distance-related features and high-level speech features.

[0007] As an optional implementation, the encoder module includes a convolutional neural network layer and a Transformer structure; wherein, The convolutional neural network layer is used to refine and fuse the speech features to capture time-frequency patterns, vocalization characteristics and reverberation information; The Transformer structure is used to maintain context information using the hidden state of the previous audio block.

[0008] As an optional implementation, the multi-task output module includes a voice activity detection head, a near-end voice judgment head, a speaker embedding extraction head and a target speaker judgment head; wherein, The voice activity detection head is used to output the voice activity probability of the audio block; The near-end speech judgment head is used to output the near-end probability of the audio block; The speaker embedding extraction head is used to output the speaker embedding vector of the audio block; The target speaker judgment head is used to integrate the voice activity probability, the proximal probability, the speaker embedding vector and a preset target speaker reference vector to output the attribution probability that the audio block belongs to the target speaker.

[0009] As an optional implementation, the preset target speaker reference vector is a global embedding mean of speaker embedding vectors in near-end audio blocks that have been identified during a call.

[0010] As an optional implementation manner, the target speaker determination head is further used for: Calculate the similarity score between the speaker embedding vector of the audio block to be identified and the target speaker reference vector; Based on a multi-dimensional feature weighted fusion mechanism, the voice activity probability, the proximal probability and the similarity score are weightedly calculated to determine the attribution probability that the audio block belongs to the target speaker.

[0011] As an optional implementation, the method further includes: An attention mechanism is introduced into the target speaker judgment head. By calculating the attention weight between the speaker embedding vector of the audio block and the target speaker reference vector, feature weighted fusion is performed on the target judgment to improve the accuracy of the audio block attribution judgment.

[0012] As an optional implementation, the loss function includes a speech activity detection loss function, a proximal judgment loss function, an embedding extraction loss function and a target speaker judgment loss function; wherein, The voice activity detection loss function is used to distinguish whether the audio block is voiced or silent; The near-end judgment loss function is used to distinguish whether the audio block is near-end or far-end; The embedding extraction loss function is used to optimize the embedding discrimination; The target speaker judgment loss function maximizes the attribution probability of the target speaker's audio block and minimizes the misjudgment probability of other speakers' audio blocks.

[0013] In a second aspect, a computer device is provided, comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, and when the processor executes the computer program, the method steps described in any one of the first aspects are implemented.

[0014] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method steps described in any one of the first aspects are implemented.

[0015] The present application provides a method for speaker speech recognition, the method comprising: acquiring call audio to be recognized in real time, and dividing the call audio to be recognized into a number of audio blocks; for each of the audio blocks, extracting the speech features of the audio block, and inputting them into a pre-trained speaker speech recognition model; outputting the probability of the audio block belonging to the target speaker through the speaker speech recognition model; if the probability of the audio block belonging to the target speaker is greater than a preset speaker attribution probability threshold, then determining that the audio block belongs to the target speaker. The technical solution provided by the embodiments of the present application brings at least the following beneficial effects: In traditional technologies, the judgment of the target speaker usually relies on a single feature and cannot make full use of multi-dimensional information, resulting in low recognition accuracy. The present application introduces a multi-task model to convert P vad (Voice Activity), P near (proximal judgment) and S (embedded similarity) for feature weighted fusion, which realizes the adaptive adjustment of different features at different stages of the call and improves the accuracy of the target speaker judgment. The traditional method uses a fixed reference vector for matching, which cannot be adaptively adjusted according to the changes in the characteristics of the target speaker during the call, resulting in a decrease in recognition accuracy during long calls. This application calculates the reference vector C of the target speaker. ref , and during the call, ref Dynamically update and continuously optimize C according to the embedding mean of high-confidence near-end audio blocks ref, thereby ensuring that the model maintains a high judgment accuracy during long calls. This application introduces an attention mechanism to monitor P according to the feature changes of the audio block. vad , P near Dynamic weight adjustment is performed on S to make feature weighting more flexible, thereby enhancing the target speaker judgment ability in complex scenarios.

[0016] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 A flowchart of a method for speaker speech recognition provided in an embodiment of the present application; Figure 2 A flowchart of a method for training a speaker speech recognition model provided in an embodiment of the present application; Figure 3 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0020] The following will describe in detail a method for speaker speech recognition provided by an embodiment of the present application in conjunction with a specific implementation method. Figure 1 A flowchart of a method for speaker speech recognition provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the specific steps are as follows: Step S101, obtaining the call audio to be identified in real time, and dividing the call audio to be identified into a plurality of audio blocks.

[0021] In implementation, the computer can obtain the audio of the call to be recognized in real time during the user's call. The audio of the call to be recognized is input in the form of continuous time series data. For more efficient processing, the audio of the call to be recognized can be segmented into fixed lengths (such as 0.5 seconds to 2 seconds). The segmented audio blocks can retain the time series context information to ensure that key voice features are not lost during feature extraction.

[0022] Step S102: for each audio block, extract the speech features of the audio block and input them into a pre-trained speaker speech recognition model.

[0023] During implementation, the computer can extract key features of the audio block, including low-level acoustic features, high-level speech features, distance-related features, etc., for subsequent judgment by the speaker speech recognition model.

[0024] Step S103: outputting the probability that the audio block belongs to the target speaker through the speaker speech recognition model.

[0025] During implementation, the computer may output the attribution probability that the audio block belongs to the target speaker through a speaker speech recognition model.

[0026] Step S104: if the attribution probability of the audio block belonging to the target speaker is greater than a preset speaker attribution probability threshold, it is determined that the audio block belongs to the target speaker.

[0027] In implementation, the computer can calculate the probability of the target speaker belonging to the target speaker (P target ) to determine whether the current audio block belongs to the target speaker. For example, the preset speaker attribution probability threshold (such as 0.75) can be adjusted according to the ROC curve of the model or the performance on historical data. target When the probability of speaker belonging is higher than the threshold, the current audio block is determined to belong to the target speaker.

[0028] As an optional implementation, Figure 2 A flowchart of a method for training a speaker speech recognition model provided in an embodiment of the present application, such as Figure 2 As shown, the specific steps are as follows: Step S201, obtain historical call audio, and divide the historical call audio into several audio blocks.

[0029] In practice, technicians can use computers to decompose historical call audio into multiple fixed-length audio chunks (chunks) to provide data for subsequent feature extraction and model training. The audio chunks can be segmented using fixed-length segmentation, overlapping window sliding segmentation, and other strategies. The segmentation length can be set between 0.5 seconds and 2 seconds to ensure that each chunk contains sufficient voice features without losing time information.

[0030] Step S202: Perform training annotation on each audio block.

[0031] In implementation, training labels may be added for each audio block.

[0032] As an optional implementation, the specific method of training and labeling each audio block in step S202 is: Each audio block is labeled with the vocal state, which includes the vocal state and the silent state. Based on the pre-trained NEAR pre-model, each audio block is labeled near-end according to the sound energy, reverberation level, and high-frequency retention of the audio block. Each audio block is labeled with the target speaker.

[0033] In implementation, the VAD (Voice Activity Detection) model can be used to determine whether there is human voice in each audio block, providing labels for subsequent multi-task model training. Specifically, the time domain and frequency domain features of the audio block (such as energy, zero crossing rate, spectral features, etc.) can be used to determine whether the audio block contains human voice. Determine whether the audio block comes from the near end of the device as a NEAR annotation (for example: 1 for near end, 0 for far end) to improve the robustness of target speaker recognition. Through the pre-trained NEAR model, the near end or far end can be judged based on the following features: the near end block has higher sound energy, the far end block is accompanied by stronger reverberation, the high frequency part of the near end block is better preserved, and the high frequency information of the far end block is more attenuated. The speakers in the audio block can be clustered or authenticated through speaker clustering (SV clustering) or speaker verification model (SV model). For example, unsupervised clustering techniques such as DBSCAN and K-means can be used, or the speaker verification (SpeakerVerification) model can be used to annotate the target speaker. The target speaker reference vector (C ref ) is compared with the embedding vector of the current block to determine whether the audio block belongs to the target speaker. The NEAR pre-model is an auxiliary model that is trained independently of the main speaker speech recognition model and is used to determine whether the audio block is near-end speech. The model is trained using supervised learning. The input is the energy distribution, high-frequency energy ratio, reverberation characteristics, etc. of the audio block, and the output is the probability value of the current audio block belonging to the near-end (device sound) or far-end (other party or background) speech. In the training data, the near-end / far-end labels can be obtained through device recording and scene annotation. The NEAR pre-model can use a lightweight CNN network structure to ensure real-time performance. After the training is completed, the model weights are frozen and used as an auxiliary annotator in the subsequent main model training. In the main speaker speech recognition model training stage, the NEAR pre-model is used to perform a priori annotation on each audio block and generate the target label for the near-end speech judgment task in the training set. At the same time, in the model inference stage, the function of the NEAR pre-model is integrated into the "near-end speech judgment head" of the main model to achieve end-to-end processing. The use of this pre-model ensures high confidence in the annotation of the near-end speech data, providing a reference vector C for the target speaker. ref It provides a foundation for the accurate construction of the recognition system, thereby improving the judgment accuracy and stability of the entire recognition system.

[0034] Step S203: extracting speech features of each audio block.

[0035] In practice, the computer can extract multi-dimensional features from each audio block, such as acoustic features (such as MFC and Fbank), distance judgment features (such as high-frequency energy ratio and reverberation features), and high-level features (such as high-level features extracted by pre-trained models such as Wav2Vec and HuBERT). Specifically, the time-frequency features can be extracted and embedded vectors can be generated through a convolutional neural network (CNN) or a Transformer model, and then the speech features will be passed as input to the speaker speech recognition model for joint learning.

[0036] Step S204: construct a speaker speech recognition model, which includes an encoder module and a multi-task output module.

[0037] In the implementation, a speaker speech recognition model is constructed to simultaneously complete voice activity detection, proximal judgment, speaker feature extraction, and target speaker judgment. The speaker speech recognition model includes an encoder module and a multi-task output module.

[0038] As an optional implementation, the encoder module in step S204 includes a convolutional neural network layer and a Transformer structure; wherein, Convolutional neural network layer, used to refine and fuse speech features, capturing time-frequency patterns, vocalization characteristics, and reverberation information; The Transformer structure is used to maintain contextual information using the hidden state of the previous audio block.

[0039] In implementation, the convolutional neural network (CNN) structure can be used to extract features from audio blocks, capture local time-frequency information, phonation patterns, and reverberation characteristics, and generate more discriminative feature representations. The convolution kernel slides in the frequency domain and time domain, and can capture local time-frequency patterns in speech features. CNN extracts multi-dimensional features through multi-layer convolution and pooling operations, including time-frequency features: identifying key information of speech spectrum; phonation features: distinguishing pronunciation patterns of different speakers; reverberation information: identifying reverberation features generated when speaking near or far. Multi-channel features are fused after pooling and activation function processing to form a time-frequency representation of audio blocks. The Transformer structure maintains the context information of the previous audio block, captures the temporal dependency across blocks, and provides a global context for the target speaker's judgment. Transformer captures the dependency between speech blocks in the time dimension by encoding the hidden state (HiddenState) of the previous block. At the same time, state information is shared between each block, so as to achieve a comprehensive understanding of the call context when the target speaker makes judgments.

[0040] As an optional implementation, the multi-task output module in step S204 includes a voice activity detection head, a near-end voice judgment head, a speaker embedding extraction head and a target speaker judgment head; wherein, A voice activity detection head, used to output the voice activity probability of the audio block; A near-end speech judgment head, used to output the near-end probability of the audio block; Speaker embedding extraction head, used to output the speaker embedding vector of the audio block; The target speaker judgment head is used to integrate the speech activity probability, the proximal probability, the speaker embedding vector and the preset target speaker reference vector, and output the attribution probability that the audio block belongs to the target speaker.

[0041] In implementation, a voice activity detection head (VAD head) can be used to determine whether the current audio block is a "voiced block" or a "silent block", and output the voice activity probability P of the block being voiced. vad Specifically, CNN or RNN models can be used to determine the speech activity state of an audio block based on the time domain and frequency domain features of the audio block (such as energy, zero crossing rate, MFCC, etc.). Output P through a binary classification model (voice / silence) vad , indicating the probability that the audio block is a "voice block". The output of the voice activity detection head is a probability value in the range of [0, 1]. For example, P vad When the value is >0.5, it can be determined as a "voiced block", otherwise it is a "silent block". The near-end speech judgment head (NEAR head) can be used to determine whether the current audio block comes from the "near" or "far" end of the device, and the probability P of the block being a near-end block is output. near Specifically, the CNN model can be used to combine features such as sound energy, high-frequency retention, and reverberation level to classify the audio blocks into two categories (near-end / far-end) and output P near For example, a block with a high-frequency energy ratio > 0.6 and a reverberation score < 0.3 is more likely to be a near-end block. For example, the near-end speech judgment head outputs a probability value in the range of [0, 1], P near When >0.5, the block is considered to be from the near end, P near <0.5, it comes from the remote end. The speaker embedding vector e in the current audio block can be extracted through the speaker embedding extraction head (SV head) for speaker identity identification. Specifically, the CNN+Transformer model can be used for feature extraction to convert the time-frequency features into a fixed-dimensional speaker embedding vector (usually 256 or 512 dimensions). The pre-trained SpeakerVerification (SV) model is used for speaker embedding extraction to ensure that the embedding vector has a high degree of discrimination. e is used as an important feature input for target speaker judgment and compared with the target speaker reference vector (C ref) to perform similarity comparison. The target speaker judgment head (TARGET head) combines multi-dimensional information for comprehensive judgment and outputs the probability P that the current audio block belongs to the target speaker target .

[0042] As an optional implementation, the preset target speaker reference vector is a global embedding mean of speaker embedding vectors in near-end audio blocks that have been identified during a call.

[0043] In implementation, a target speaker reference vector C can be generated by extracting and aggregating the embedding vectors of multiple near-end audio blocks during a call. ref , as a standard reference for subsequent judgment of the target speaker. During a call, the near-end speech judgment head (NEAR head) is used to perform near-end / far-end judgment on each audio block. Only the audio block judged as "near-end" (P near >0.5) will be used to generate the target speaker reference vector. The speaker embedding vector e is extracted from the near-end audio block by the speaker embedding extraction head (SV head). Each near-end audio block can generate an embedding vector, which together reflect the characteristic information of the target speaker. By taking a weighted average or a simple average of all the identified near-end audio block embedding vectors, a global embedding mean is obtained as the target speaker reference vector C ref As more near-end audio blocks are identified, the model can continuously introduce new embedding vectors and use a sliding window or time-weighted approach to C ref Dynamic update; this update mechanism can ensure that C ref Adaptive adjustments are made when the target speaker's characteristics change, thereby improving the robustness of recognition. For example, in the first few seconds of a call, the model extracts embedding vectors from high-confidence near-end audio blocks, which are averaged to form the initial C ref , and used as the reference vector for subsequent blocks. Whenever a newly recognized near-end audio block is determined to belong to the target speaker, its embedding vector is included in C ref Calculation can be performed by using sliding window mean or time-weighted average. ref Dynamically update to adapt to changes in the target speaker's characteristics.

[0044] As an optional implementation, mean filtering or median filtering may be used to denoise the embedding vector to eliminate abnormal embedding or noise interference.

[0045] As an optional implementation, the target speaker determination head is also used for: Calculate the similarity score between the speaker embedding vector of the audio block to be identified and the target speaker reference vector; Based on the multi-dimensional feature weighted fusion mechanism, the voice activity probability, proximal probability and similarity score are weighted and calculated to determine the probability that the audio block belongs to the target speaker.

[0046] In implementation, the speaker embedding vector e of the current audio block can be calculated by comparing it with the target speaker reference vector C ref The similarity score between them quantifies the degree of match between the current block and the target speaker's features. e is a feature vector of fixed dimension (usually 256 or 512 dimensions) representing the speaker's features of the current block. The similarity score S can be calculated using cosine similarity or dot product similarity. If the value range of S is [-1, 1], the higher the similarity, the higher the match between the current block and the target speaker's features. Through the multi-dimensional feature weighted fusion mechanism, the feature information from different sources is integrated to improve the accuracy of the target speaker's judgment, and the output audio block belongs to the target speaker's attribution probability P target The target speaker judgment head can be P vad , P near And S are weighted summed to generate the final P target Among them, the weights can be automatically optimized according to the performance of the model during training, or they can be set through experience. For example, when the background noise of the call is large, P near The weight of .

[0047] As an optional implementation, an attention mechanism is introduced in the target speaker judgment head. By calculating the attention weight between the speaker embedding vector of the audio block and the target speaker reference vector, the target judgment is weightedly fused to improve the accuracy of the audio block attribution judgment.

[0048] In implementation, the speaker embedding vector e of the current audio block and the target speaker reference vector C can be calculated through the attention mechanism ref The weight relationship between P and S is dynamically adjusted to optimize the judgment result of the target speaker. Among them, the attention mechanism is a weighted feature aggregation mechanism. By calculating the correlation (similarity) between different features, a higher weight is assigned to important features, thereby enhancing the model's attention to key information. vad , P near The weights of and S make the feature fusion more targeted, thereby improving the target speaker attribution probability P output by the target judgment head. target Through the attention mechanism, according to e(chunk) and C ref The matching degree of different features is adjusted. When S is high, the attention mechanism can give it a higher weight, making it have a greater impact on the judgment of the target speaker. In the scene with high noise, Pnear The weight is relatively increased to reduce the interference of environmental noise. For example: at the beginning of a call, the target speaker reference vector C ref Not yet fully stable, P vad and P near The weight of is higher. In the middle of the call, the target speaker reference vector C ref Gradually converge, the matching degree of S improves. The attention mechanism automatically increases the weight of S, making the judgment more dependent on the matching degree of the speaker's embedded features. near and P vad The reliability is relatively improved.

[0049] Step S205: jointly optimize the speaker speech recognition model using a multi-task loss function.

[0050] In practice, the model can be jointly optimized through a multi-task loss function, so that it has good generalization ability in multiple tasks. Specifically, in the actual training process, the model can be trained end-to-end using a joint multi-task learning method. All subtasks (including voice activity detection, proximal judgment, speaker embedding extraction, and target speaker judgment) share the front encoder module and output the prediction results through different task heads. The multi-task loss function can be composed of the following parts: Voice activity detection loss function: cross entropy loss is used to optimize the model to determine whether an audio block is a voiced block; Proximal judgment loss function: cross entropy loss is used to determine whether the audio block comes from the proximal end; Embedding extraction loss function: AM-Softmax (AdditiveMarginSoftmax) or TripletLoss is used to improve the distinguishability of the embedding vector; Target speaker judgment loss function: cross entropy loss is used to optimize the judgment result of whether the audio block belongs to the target speaker; The final total loss function of the model is L total It can be expressed as the weighted sum of multiple subtask loss functions: L total =αL vad +βL near +γL sv +σL target ; Among them, α, β, γ and σ are the weight coefficients of each subtask loss, which can be tuned by cross-validation or dynamically adjusted using a non-fixed-proportional learning strategy (such as GradNorm).

[0051] During the model training process, all audio blocks are labeled and predicted for four tasks simultaneously. Through this multi-task joint optimization strategy, not only the performance of a single task is improved, but also the information sharing and generalization capabilities between tasks are enhanced, effectively improving the recognition accuracy and robustness of the target speaker.

[0052] As an optional implementation, the loss function in step S205 includes a speech activity detection loss function, a proximal judgment loss function, an embedding extraction loss function and a target speaker judgment loss function; wherein, Voice activity detection loss function, used to distinguish whether an audio block is voiced or silent; Near-end judgment loss function, used to distinguish audio blocks as near-end or far-end; Embedding extraction loss function, used to optimize embedding discrimination; The target speaker judgment loss function maximizes the attribution probability of the target speaker's audio block and minimizes the misjudgment probability of other speakers' audio blocks.

[0053] In implementation, the speech activity detection loss function (L vad ), the optimization model distinguishes the "voice state" and "silence state" of the audio block, and improves the accuracy of voice activity detection. Specifically, the voice activity detection head determines whether the audio block contains voice activity and outputs P vad The goal is to label the voiced block as 1 and the silent block as 0, forming a binary classification task. Maximize P vad The probability of being in the voiced block is minimized, while the probability of misjudgment of the silent block is minimized, thus improving the accuracy of voice activity detection. near ), the optimization model distinguishes the "near-end block" and "far-end block" of the audio block, and improves the accuracy of the near-end judgment. The near-end speech judgment head outputs P near , to determine whether the audio block is from the "near end" or the "far end" of the device. The goal is to correctly distinguish between near-end blocks (labeled as 1) and far-end blocks (labeled as 0). By embedding the extraction loss function (L sv ), optimize the discrimination of the speaker embedding vector and improve the ability to separate the target speaker from other speakers. The speaker embedding extraction head extracts the embedding vector e of the audio block and compares it with the target speaker reference vector C ref For comparison. The speaker embedding extraction task can use the AM-Softmax (AdditiveMarginSoftmax) loss function, in which the scaling factor is used to control the distribution range and the angle interval is used to increase the distinction between the target and non-target speakers. target), optimize the model's judgment on the attribution of the target speaker block, maximize the correct attribution probability of the target block, and minimize the misjudgment probability of other speaker blocks. The target speaker judgment head outputs the probability P that the current audio block belongs to the target speaker through weighted fusion of multi-dimensional features. target . Target speaker judgment is a binary classification task, which can be optimized using the cross entropy loss function. After the target speaker's true label (1 for the target block, 0 for the non-target block) is optimized, the probability of belonging to the target speaker block will be further improved, while reducing the interference of other speaker blocks, improving the overall robustness of the system.

[0054] The embodiment of the present application provides a method for speaker speech recognition, the method comprising: acquiring the call audio to be recognized in real time, and dividing the call audio to be recognized into a number of audio blocks; for each audio block, extracting the speech features of the audio block, and inputting them into a pre-trained speaker speech recognition model; outputting the probability of the audio block belonging to the target speaker through the speaker speech recognition model; if the probability of the audio block belonging to the target speaker is greater than a preset speaker belonging probability threshold, then determining that the audio block belongs to the target speaker. The technical solution provided by the embodiment of the present application brings at least the following beneficial effects: In traditional technologies, the judgment of the target speaker usually relies on a single feature and cannot make full use of multi-dimensional information, resulting in low recognition accuracy. The present application introduces a multi-task model to convert P vad (Voice Activity), P near (proximal judgment) and S (embedded similarity) for feature weighted fusion, which realizes the adaptive adjustment of different features at different stages of the call and improves the accuracy of the target speaker judgment. The traditional method uses a fixed reference vector for matching, which cannot be adaptively adjusted according to the changes in the characteristics of the target speaker during the call, resulting in a decrease in recognition accuracy during long calls. This application calculates the reference vector C of the target speaker. ref , and during the call, ref Dynamically update and continuously optimize C according to the embedding mean of high-confidence near-end audio blocks ref , thereby ensuring that the model maintains a high judgment accuracy during long calls. This application introduces an attention mechanism to monitor P according to the feature changes of the audio block. vad , P near Dynamic weight adjustment is performed on S to make feature weighting more flexible, thereby enhancing the target speaker judgment ability in complex scenarios.

[0055] It should be understood that although Figure 1 and Figure 2The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 and Figure 2 At least part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.

[0056] It can be understood that the same / similar parts between the various embodiments of the above method in this specification can refer to each other, and each embodiment focuses on the differences from other embodiments. For related points, please refer to the description of other method embodiments.

[0057] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program that can be executed on the processor. When the processor executes the computer program, the method steps of dynamically identifying flight changes are implemented. Figure 3 A schematic diagram of the structure of a computer device provided in an embodiment of the present application is shown in FIG. Figure 3 As shown, the computer device may include a processor 301, a system bus 302, a non-volatile storage medium 303, an internal memory 304, a network interface 305, a display screen 306, and an input device 307. Among them, the non-volatile storage medium 303 stores an operating system 3031 and a computer program 3032. The processor 301 is used to execute the computer program 3032 to implement the method steps of speaker speech recognition. The system bus 302 is used to connect the processor 301, the non-volatile storage medium 303, the internal memory 304, the network interface 305, the display screen 306, and the input device 307 to ensure efficient communication between the components. The internal memory 304 is used to temporarily store the running programs and data, helping the processor 301 to quickly access the required information, thereby improving the overall system performance. The network interface 305 (such as a network card) enables the computer device to be connected to a local area network or the Internet to achieve data transmission and remote communication. The display screen 306 is used to present the speaker speech recognition information processed by the computer device to the user in the form of graphics or text. The input device 307 (such as a keyboard, a mouse, a touch screen, etc.) is used to allow the user to input voice data and commands to the computer device to achieve interactive operations with the computer device.

[0058] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0059] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0060] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data for analysis, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0061] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0062] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0063] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A method for speaker speech recognition, characterized in that: The method comprises: Acquire the call audio to be identified in real time, and divide the call audio to be identified into a plurality of audio blocks; For each of the audio blocks, extract the speech features of the audio block and input them into a pre-trained speaker speech recognition model; Outputting the probability that the audio block belongs to the target speaker through the speaker speech recognition model; If the attribution probability of the audio block belonging to the target speaker is greater than a preset speaker attribution probability threshold, it is determined that the audio block belongs to the target speaker.

2. The method according to claim 1, characterized in that The training process of the speaker speech recognition model includes: Obtaining historical call audio, and dividing the historical call audio into a plurality of audio blocks; Performing training annotation on each of the audio blocks; For each of the audio blocks, extracting speech features of the audio block; Constructing the speaker speech recognition model, wherein the speaker speech recognition model includes an encoder module and a multi-task output module; The speaker speech recognition model is jointly optimized using a multi-task loss function.

3. The method according to claim 2, characterized in that The training and labeling of each audio block includes: Marking the vocal state of each of the audio blocks, where the vocal state includes a voiced state and a silent state; Based on the pre-trained NEAR pre-model, each of the audio blocks is proximally labeled according to the sound energy, reverberation level, and high-frequency retention of the audio block; A target speaker is labeled for each of the audio blocks.

4. The method according to claim 1 or 3, characterized in that: The speech features include acoustic features, distance-related features and high-level speech features.

5. The method according to claim 2, characterized in that: The encoder module includes a convolutional neural network layer and a Transformer structure; wherein, The convolutional neural network layer is used to refine and fuse the speech features to capture time-frequency patterns, vocalization characteristics and reverberation information; The Transformer structure is used to maintain context information using the hidden state of the previous audio block.

6. The method according to claim 2, characterized in that The multi-task output module includes a voice activity detection head, a near-end voice judgment head, a speaker embedding extraction head and a target speaker judgment head; wherein, The voice activity detection head is used to output the voice activity probability of the audio block; The near-end speech judgment head is used to output the near-end probability of the audio block; The speaker embedding extraction head is used to output the speaker embedding vector of the audio block; The target speaker judgment head is used to integrate the voice activity probability, the proximal probability, the speaker embedding vector and a preset target speaker reference vector to output the attribution probability that the audio block belongs to the target speaker.

7. The method according to claim 6, characterized in that The preset target speaker reference vector is the global embedding mean of speaker embedding vectors in the near-end audio blocks that have been identified during the call.

8. The method according to claim 6, characterized in that The target speaker judgment head is also used for: Calculate the similarity score between the speaker embedding vector of the audio block to be identified and the target speaker reference vector; Based on a multi-dimensional feature weighted fusion mechanism, the voice activity probability, the proximal probability and the similarity score are weightedly calculated to determine the attribution probability that the audio block belongs to the target speaker.

9. The method according to claim 6, characterized in that The method further comprises: An attention mechanism is introduced into the target speaker judgment head. By calculating the attention weight between the speaker embedding vector of the audio block and the target speaker reference vector, feature weighted fusion is performed on the target judgment to improve the accuracy of the audio block attribution judgment.

10. The method according to claim 2, characterized in that The loss function includes a speech activity detection loss function, a proximal judgment loss function, an embedding extraction loss function and a target speaker judgment loss function; wherein, The voice activity detection loss function is used to distinguish whether the audio block is voiced or silent; The near-end judgment loss function is used to distinguish whether the audio block is near-end or far-end; The embedding extraction loss function is used to optimize the embedding discrimination; The target speaker judgment loss function maximizes the attribution probability of the target speaker's audio block and minimizes the misjudgment probability of other speakers' audio blocks.

Citation Information

Patent Citations

  • Speech recognition method and device, electronic equipment and storage medium

    CN116913268A

  • Audio signal processing device, audio signal processing method, and audio signal processing program

    CN118633300A

  • Junk call identification method and device, computer equipment and storage medium

    CN119629636A

  • Identifying far-end sound

    US20090150149A1

  • Speaking object detection in multi-human-machine interaction scenario

    WO2024032159A1

Cited By

  • Speaker recognition method and device and speaker feature vector extraction method and device

    CN121122288A

  • Audio processing model training method and device, equipment and storage medium

    CN121306103A