Voiceprint recognition method, voiceprint recognition device, voiceprint recognition equipment, medium and product
Through the multi-level cascading voiceprint recognition model combined with waveform direct comparison and depth feature analysis, the problem of insufficient real-time performance of voiceprint recognition methods based on depth features is solved, and more efficient voiceprint recognition performance and real-time performance are achieved.
Patent Information
- Application Number
- CN202510210083.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
AI Technical Summary
Although the voiceprint recognition method based on deep features improves the accuracy of voiceprint recognition, the real-timeness still needs to be improved.
A multi-level cascading vocalprint recognition model is adopted, and a multi-level cascading architecture combining direct waveform comparison and depth feature analysis is used to directly compare waveforms through the waveform comparison layer to improve real-time performance, and multi-level collaborative recognition is used when waveform comparison is not passed.
It significantly improves the real-time and performance of voiceprint recognition, and can more effectively deal with complex scenarios such as audio content differences, environmental noise and changes in speaker status.
Smart Images

Figure CN120148520A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biometric identification, and in particular to a voiceprint recognition method, device, equipment, medium and product. Background Art
[0002] In the current field of voiceprint recognition, the mainstream methods are mainly divided into two categories: text-dependent and text-independent. The text-dependent methods require the speaker to read fixed content, which has high accuracy but is greatly limited in practical applications. Although the text-independent methods have a wider range of application scenarios, they face the problem of inconsistent features caused by content differences. Firstly, there is the uncertainty of the speech content. Different text contents will lead to significant differences in acoustic features, including differences in speech content and length. Secondly, there are the influences of environmental factors, including signal distortion caused by background noise, equipment differences, etc. In addition, the emotional state, health status, etc. of the speaker will also affect the stability of the voice features.
[0003] With the development of deep learning technology, although the voiceprint recognition method based on deep features has improved the accuracy of voiceprint recognition, the real-time performance still needs to be improved. Summary of the Invention
[0004] The present invention provides a voiceprint recognition method, device, equipment, medium and product, which solves the defect that although the voiceprint recognition method based on deep features has improved the accuracy of voiceprint recognition, the real-time performance still needs to be improved.
[0005] The present invention provides a voiceprint recognition method, including: Obtaining a pair of audio to be recognized; Inputting the pair of audio to be recognized into a pre-trained multi-level cascaded voiceprint recognition model to obtain a voiceprint recognition result output by the multi-level cascaded voiceprint recognition model; Wherein, the multi-level cascaded voiceprint recognition model includes an input layer, a waveform comparison layer, an acoustic feature comparison layer and a decision layer; The input layer is used to input the pair of audio to be recognized into the waveform comparison layer and the acoustic feature comparison layer respectively; The waveform comparison layer is used to compare the waveform similarity of the pair of audio to be recognized to obtain a waveform similarity score, and output the waveform similarity score to the decision layer; The acoustic feature comparison layer is used to compare the acoustic feature similarity of the pair of audio to be recognized in response to an instruction from the decision layer to obtain an acoustic feature similarity score, and output the acoustic feature similarity score to the decision layer, where the acoustic feature similarity score includes at least one of a spectrum feature similarity score, a text content similarity score, a speaker vector similarity score and a cloned audio similarity score; The decision-making layer is used to determine whether the waveform similarity score reaches a preset value. If so, the waveform similarity score is used as the voiceprint recognition result. Otherwise, the acoustic feature similarity score is obtained and used as the voiceprint recognition result.
[0006] As an embodiment, the acoustic feature comparison layer includes a spectral feature analysis module, a text content recognition module, a speaker vector recognition module, and a voice cloning comparison module. The spectral feature analysis module is used to extract the spectral features of the audio pair to be recognized to determine the spectral feature similarity score, and transmit the spectral features of the audio pair to be recognized to the text content recognition module and the speaker vector recognition module. The text content recognition module is used to judge whether the text contents of the audio pair to be recognized are consistent according to the spectral features of the audio pair to be recognized, and obtain the text content similarity score. The speaker vector recognition module is used to judge whether the speaker vectors of the audio pair to be recognized are consistent according to the spectral features of the audio pair to be recognized, and obtain the speaker vector similarity score. The voice cloning comparison module is used to extract the prosody features and speaker features of the audio pair to be recognized, generate a cloned audio pair with consistent text content according to the prosody features, the speaker features, and the preset text, and judge the similarity of the cloned audio pair to obtain the cloned audio similarity score.
[0007] As an embodiment, the spectral feature analysis module includes a preprocessing and framing sub-module, a short-time Fourier transform sub-module, a Mel spectral feature extraction sub-module, a Mel spectral feature calculation sub-module, a dynamic feature extraction sub-module, and a feature normalization sub-module which are arranged in sequence.
[0008] As an embodiment, the text content recognition module includes an encoder, a decoder, and a content similarity calculation sub-module. The encoder is used to extract the deep features of the audio pair to be recognized from the spectral features of the audio pair to be recognized based on relative position encoding and an improved self-attention mechanism. The decoder is used to determine the text content of the audio pair to be recognized according to the deep features of the audio pair to be recognized. The content similarity calculation sub-module is used to calculate the similarity of the text content of the audio pair to be recognized based on a mixed metric of edit distance and semantic similarity to obtain the text content similarity score.
[0009] As an embodiment, the speaker vector recognition module includes a time delay neural network sub-module, a statistical pooling sub-module, and a feature vector extraction and judgment sub-module. The time-delay neural network sub-module is used to extract features of different time scales of the audio pair to be recognized from the spectral features of the audio pair to be recognized; The statistical pooling sub-module is used to calculate the statistical features of the audio pair to be recognized based on the mean and variance according to the features of different time scales of the audio pair to be recognized; The feature vector extraction and judgment sub-module is used to map the statistical features of the audio pair to be recognized to a low-dimensional discriminant space, obtain the speaker vector of the audio pair to be recognized and calculate the similarity of the speaker vectors, and obtain the speaker vector similarity score.
[0010] As an embodiment, the waveform comparison layer is implemented based on a deep siamese network structure, and the deep siamese network structure is used to extract the time-domain features of the audio pair to be recognized based on a multi-layer one-dimensional convolutional network, perform weighted processing on the time-domain features of the audio pair to be recognized based on an attention mechanism, and calculate the similarity of the weighted time-domain features of the audio pair to be recognized based on the embedding space obtained by metric learning, so as to obtain the waveform similarity score.
[0011] The present invention also provides a voiceprint recognition device, including: An acquisition module, which is used to acquire an audio pair to be recognized; A recognition module, which is used to input the audio pair to be recognized into a pre-trained multi-level cascaded voiceprint recognition model to obtain a voiceprint recognition result output by the multi-level cascaded voiceprint recognition model; Wherein, the multi-level cascaded voiceprint recognition model includes an input layer, a waveform comparison layer, an acoustic feature comparison layer and a decision layer; The input layer is used to input the audio pair to be recognized into the waveform comparison layer and the acoustic feature comparison layer respectively; The waveform comparison layer is used to compare the waveform similarity of the audio pair to be recognized, obtain a waveform similarity score, and output the waveform similarity score to the decision layer; The acoustic feature comparison layer is used to compare the acoustic feature similarity of the audio pair to be recognized, obtain an acoustic feature similarity score, and output the acoustic feature similarity score to the decision layer, and the acoustic feature similarity score includes at least one of a spectral feature similarity score, a text content similarity score, a speaker vector similarity score and a cloned audio similarity score; The decision layer is used to determine whether the waveform similarity score reaches a preset value. If so, the waveform similarity score is used as the voiceprint recognition result. Otherwise, the acoustic feature similarity score is obtained and the acoustic feature similarity score is used as the voiceprint recognition result.
[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the voiceprint recognition method as described in any one of the above is implemented.
[0013] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the voiceprint recognition method as described in any one of the above is implemented.
[0014] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the voiceprint recognition method as described in any one of the above is implemented.
[0015] For the voiceprint recognition method, device, equipment, medium, and product provided by the present invention, the multi-level cascaded voiceprint recognition model adopts a multi-level cascaded architecture that combines waveform direct comparison and depth feature analysis. The waveform comparison layer directly compares waveforms, improving the real-time performance of voiceprint recognition. In the case where the waveform comparison fails, a multi-level collaborative recognition method is adopted, effectively improving the performance and reliability of voiceprint recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 is one of the schematic flowcharts of the voiceprint recognition method provided by the present invention.
[0018] Figure 2 is the schematic data processing flowchart of the waveform comparison layer provided by the present invention.
[0019] Figure 3 is the second schematic flowchart of the voiceprint recognition method provided by the present invention.
[0020] Figure 4 is the schematic structural diagram of the spectrum feature analysis module provided by the present invention.
[0021] Figure 5 is the schematic structural diagram of the text content recognition module provided by the present invention.
[0022] Figure 6 is the schematic structural diagram of the speaker vector recognition module provided by the present invention.
[0023] Figure 7It is a schematic diagram of the data processing flow of the voice cloning comparison module provided by the present invention.
[0024] Figure 8 It is a schematic structural diagram of the voiceprint recognition device provided by the present invention.
[0025] Figure 9 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0026] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0027] It should be noted that all actions of obtaining signals, information or data in the present invention are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located and obtaining the authorization given by the owner of the corresponding device.
[0028] The prior art usually adopts a single feature extraction path, and cannot dynamically select the optimal processing strategy according to the characteristics of the input audio, resulting in low processing efficiency. Even highly similar audios need to go through the complete feature extraction process, and the accuracy drops significantly when processing audios with large content differences. There is a lack of an effective fault tolerance mechanism, and environmental noise easily leads to recognition failure.
[0029] In response to this, the present invention proposes a voiceprint recognition method based on a multi-level cascaded voiceprint recognition model, adopting a strategy of combining a fast path and an accurate path, and realizing accurate and efficient voiceprint recognition through the coordination of multiple levels such as direct waveform comparison, spectral feature analysis, content recognition and voice cloning. Following the principle from simple to difficult, that is, first try methods with lower computational complexity and only use more complex processing processes when necessary, it is applicable to cross-applications in the fields of artificial intelligence and biometric recognition, especially the application of multi-level cascaded architectures and deep learning in voiceprint recognition, especially in dealing with complex scenarios such as audio content differences, environmental noise and speaker state changes.
[0030] Figure 1 It is one of the schematic flowcharts of the voiceprint recognition method provided by the present invention. As Figure 1 shown, the present invention provides a voiceprint recognition method, including step S100-step S200.
[0031] Step S100, obtain the audio pair to be recognized. Given two input audio sequences and , where n and m respectively represent the sequence lengths.
[0032] Step S200: Input the audio pair to be recognized into a pre-trained multi-level cascaded voiceprint recognition model to obtain the voiceprint recognition result output by the multi-level cascaded voiceprint recognition model. The multi-level cascaded voiceprint recognition model can be used to identify whether the content of the audio pair is the same or whether the audio pair is spoken by the same speaker. In the embodiment of the present invention, taking the identification of whether the speakers are the same as an example, the multi-level cascaded voiceprint recognition model can be expressed as F(A, B) → {0, 1}, where the output of 0 indicates different speakers, and the output of 1 indicates the same speaker.
[0033] Among them, the multi-level cascaded voiceprint recognition model includes an input layer, a waveform comparison layer, an acoustic feature comparison layer, and a decision layer.
[0034] The input layer is used to input the audio pair to be recognized into the waveform comparison layer and the acoustic feature comparison layer respectively.
[0035] The waveform comparison layer is used to compare the waveform similarity of the audio pair to be recognized, obtain a waveform similarity score, and output the waveform similarity score to the decision layer.
[0036] The acoustic feature comparison layer is used to compare the acoustic feature similarity of the audio pair to be recognized in response to the instruction of the decision layer, obtain an acoustic feature similarity score, and output the acoustic feature similarity score to the decision layer. The acoustic feature similarity score includes at least one of a spectral feature similarity score, a text content similarity score, a speaker vector similarity score, and a cloned audio similarity score.
[0037] The decision layer is used to determine whether the waveform similarity score reaches a preset value. If so, use the waveform similarity score as the voiceprint recognition result; otherwise, obtain the acoustic feature similarity score and use the acoustic feature similarity score as the voiceprint recognition result.
[0038] The data processing flow of the multi-level cascaded voiceprint recognition model is that first, the input layer receives the original pair of audio signal sequences. The audio signal sequences are passed to the waveform comparison layer and the acoustic feature comparison layer in parallel. The waveform comparison layer directly compares the waveforms of the audio signal sequences. To improve the processing speed, the acoustic feature comparison layer can simultaneously perform preliminary processing on the audio signal sequences. The decision layer determines whether the waveform similarity score reaches a preset value. If so, it directly uses the waveform similarity score as the voiceprint recognition result and controls the stop instruction of the acoustic feature comparison layer. Otherwise, it sends an instruction to the acoustic feature comparison layer. The acoustic feature comparison layer performs speech content recognition and speaker vector recognition based on the features of the audio signal sequences obtained from the preliminary processing, completing the analysis process from acoustic features to speaker identity. Preferably, complex situations that are difficult to directly determine can be processed through acoustic wave cloning comparison.
[0039] It can be understood that the multi-level cascaded voiceprint recognition model adopts a multi-level cascaded architecture that combines direct waveform comparison and deep feature analysis. The waveform comparison layer directly compares the waveforms, improving the real-time performance of voiceprint recognition. In the case where the waveform comparison fails, a multi-level collaborative recognition method is adopted, effectively improving the performance and reliability of voiceprint recognition.
[0040] As Figure 2 shown, based on the above embodiments, as an optional embodiment, the waveform comparison layer is implemented based on a deep siamese network structure. The deep siamese network structure is used to extract the time-domain features of the audio pair to be recognized based on a multi-layer one-dimensional convolutional network, perform weighted processing on the time-domain features of the audio pair to be recognized based on an attention mechanism, and calculate the similarity of the weighted time-domain features of the audio pair to be recognized in the embedding space obtained by metric learning to obtain the waveform similarity score.
[0041] The waveform comparison layer adopts a deep siamese network structure. The input is the waveform of the original audio signal sequence. The time-domain features are extracted through a multi-layer one-dimensional convolutional network. In the network design, special consideration is given to the extraction of multi-scale features, and different convolutional kernel sizes are used to capture sound features at different time scales. To retain detailed information, a residual connection structure is added to the network. After feature extraction, an attention mechanism is used to weight the features at different time positions to highlight important sound segments. Finally, the similarity of the two audio segments is calculated in the embedding space obtained by metric learning. The design intention of the waveform comparison layer is to quickly screen out highly similar audio segments, thus avoiding subsequent more complex processing processes.
[0042] In actual voiceprint recognition application scenarios, two key challenges are faced: one is the need to process a large number of voiceprint comparison requests, requiring the model to have the ability to respond quickly; the other is that the quality, length, and content of the input audio vary, and the model needs to be sufficiently robust. Traditional voiceprint recognition methods usually require a complex feature extraction process, such as calculating Mel spectrum features, extracting acoustic features, etc. Although these processes can provide rich feature information, the computational cost is relatively high and it is not suitable as the first processing link of the model.
[0043] Based on such actual needs, the present invention designs a waveform comparison layer as a fast processing channel of the model. When two audio segments are from the same speaker and have highly similar content, they will show obvious similarity in the time-domain waveform, and a quick judgment can be obtained by directly comparing the waveform features without the need for complex frequency-domain conversion and feature extraction. This design idea is similar to the "fast channel" mechanism in the human auditory system, that is, in some cases, people can complete the preliminary voice recognition through fast feature matching without the need to analyze all features of the voice in detail.
[0044] In the specific design of the waveform comparison layer, the present invention adopts a deep siamese network structure instead of traditional signal processing methods. This choice is based on the following considerations: First, the deep siamese network structure can adaptively learn time-domain features of different scales and has stronger feature extraction ability than fixed signal processing algorithms; second, the structural characteristics of the deep siamese network structure enable the model to focus on learning the similarity metric between audio pairs; finally, through the design of sharing weights, the number of model parameters is effectively controlled, which is conducive to realizing fast inference.
[0045] In the design of the deep siamese network structure, the present invention particularly focuses on the ability to capture multi-scale features. By using one-dimensional convolutional layers with different kernel sizes, it is possible to simultaneously focus on local acoustic details and larger-scale sound structure features. The introduction of residual connections ensures that shallow features can be directly transmitted to the deep layer, which is crucial for maintaining the detail features of the audio. The design of the attention mechanism can automatically identify and focus on the key time segments in voiceprint discrimination, thereby improving the efficiency and accuracy of feature extraction.
[0046] Given the input audio signal sequences A and B, first perform preprocessing normalization: , , where μ and σ represent the mean and standard deviation of the sequence respectively.
[0047] The multi-scale feature extraction network E consists of L one-dimensional convolutional layers, and the output of each layer can be expressed as: h_l = f_l(h_{l-1}) = ReLU(Conv1D(h_{l-1}) + ResConn(h_{l-1})), where h_0 = Â or , ResConn represents the residual connection. Each convolutional layer uses convolutional kernels of different scales: k_l ∈ {3, 5, 7, 9} to capture multi-scale time-domain features.
[0048] The attention mechanism is used to weight the time-domain features: α = Softmax(w^T tanh(Wh_L + b)) e = ∑(α_i * h_L_i), where w, W, b are learnable parameters, and e is the final embedding vector.
[0049] The similarity calculation uses the cosine distance: s_w = cosine(e_A, e_B) = (e_A · e_B) / (||e_A|| ||e_B||).
[0050] The training objective uses the contrastive loss function: Loss = y * max(0, m - s_w) + (1 - y) * max(0, s_w - m), where y ∈ {0, 1} is the label and m is the margin parameter.
[0051] It can be understood that by designing the waveform contrast layer, the present invention can not only significantly improve the processing efficiency of the model, but also provide a reliable pre-screening mechanism for the subsequent precise recognition module. Practice shows that in the case of good audio quality and moderate length, the waveform contrast layer alone can achieve a high recognition accuracy. For those cases where the recognition result cannot be determined by the waveform contrast layer, the subsequent precise recognition process is started, and the hierarchical processing strategy effectively balances the efficiency and accuracy requirements of the model.
[0052] As Figure 3 shown, on the basis of the above embodiments, as an optional embodiment, the acoustic feature contrast layer includes a spectral feature analysis module, a text content recognition module, a speaker vector recognition module, and a voice cloning comparison module.
[0053] The spectral feature analysis module is used to extract the spectral features of the audio pair to be recognized to determine the spectral feature similarity score, and transmit the spectral features of the audio pair to be recognized to the text content recognition module and the speaker vector recognition module.
[0054] The text content recognition module is used to determine whether the text content of the audio pair to be recognized is consistent according to the spectral features of the audio pair to be recognized, and obtain the text content similarity score.
[0055] The speaker vector recognition module is used to determine whether the speaker vectors of the audio pair to be recognized are consistent according to the spectral features of the audio pair to be recognized, and obtain the speaker vector similarity score.
[0056] The voice cloning comparison module is used to extract the prosodic features and speaker features of the audio pair to be recognized, generate a cloned audio pair with consistent text content according to the prosodic features, the speaker features and the preset text, determine the similarity of the cloned audio pair, and obtain the cloned audio similarity score.
[0057] Two input audio sequences and , the processing process of the waveform comparison layer can be expressed as W(A, B) → s_w, and the output similarity score s_w ∈ [0, 1]. The processing process of the spectral feature analysis module can be expressed as M(A, B) → (S_A, S_B), and the output Mel spectral feature matrix is obtained. S_A and S_B respectively represent the Mel spectral feature matrices of the audio signal sequences A and B. The processing process of the text content recognition module can be expressed as C(S_A, S_B) → (T_A, T_B, s_c), where T_A represents the text content of the audio signal sequence A, T_B represents the text content of the audio signal sequence B, and s_c represents the similarity score between T_A and T_B, that is, the text content similarity score. The processing process of the speaker vector recognition module can be expressed as V(S_A, S_B) → s_v, and s_v represents the speaker vector similarity score. The processing process of the voice cloning comparison module can be expressed as K(A, B) → s_k, and s_k represents the cloned audio similarity score. Optionally, the processing process of the decision layer can be expressed as D(W(A, B), M(A, B), C(S_A, S_B), V(S_A, S_B), K(A, B)), and the processing process of the multi-level cascaded voiceprint recognition model can be expressed as F(A, B) = D(W(A, B), M(A, B), C(S_A, S_B), V(S_A, S_B), K(A, B)).
[0058] It can be understood that the present invention extracts high-quality acoustic feature representations through the spectral feature analysis module, analyzes the consistency of the speech content through the text content recognition module, generates stable identity features through the speaker vector recognition module, processes complex situations with large content differences through the voice cloning comparison module, and constructs a complete voiceprint recognition link by using the collaborative interaction of multiple levels such as spectral feature analysis, content recognition and voice cloning, so as to achieve accurate and efficient voiceprint recognition.
[0059] As Figure 4 shown, on the basis of the above embodiments, as an alternative embodiment, the spectral feature analysis module includes a preprocessing and framing sub-module, a short-time Fourier transform sub-module, a Mel spectral feature extraction sub-module, a Mel spectral feature calculation sub-module, a dynamic feature extraction sub-module, and a feature normalization sub-module that are arranged in sequence.
[0060] The performance of an audio signal in the time domain often cannot intuitively reflect its internal characteristics. In the frequency domain, various characteristics of the audio signal will present a clearer structure. Especially in the task of voiceprint recognition, personalized characteristics such as the vocal tract characteristics and pronunciation habits of the speaker are more prominent in the frequency domain. Therefore, the core idea of designing the spectral feature analysis module is to convert the time-domain signal into a more discriminative frequency-domain representation space. Although traditional spectral analysis methods such as Fourier transform can complete time-frequency conversion, they cannot well simulate the characteristics of the human auditory system. Research shows that human perception of sound presents a non-linear distribution in the frequency dimension and has finer resolution ability in the low-frequency region. Based on this understanding, the present invention selects Mel spectral features as the core sound representation method.
[0061] In module design, the balance of time-frequency resolution, the selection of window size directly affects the time and frequency resolution ability of spectral features; for the non-linear mapping of frequency scales, it is necessary to design a suitable Mel filter bank to simulate the auditory characteristics of the human ear; the normalization processing of features can reduce the influence of environmental and device differences.
[0062] Given an input audio sequence , the spectral feature extraction process can be formalized into the following steps.
[0063] Preprocessing normalization: .
[0064] Framing processing: , , where w is the window size (usually 25 ms) and h(w) is the Hamming window function: .
[0065] Short-time Fourier transform (STFT): , .
[0066] Mel filter bank design: , where the center frequency of each filter follows the Mel scale conversion: .
[0067] Mel spectral feature calculation: , where S is the final Mel spectral feature matrix.
[0068] Dynamic feature extraction: .
[0069] Feature normalization: , where μ_S and σ_S are the mean and standard deviation of the features respectively.
[0070] The output feature matrix of the spectral feature analysis module will be simultaneously passed to the speech content recognition module and the speaker vector extraction module as input features.
[0071] Optionally, to improve the practicality of the model, the present invention also introduces a series of engineering optimization measures, implements an efficient parallel computing mechanism in the spectral feature analysis layer, significantly reduces the processing delay, adopts quantization compression technology in model deployment, and reduces the computing resource requirements. At the same time, by establishing a feature cache pool and a dynamic threshold adjustment mechanism, the processing efficiency and adaptability of the model are further improved.
[0072] It can be understood that the present invention converts the audio into a Mel spectrogram, uses the same preprocessing parameters as the Whisper model to ensure the consistency of features. The generation of the spectrogram involves the precise setting of multiple parameters such as window size, overlap rate, and frequency resolution. The selection of parameters needs to balance the time-frequency resolution. The present invention preferably uses a Hamming window of 25 ms, a frame shift of 10 ms, and 80 Mel filter banks, which not only ensures sufficient frequency resolution but also captures the dynamic features of the speech.
[0073] As Figure 5 shown, on the basis of the above embodiments, as an optional embodiment, the text content recognition module includes an encoder, a decoder, and a content similarity calculation sub-module.
[0074] The encoder is used to extract the deep features of the audio pair to be recognized from the spectral features of the audio pair to be recognized based on relative position encoding and an improved self-attention mechanism.
[0075] The decoder is used to determine the text content of the audio pair to be recognized according to the deep features of the audio pair to be recognized.
[0076] The content similarity calculation sub-module is used to calculate the similarity of the text content of the audio pair to be recognized based on a hybrid metric of edit distance and semantic similarity to obtain the text content similarity score.
[0077] The text content recognition module plays a crucial hub role in the entire voiceprint recognition model. When speakers express the same content, the comparison of their voice features will be more reliable and accurate. However, in practical applications, it is difficult to require speakers to follow fixed text content. Therefore, the core task of this module is to accurately understand the spoken content, provide content similarity information for subsequent feature comparison, and thus guide the model to select the optimal recognition strategy.
[0078] Traditional ASR models often have large computational requirements and high latency, making them unsuitable as an auxiliary module for voiceprint recognition. Secondly, they need to handle diverse speech inputs, including different languages, accents, and speaking styles. Finally, it is about how to extract and quantify the similarity of content so that it can effectively guide the subsequent recognition process.
[0079] To this end, the present invention adopts a lightweight Conformer structure as the basic model, which can be expressed as a mapping function: C(S_A, S_B) → (T_A, T_B, s_c), where S_A and S_B are the input Mel spectrogram features, T_A and T_B are the recognized text contents, and s_c is the content similarity score.
[0080] The processing flow of the text content recognition module includes the following steps: Speech feature serialization: Enhance the temporal information of the input Mel spectrogram features through positional encoding E = S + RPE(S), where RPE is the relative positional encoding function; Feature extraction: Extract deep features through the Conformer layer H = Conformer(E) = MHAtt(Conv(E)) + FFN(E); Text decoding: Use the CTC decoder to output the text sequence T = CTC_decode(H); Content similarity calculation: Based on a hybrid metric of edit distance and semantic similarity .
[0081] To improve the practicality of the module, the present invention also introduces several key optimization mechanisms. First is the streaming processing architecture, which achieves low-latency online recognition by restricting the calculation range of self-attention. Secondly is the multi-task learning framework, which adds auxiliary tasks such as language type recognition and speaking style classification in addition to the basic speech recognition task, improving the model's feature extraction ability. Finally is the adaptive fusion strategy, which dynamically adjusts the weight of the recognized text in the final decision according to its credibility.
[0082] In terms of the training strategy, the present invention adopts a three-stage method: First, pre-train the model on a large-scale general speech dataset to establish basic speech recognition capabilities; then fine-tune it on the data of the target scenario to adapt to a specific application environment; finally, continuously optimize the model parameters through online learning to improve the adaptability of the model.
[0083] It can be understood that the present invention realizes content recognition and comparison based on ASR, adopts a multi-task learning framework of CTC loss and attention mechanism, and can utilize the advantages of both decoding methods simultaneously. The text content recognition module not only provides accurate content recognition results, but also significantly improves the overall performance of the model through content similarity information. Especially when processing short speech segments, the content similarity information can effectively guide the model to select the most suitable feature comparison strategy.
[0084] As Figure 6 shown, on the basis of the above embodiments, as an optional embodiment, the speaker vector recognition module includes a time-delay neural network sub-module, a statistical pooling sub-module, and a feature vector extraction and judgment sub-module.
[0085] The time-delay neural network sub-module is used to extract features of different time scales of the audio pair to be recognized from the spectral features of the audio pair to be recognized.
[0086] The statistical pooling sub-module is used to calculate the statistical features of the audio pair to be recognized based on the mean and variance according to the features of different time scales of the audio pair to be recognized.
[0087] The feature vector extraction and judgment sub-module is used to map the statistical features of the audio pair to be recognized to a low-dimensional discriminant space, obtain the speaker vector of the audio pair to be recognized, and calculate the speaker vector similarity to obtain the speaker vector similarity score.
[0088] Each speaker has unique voice characteristics, which can be encoded as a compact and discriminative vector representation, similar to human fingerprint recognition. However, different from fingerprint recognition, the voice characteristics of speakers are affected by various factors, such as emotional state, health condition, surrounding environment, etc., which makes it particularly challenging to extract stable and reliable features.
[0089] When designing the speaker vector recognition module, the present invention focuses on three key issues: First is how to extract identity features that are irrelevant to the speech content. A speaker says different things in different scenarios, but their identity features should remain stable. The second issue is temporal modeling. The features of a speaker are not only reflected in the acoustic features at a single time point but also in the pronunciation habits and intonation changes over a longer time span. The third issue is the discriminability of features. The extracted vectors should be able to maximize the differences between different speakers while minimizing the differences between different speech segments of the same speaker.
[0090] Based on these considerations, the present invention adopts an improved TDNN (Time Delay Neural Network) structure as the basic framework. The uniqueness of TDNN lies in its ability to effectively model context dependencies at different time scales, which is very similar to the process of humans recognizing speakers.
[0091] The speaker vector recognition module can be described as a mapping function: V(S_A, S_B) → s_v, where S is the input sequence of Mel spectrogram features, and xe is the extracted speaker vector. Its processing flow includes the following steps: Hierarchical feature extraction: Extract features at different time scales through a multi-layer TDNN structure H_l = TDNN_l(H_{l - 1}, d_l), where d_l represents the time delay parameter of the l-th layer.
[0092] Statistical pooling: Integrate frame-level features to obtain a global representation μ = E[H_L], σ = √(E[(H_L - μ)²]) G = [μ; σ].
[0093] Projection mapping: Map the statistical features to a low-dimensional discriminant space to obtain the speaker vector xe = normalize(W * G + b).
[0094] Speaker vector similarity calculation: Substitute two speaker vectors into the cosine similarity calculation formula to obtain the speaker vector similarity score s_v.
[0095] It can be understood that the present invention provides x-vector feature extraction for the case of consistent content, uses the X-VECTOR structure to extract speaker vectors, the network contains multiple delay layers, and each layer captures context information at different time scales. After the statistical pooling layer, the variable-length frame-level features are mapped to fixed-dimensional speaker vectors through a multi-layer fully connected network. The training of the speaker vector recognition module adopts an improved AAM-Softmax loss function to improve the discriminability of features by increasing the angular margin.
[0096] As Figure 7As shown, based on the above embodiments, as an alternative embodiment, the existence of the voice cloning comparison module provides a powerful fallback mechanism for the model, especially playing a key role when dealing with speech segments with significant content differences.
[0097] When designing this module, three main challenges are faced. First is the issue of cloning quality. Voice cloning not only needs to restore the timbre characteristics of the speaker but also maintain a natural and fluent speech rhythm. Second is the problem of computational efficiency. Traditional voice cloning models usually have a large computational amount and are difficult to meet the requirements of real-time processing. Finally, it is how to ensure that the cloning process does not introduce additional noise and distortion, affecting the final comparison result.
[0098] To address these challenges, the present invention designs a voice cloning framework based on the improved CosyVoice2 architecture. The unique feature is the introduction of an adaptive layer normalization mechanism, which can better maintain the personalized characteristics of the speaker. In formal expression, this module can be described as: K(A, B) → s_k, where A and B are the input audio pairs, and s_k is the similarity score based on the cloned audio.
[0099] The core processing flow of the voice cloning comparison module includes three main stages: Voice feature extraction stage: Extract the prosody features and speaker features of the source audio, and use the self-attention mechanism to capture long-term dependencies; Audio generation stage: Generate an acoustic feature sequence based on the target text, and inject the speaker features through adaptive layer normalization; Similarity calculation stage: Generate cloned audios with the same content and calculate the similarity between the cloned audio pairs.
[0100] In the implementation process, short text segments of 16 - 32 phonemes are used as the standard text, which can not only ensure the cloning effect but also control the computational cost, achieving a balance between cloning quality and efficiency.
[0101] Practice has proved that the voice cloning comparison module performs excellently when dealing with speech segments with significant content differences. Especially in scenarios where traditional methods are difficult to handle, such as when the speaker uses different languages or dialects, by generating cloned audios with the same content for comparison, the recognition accuracy can be significantly improved.
[0102] It can be understood that the present invention uses an improved CosyVoice2 model. By introducing mechanisms such as adaptive layer normalization and dynamic convolution, the quality and efficiency of voice cloning are improved. The training of the cloning model uses a multi-speaker dataset, and an adversarial training strategy is used to improve the generalization ability of the model.
[0103] The present invention is comprehensively tested on multiple real-scene data sets, including efficiency tests, accuracy tests, and robustness tests.
[0104] Efficiency test: In a business test set of 100,000 audio pairs, the average processing time was reduced from 780ms of the traditional method to 290ms. Among them, 43.5% of the audio pairs completed the judgment in the waveform comparison stage, with an average time of only 85ms.
[0105] Accuracy test: On the business test set with content difference >50%, the accuracy rate reached 92.3%, exceeding the 73.8% of the baseline model.
[0106] Robustness test: On the service test set with various environmental noises added, the average accuracy remained at 91.2%. On the 2-second short audio service test set, the accuracy reached 88.7%, an increase of 15.3% over the baseline model.
[0107] In summary, the present invention adopts a multi-level cascade architecture combining fast paths and precise paths. Through direct waveform comparison as the first fast screening, conclusions can be drawn directly for highly similar audio, avoiding the computational overhead of complete feature extraction in traditional methods, and effectively improving computational efficiency. Actual measurements show that in actual application scenarios, about 40-50% of the comparison tasks can be completed in the waveform comparison stage, which reduces the computational time by more than 65% compared with traditional methods. The present invention innovatively uses content recognition results as the basis for path selection, adopts different feature extraction strategies for different degrees of content difference, and solves the comparison problem in scenarios with large content differences through sound cloning technology. In scenarios with content differences exceeding 50%, the recognition accuracy is improved by about 25%, and the overall equal error rate (EER) is reduced by about 40% compared with traditional voiceprint recognition methods. The multi-level collaborative processing mechanism significantly improves the robustness in complex environments. In a noisy environment with a signal-to-noise ratio of 0-15dB, the recognition accuracy can still be maintained above 90%. It is insensitive to audio length and can simultaneously process audio clips from 2 seconds to several minutes.
[0108] The voiceprint recognition device provided by the present invention is described below. The voiceprint recognition device described below and the voiceprint recognition method described above can be referenced to each other.
[0109] Figure 8 is a timing diagram of the voiceprint recognition device provided by the present invention, such as Figure 8 As shown, the present invention also provides a voiceprint recognition device, including the following modules.
[0110] An acquisition module 810 is used to acquire an audio pair to be identified; An identification module 820, configured to input the audio pair to be identified into a pre-trained multi-level cascaded voiceprint recognition model, and obtain a voiceprint recognition result output by the multi-level cascaded voiceprint recognition model; Wherein, the multi-level cascaded voiceprint recognition model includes an input layer, a waveform comparison layer, an acoustic feature comparison layer, and a decision layer; The input layer is configured to input the audio pair to be identified into the waveform comparison layer and the acoustic feature comparison layer respectively; The waveform comparison layer is configured to compare the waveform similarity of the audio pair to be identified, obtain a waveform similarity score, and output the waveform similarity score to the decision layer; The acoustic feature comparison layer is configured to compare the acoustic feature similarity of the audio pair to be identified, obtain an acoustic feature similarity score, and output the acoustic feature similarity score to the decision layer. The acoustic feature similarity score includes at least one of a spectrum feature similarity score, a text content similarity score, a speaker vector similarity score, and a cloned audio similarity score; The decision layer is configured to determine whether the waveform similarity score reaches a preset value. If so, use the waveform similarity score as the voiceprint recognition result. Otherwise, obtain the acoustic feature similarity score and use the acoustic feature similarity score as the voiceprint recognition result.
[0111] As an embodiment, the acoustic feature comparison layer includes a spectrum feature analysis module, a text content recognition module, a speaker vector recognition module, and a voice cloning comparison module; The spectrum feature analysis module is configured to extract the spectrum features of the audio pair to be identified to determine the spectrum feature similarity score, and transmit the spectrum features of the audio pair to be identified to the text content recognition module and the speaker vector recognition module; The text content recognition module is configured to determine whether the text content of the audio pair to be identified is consistent according to the spectrum features of the audio pair to be identified, and obtain the text content similarity score; The speaker vector recognition module is configured to determine whether the speaker vectors of the audio pair to be identified are consistent according to the spectrum features of the audio pair to be identified, and obtain the speaker vector similarity score; The voice cloning comparison module is configured to extract the prosody features and speaker features of the audio pair to be identified, generate a cloned audio pair with consistent text content according to the prosody features, the speaker features, and a preset text, and judge the similarity of the cloned audio pair to obtain the cloned audio similarity score.
[0112] As an embodiment, the spectrum feature analysis module includes a preprocessing and framing sub-module, a short-time Fourier transform sub-module, a Mel spectrum feature extraction sub-module, a Mel spectrum feature calculation sub-module, a dynamic feature extraction sub-module, and a feature normalization sub-module that are sequentially arranged.
[0113] As an embodiment, the text content recognition module includes an encoder, a decoder, and a content similarity calculation sub-module; The encoder is used to extract the deep features of the audio pair to be recognized from the spectrum features of the audio pair to be recognized based on relative position encoding and an improved self-attention mechanism; The decoder is used to determine the text content of the audio pair to be recognized according to the deep features of the audio pair to be recognized; The content similarity calculation sub-module is used to calculate the similarity of the text content of the audio pair to be recognized based on a mixed metric of edit distance and semantic similarity to obtain the text content similarity score.
[0114] As an embodiment, the speaker vector recognition module includes a time-delay neural network sub-module, a statistical pooling sub-module, and a feature vector extraction and judgment sub-module; The time-delay neural network sub-module is used to extract the features of different time scales of the audio pair to be recognized from the spectrum features of the audio pair to be recognized; The statistical pooling sub-module is used to calculate the statistical features of the audio pair to be recognized based on the mean and variance according to the features of different time scales of the audio pair to be recognized; The feature vector extraction and judgment sub-module is used to map the statistical features of the audio pair to be recognized to a low-dimensional discriminant space, obtain the speaker vector of the audio pair to be recognized, and calculate the speaker vector similarity to obtain the speaker vector similarity score.
[0115] As an embodiment, the waveform comparison layer is implemented based on a deep siamese network structure, and the deep siamese network structure is used to extract the time-domain features of the audio pair to be recognized based on a multi-layer one-dimensional convolutional network, perform weighted processing on the time-domain features of the audio pair to be recognized based on an attention mechanism, and calculate the similarity of the weighted time-domain features of the audio pair to be recognized based on the embedding space obtained by metric learning to obtain the waveform similarity score.
[0116] It should be noted that the voiceprint recognition device provided by the present invention can execute the video ringback tone interaction method described in any of the above embodiments during specific operation, and has the corresponding technical effects of the method, which will not be elaborated in this embodiment.
[0117] Figure 9 Illustrates a schematic diagram of the physical structure of an electronic device, such asFigure 9 As shown, the electronic device may include: a processor 910, a communications interface 920, a memory 930, and a communication bus 940. Among them, the processor 910, the communications interface 920, and the memory 930 complete communication with each other through the communication bus 940. The processor 910 may call the logical instructions in the memory 930 to execute a voiceprint recognition method, which includes: obtaining an audio pair to be recognized; inputting the audio pair to be recognized into a pre-trained multi-level cascaded voiceprint recognition model to obtain a voiceprint recognition result output by the multi-level cascaded voiceprint recognition model; Among them, the multi-level cascaded voiceprint recognition model includes an input layer, a waveform comparison layer, an acoustic feature comparison layer, and a decision layer; The input layer is used to input the audio pair to be recognized into the waveform comparison layer and the acoustic feature comparison layer respectively; The waveform comparison layer is used to compare the waveform similarity of the audio pair to be recognized to obtain a waveform similarity score, and output the waveform similarity score to the decision layer; The acoustic feature comparison layer is used to compare the acoustic feature similarity of the audio pair to be recognized to obtain an acoustic feature similarity score, and output the acoustic feature similarity score to the decision layer. The acoustic feature similarity score includes at least one of a spectral feature similarity score, a text content similarity score, a speaker vector similarity score, and a cloned audio similarity score; The decision layer is used to determine whether the waveform similarity score reaches a preset value. If so, use the waveform similarity score as the voiceprint recognition result. Otherwise, obtain the acoustic feature similarity score and use the acoustic feature similarity score as the voiceprint recognition result.
[0118] In addition, when the logical instructions in the above-mentioned memory 930 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0119] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the voiceprint recognition method provided by each of the above methods, and the method includes: obtaining a pair of audio to be recognized; inputting the pair of audio to be recognized into a pre-trained multi-level cascaded voiceprint recognition model to obtain a voiceprint recognition result output by the multi-level cascaded voiceprint recognition model; wherein, the multi-level cascaded voiceprint recognition model includes an input layer, a waveform comparison layer, an acoustic feature comparison layer, and a decision layer; The input layer is configured to input the pair of audio to be recognized into the waveform comparison layer and the acoustic feature comparison layer respectively; The waveform comparison layer is configured to compare the waveform similarity of the pair of audio to be recognized to obtain a waveform similarity score, and output the waveform similarity score to the decision layer; The acoustic feature comparison layer is configured to compare the acoustic feature similarity of the pair of audio to be recognized to obtain an acoustic feature similarity score, and output the acoustic feature similarity score to the decision layer. The acoustic feature similarity score includes at least one of a spectral feature similarity score, a text content similarity score, a speaker vector similarity score, and a cloned audio similarity score; The decision layer is configured to determine whether the waveform similarity score reaches a preset value. If so, use the waveform similarity score as the voiceprint recognition result. Otherwise, obtain the acoustic feature similarity score and use the acoustic feature similarity score as the voiceprint recognition result.
[0120] In yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the voiceprint recognition method provided by each of the above methods, and the method includes: obtaining a pair of audio to be recognized; inputting the pair of audio to be recognized into a pre-trained multi-level cascaded voiceprint recognition model to obtain a voiceprint recognition result output by the multi-level cascaded voiceprint recognition model; wherein, the multi-level cascaded voiceprint recognition model includes an input layer, a waveform comparison layer, an acoustic feature comparison layer, and a decision layer; The input layer is configured to input the pair of audio to be recognized into the waveform comparison layer and the acoustic feature comparison layer respectively; The waveform comparison layer is configured to compare the waveform similarity of the pair of audio to be recognized to obtain a waveform similarity score, and output the waveform similarity score to the decision layer; The acoustic feature comparison layer is used to compare the acoustic feature similarity of the audio pair to be recognized, obtain an acoustic feature similarity score, and output the acoustic feature similarity score to the decision layer. The acoustic feature similarity score includes at least one of a spectral feature similarity score, a text content similarity score, a speaker vector similarity score, and a cloned audio similarity score; The decision layer is used to determine whether the waveform similarity score reaches a preset value. If so, the waveform similarity score is used as the voiceprint recognition result. Otherwise, the acoustic feature similarity score is obtained and used as the voiceprint recognition result.
[0121] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.
[0122] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, also by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A voiceprint recognition method, characterized in that: include: Get the audio pair to be recognized; Inputting the audio pair to be recognized into a pre-trained multi-level cascade voiceprint recognition model to obtain a voiceprint recognition result output by the multi-level cascade voiceprint recognition model; The multi-level cascade voiceprint recognition model includes an input layer, a waveform comparison layer, an acoustic feature comparison layer and a decision layer; The input layer is used to input the audio pair to be recognized to the waveform comparison layer and the acoustic feature comparison layer respectively; The waveform comparison layer is used to compare the waveform similarities of the audio pairs to be identified, obtain waveform similarity scores, and output the waveform similarity scores to the decision layer; The acoustic feature comparison layer is used to respond to the instruction of the decision layer, compare the acoustic feature similarity of the audio pair to be identified, obtain an acoustic feature similarity score, and output the acoustic feature similarity score to the decision layer, wherein the acoustic feature similarity score includes at least one of a spectral feature similarity score, a text content similarity score, a speaker vector similarity score, and a cloned audio similarity score; The decision layer is used to determine whether the waveform similarity score reaches a preset value. If so, the waveform similarity score is used as the voiceprint recognition result. Otherwise, the acoustic feature similarity score is obtained and the acoustic feature similarity score is used as the voiceprint recognition result.
2. The voiceprint recognition method according to claim 1, characterized in that: The acoustic feature comparison layer includes a spectrum feature analysis module, a text content recognition module, a speaker vector recognition module and a sound clone comparison module; The spectrum feature analysis module is used to extract the spectrum features of the audio pair to be identified to determine the spectrum feature similarity score, and transmit the spectrum features of the audio pair to be identified to the text content recognition module and the speaker vector recognition module; The text content recognition module is used to determine whether the text content of the audio pair to be recognized is consistent according to the frequency spectrum characteristics of the audio pair to be recognized, and obtain the text content similarity score; The speaker vector recognition module is used to determine whether the speaker vectors of the audio pair to be recognized are consistent according to the frequency spectrum characteristics of the audio pair to be recognized, and obtain the speaker vector similarity score; The sound cloning comparison module is used to extract the prosodic features and speaker features of the audio pair to be identified, generate a clone audio pair with consistent text content based on the prosodic features, the speaker features and a preset text, determine the similarity of the clone audio pair, and obtain the clone audio similarity score.
3. The voiceprint recognition method according to claim 2, characterized in that: The spectrum feature analysis module includes a preprocessing and framing submodule, a short-time Fourier transform submodule, a Mel spectrum feature extraction submodule, a Mel spectrum feature calculation submodule, a dynamic feature extraction submodule and a feature normalization submodule which are arranged in sequence.
4. The voiceprint recognition method according to claim 2, characterized in that: The text content recognition module includes an encoder, a decoder and a content similarity calculation submodule; The encoder is used to extract deep features of the audio pair to be identified from the spectrum features of the audio pair to be identified based on relative position encoding and an improved self-attention mechanism; The decoder is used to determine the text content of the audio pair to be identified according to the deep features of the audio pair to be identified; The content similarity calculation submodule is used to calculate the similarity of the text content of the audio pair to be recognized based on a hybrid metric of edit distance and semantic similarity to obtain the text content similarity score.
5. The voiceprint recognition method according to claim 2, characterized in that: The speaker vector recognition module includes a time-delay neural network submodule, a statistical pooling submodule, and a feature vector extraction and judgment submodule; The time-delay neural network submodule is used to extract the features of different time scales of the audio pair to be identified from the frequency spectrum features of the audio pair to be identified; The statistical pooling submodule is used to calculate the statistical features of the audio pair to be identified based on the mean and variance according to the features of different time scales of the audio pair to be identified; The feature vector extraction and judgment submodule is used to map the statistical features of the audio pair to be identified to a low-dimensional discriminant space, obtain the speaker vector of the audio pair to be identified and calculate the speaker vector similarity to obtain the speaker vector similarity score.
6. The voiceprint recognition method according to any one of claims 1 to 5, characterized in that: The waveform comparison layer is implemented based on a deep twin network structure, which is used to extract the time domain features of the audio pair to be identified based on a multi-layer one-dimensional convolutional network, perform weighted processing on the time domain features of the audio pair to be identified based on an attention mechanism, and calculate the similarity of the time domain features of the audio pair to be identified after weighted processing based on the embedding space obtained by metric learning to obtain the waveform similarity score.
7. A voiceprint recognition device, characterized in that: include: An acquisition module, used to acquire an audio pair to be recognized; A recognition module, used for inputting the audio pair to be recognized into a pre-trained multi-level cascade voiceprint recognition model to obtain a voiceprint recognition result output by the multi-level cascade voiceprint recognition model; The multi-level cascade voiceprint recognition model includes an input layer, a waveform comparison layer, an acoustic feature comparison layer and a decision layer; The input layer is used to input the audio pair to be recognized to the waveform comparison layer and the acoustic feature comparison layer respectively; The waveform comparison layer is used to compare the waveform similarities of the audio pairs to be identified, obtain waveform similarity scores, and output the waveform similarity scores to the decision layer; The acoustic feature comparison layer is used to compare the acoustic feature similarities of the audio pairs to be identified, obtain an acoustic feature similarity score, and output the acoustic feature similarity score to the decision layer, wherein the acoustic feature similarity score includes at least one of a spectral feature similarity score, a text content similarity score, a speaker vector similarity score, and a cloned audio similarity score; The decision layer is used to determine whether the waveform similarity score reaches a preset value. If so, the waveform similarity score is used as the voiceprint recognition result. Otherwise, the acoustic feature similarity score is obtained and the acoustic feature similarity score is used as the voiceprint recognition result.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the voiceprint recognition method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the voiceprint recognition method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the voiceprint recognition method according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Edge processing system for real-time voiceprint comparison and event association
CN121237098A