Voiceprint recognition method and system based on deep voice embedding

Through deep voice embedding technology, resampling, paragraph segmentation and data enhancement of voice data, and feature aggregation is performed in combination with attention mechanism, solving the problems of insufficient accuracy and weak noise resistance of traditional voiceprint recognition technology, and achieving efficient identity recognition in complex environments.

CN120581014AActive Publication Date: 2025-09-02GUANGZHOU LANGO ELECTRONICS TECH CO LTD +1

Patent Information

Application Number
CN202510751007.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-02
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

Traditional voiceprint recognition technology has weak noise resistance in complex environments inadequate accuracy, insufficient real-time audio processing capabilities.

Method used

The voiceprint recognition method based on deep speech embedding is adopted, and feature aggregation is performed through speech data resampling, paragraph segmentation, data augmentation and feature extraction, combined with attention mechanism, and feature extraction and matching is performed using deep learning network.

Benefits of technology

It improves the accuracy and robustness of voiceprint recognition, enhances the recognition ability and adaptability in complex environments, and achieves fast and accurate identity recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120581014A_ABST
    Figure CN120581014A_ABST
Patent Text Reader

Abstract

The invention provides a voiceprint recognition method and system based on deep voice embedding, and the method comprises the steps: obtaining voice data, and obtaining buffering waveform data based on the voice data; performing resampling based on the buffer waveform data and a preset sampling rate to obtain sampling voice data; performing paragraph segmentation based on the sampled voice data and preset window time to obtain segmented voice data; performing data enhancement processing based on the segmented voice data to obtain enhanced voice data; performing feature extraction on the enhanced voice data based on a pre-training model to obtain a high-dimensional feature vector; calculating frame-level statistical information based on the high-dimensional feature vector, and obtaining a paragraph feature vector based on the attention mechanism and the frame-level statistical information; and feature vector matching is carried out based on the paragraph feature vectors, so that identity mapping is carried out, and voiceprint recognition is realized. According to the voiceprint recognition method and system based on deep voice embedding provided by the invention, the accuracy of voiceprint recognition is ensured, and the efficiency and reliability of voiceprint recognition are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent processing of speech data, and in particular to a voiceprint recognition method and system based on deep speech embedding. Background Art

[0002] Voiceprint recognition technology is a method of verifying identity by analyzing and identifying individual voice characteristics. Because everyone has different physiological characteristics such as vocal cord structure and oral shape, each person's voice has a unique voiceprint. Voiceprint recognition technology uses these unique biometric characteristics to achieve identity verification.

[0003] However, traditional password and biometric recognition methods have gradually exposed their limitations in performance and timeliness in practical applications, mainly manifesting in the following deficiencies: (1) insufficient voiceprint recognition accuracy; (2) processing of speech features mostly targets offline audio, with insufficient consideration given to real-time audio processing; (3) weak ability to resist interference from environmental noise in complex scenarios. Summary of the Invention

[0004] The present invention aims to provide a voiceprint recognition method and system based on deep speech embedding to solve the above technical problems and effectively improve the recognition ability and accuracy of the voiceprint recognition system in complex environments.

[0005] In order to solve the above technical problems, the present invention provides a voiceprint recognition method based on deep speech embedding, comprising the following steps:

[0006] Acquire voice data, and obtain buffered waveform data based on the voice data;

[0007] Resampling is performed based on the buffered waveform data and a preset sampling rate to obtain sampled voice data;

[0008] Perform segmentation based on the sampled speech data and the preset window time to obtain segmented speech data;

[0009] Perform data enhancement processing based on the segmented speech data to obtain enhanced speech data;

[0010] Perform feature extraction on the enhanced speech data based on the pre-trained model to obtain a high-dimensional feature vector;

[0011] Calculate frame-level statistics based on high-dimensional feature vectors, and obtain paragraph feature vectors based on the attention mechanism and frame-level statistics;

[0012] Feature vector matching is performed based on paragraph feature vectors to perform identity mapping and realize voiceprint recognition.

[0013] The above scheme obtains high-dimensional feature vectors after noise reduction and data enhancement by resampling, paragraph segmentation, data enhancement and feature extraction operations on voice data, which can effectively suppress background noise and enhance the quality of voice data. It can extract fine high-dimensional feature vectors from voice data, thereby improving the accuracy and robustness of subsequent voiceprint recognition, and enhancing the recognition ability and adaptability of voiceprint recognition in complex environments; and based on the attention mechanism, it performs feature aggregation on frame-level statistical information, accurately aggregates frame-level features into paragraph-level representation, and finally performs voice identity mapping, which can quickly and accurately identify the identity of the speaker, ensure the accuracy of voiceprint recognition, and improve the efficiency and reliability of recognition.

[0014] Furthermore, voice data is acquired, and buffered waveform data is obtained based on the voice data, including: acquiring initial voice information, acquiring voice data based on a preset format processing method and the initial voice information; and acquiring buffered waveform data based on a local buffering method and the voice data.

[0015] In the above solution, the voice data is pre-processed by processing it in a preset format and buffering it locally, thereby completing the pre-processing of the voice data. This enables the reception of voice data in different formats and from different sources and subsequent recognition.

[0016] Furthermore, resampling is performed based on the buffered waveform data and the preset sampling rate to obtain sampled voice data. The specific calculation process satisfies the following formula:

[0017]

[0018] Where: x represents the buffered waveform data, y(n) represents the resampled speech data, f out represents the target sampling rate, f in Represents the original sampling rate, and n represents an integer.

[0019] In the above solution, the buffered waveform is resampled by setting a certain sampling rate to ensure that sufficient sound information can be captured in each time interval, while maintaining sufficient audio quality and reducing the complexity of voice data processing and storage.

[0020] Furthermore, segmentation is performed based on the sampled speech data and the preset window time to obtain segmented speech data. The specific calculation process satisfies the following formula:

[0021] L=T×SA

[0022] Where: L represents the window length for segmenting speech data, T represents the preset window time, and SA represents the sampling rate after resampling.

[0023] In the above scheme, the input sound segments are cut into fixed-length sound segments by using a fixed time window, ensuring that each segment contains sufficient voice information while avoiding data redundancy and increased processing complexity caused by overly long segments.

[0024] Furthermore, data enhancement processing is performed based on the segmented speech data to obtain enhanced speech data, including:

[0025] Noise estimation is performed based on the segmented speech data and noise estimation algorithm to obtain the noise spectrum. The specific calculation process satisfies the following formula:

[0026]

[0027] Where X(t,f) represents the segmented speech data, Represents the estimated noise spectrum, EstimateNoise represents the minimum mean square error estimation function;

[0028] Filtering is performed based on the noise spectrum and segmented speech data to obtain filtered speech data. The specific calculation process satisfies the following formula:

[0029]

[0030] Where t represents a moment in the time domain, f represents a frequency in the frequency domain, Y(t,f) represents the filtered speech data, and S(t,f) represents the spectrum of the speech signal. represents the estimated noise spectrum;

[0031] Data enhancement and smoothing are performed based on the filtered speech data to obtain enhanced speech data.

[0032] In the above scheme, noise estimation is performed on the segmented speech data, that is, the characteristics of the background noise are estimated using a noise estimation algorithm, and the spectral characteristics of the background noise are identified by analyzing the silent segments or low-energy segments of the audio signal. Then, the filtering parameters are adaptively adjusted according to the noise spectrum for filtering processing, and the effective components of the signal are retained to obtain enhanced speech data, which effectively suppresses the background noise and enhances the quality of the speech data, thereby improving the accuracy and robustness of subsequent voiceprint recognition.

[0033] Furthermore, feature extraction is performed on the enhanced speech data based on the pre-trained model to obtain a high-dimensional feature vector, including: introducing skip connections in the deep learning network model, and using a large-scale speech data set to pre-train the deep learning network model to obtain a pre-trained model; performing multi-layer convolution and pooling operations based on the pre-trained model and the enhanced speech data to obtain a high-dimensional feature vector.

[0034] In the above scheme, by introducing jump connections in the deep learning network model, the gradient vanishing problem in the deep network can be alleviated, so that deep-level speech features can be extracted more effectively; at the same time, a large-scale speech data set is used to pre-train the deep learning network model, so that the system can learn rich speech features from big data, which helps to improve the accuracy of feature extraction; and the enhanced speech data is input into the pre-training model through multi-layer convolution and pooling operations to extract fine high-dimensional speech feature vectors, that is, high-dimensional feature vectors containing key information of the speech signal, which is helpful for subsequent voiceprint recognition.

[0035] Furthermore, skip connections are introduced into the deep learning network model, and a large-scale speech dataset is used to pre-train the deep learning network model to obtain a pre-trained model, including: randomly deforming, cropping or rotating the speech data based on the large-scale speech dataset to obtain a complex speech dataset in a complex noise environment; introducing skip connections into the deep learning network model, and using a complex speech dataset to pre-train the deep learning network model to obtain a pre-trained model.

[0036] In the above scheme, by randomly deforming, cropping or rotating large-scale speech data sets, the simulation of complex noise environments is achieved on the basis of diversified data, which increases the diversity of data and enables the deep learning network model to learn a wider range of voiceprint features, thereby realizing speech data recognition and processing in complex noise environments, thereby improving the recognition ability and adaptability in complex actual scenarios.

[0037] Furthermore, frame-level statistics are calculated based on the high-dimensional feature vector, and paragraph feature vectors are obtained based on the attention mechanism and frame-level statistics, including:

[0038] Statistical information is obtained based on the statistical pooling function and high-dimensional feature vectors. The statistical information includes the mean and standard deviation. The specific calculation process satisfies the following formula:

[0039]

[0040] Where x i represents the high-dimensional feature vector of the i-th frame, N represents the number of frames, mean(x i ) represents the mean, std(x i ) represents the standard deviation;

[0041] The attention weight of high-dimensional feature vectors is obtained based on deep neural networks. The specific calculation process satisfies the following formula:

[0042]

[0043] Where: W and b represent the preset learning parameters, A(x i) represents the attention weight of the high-dimensional feature vector of the i-th frame, and T represents the period;

[0044] The weighted average feature vector is calculated based on the attention pooling mechanism and attention weight to obtain the paragraph feature vector. The specific calculation process satisfies the following formula:

[0045]

[0046] Where: P(x) represents the paragraph feature vector.

[0047] In the above scheme, the statistical information of the high-order feature vector at the frame level is first calculated through the statistical pooling function, and the attention weight of the high-dimensional feature vector is obtained through deep neural network learning. Finally, the attention mechanism is adopted to learn the weights of different frame-level features, perform weighted averaging on the feature vectors, emphasize important features, ignore irrelevant noise, and through the combination of statistical pooling function and attention mechanism, the frame-level features are more accurately aggregated to obtain the paragraph feature vector, capturing the important information in the speech signal.

[0048] Furthermore, feature vector matching is performed based on the paragraph feature vector, thereby performing identity mapping and realizing voiceprint recognition, including:

[0049] Based on the paragraph feature vector and the feature vector database, feature vector matching is performed to obtain matching similarity. The specific calculation process satisfies the following formula:

[0050]

[0051] Where x represents the paragraph feature vector, y represents the feature vector in the feature vector database, and similarity(x,y) represents the matching similarity;

[0052] Matching mapping is performed based on preset matching principles and matching similarity to obtain the corresponding identity mapping label and realize voiceprint recognition.

[0053] In the above scheme, by calculating the similarity of the feature vectors and mapping them to the corresponding speaker's voiceprint label, voiceprint recognition is achieved, and the speaker's identity is identified quickly and accurately, thereby improving the efficiency and reliability of the system.

[0054] The above scheme obtains high-dimensional feature vectors after noise reduction and data enhancement by resampling, paragraph segmentation, data enhancement and feature extraction operations on voice data, which can effectively suppress background noise and enhance the quality of voice data, so that fine high-dimensional feature vectors can be extracted from the voice data; and large-scale voice data is used for processing to simulate more complex voice data to train deep learning network models, enhance noise resistance and adaptability, so that it can maintain a high recognition rate in complex environments, thereby improving the accuracy and robustness of subsequent voiceprint recognition, and enhancing the recognition ability and adaptability of voiceprint recognition in complex environments; and based on the attention mechanism, frame-level statistical information is feature aggregated, and frame-level features are accurately aggregated to paragraph-level representation to finally perform voice identity mapping, which can quickly and accurately identify the identity of the speaker, ensure the accuracy of voiceprint recognition, and improve recognition efficiency and reliability.

[0055] The present invention also provides a voiceprint recognition system based on deep speech embedding, comprising:

[0056] A voice acquisition module, used to acquire voice data and obtain buffered waveform data based on the voice data;

[0057] A data sampling module is used to resample the buffered waveform data based on a preset sampling rate to obtain sampled voice data;

[0058] A paragraph segmentation module is used to perform paragraph segmentation based on the sampled speech data and a preset window time to obtain segmented speech data;

[0059] A data enhancement module is used to perform data enhancement processing based on the segmented speech data to obtain enhanced speech data;

[0060] A feature extraction module is used to extract features from the enhanced speech data based on the pre-trained model to obtain a high-dimensional feature vector;

[0061] Feature aggregation module, which is used to calculate frame-level statistics based on high-dimensional feature vectors and obtain paragraph feature vectors based on the attention mechanism and frame-level statistics;

[0062] The label mapping module is used to perform feature vector matching based on paragraph feature vectors, thereby performing identity mapping and realizing voiceprint recognition.

[0063] The above scheme provides a voiceprint recognition system, which resamples, segments, enhances and extracts features on voice data through a voice acquisition module, a data sampling module, a paragraph segmentation module, a data enhancement module and a feature extraction module, and obtains a high-dimensional feature vector after noise reduction and data enhancement. It can effectively suppress background noise and enhance the quality of voice data. It can extract fine high-dimensional feature vectors from voice data, thereby improving the accuracy and robustness of subsequent voiceprint recognition and enhancing the recognition ability and adaptability of voiceprint recognition in complex environments. The feature aggregation module adopts an attention mechanism to perform feature aggregation on frame-level statistical information, accurately aggregates frame-level features to paragraph-level representation, and finally performs voice identity mapping through a label mapping module, which can quickly and accurately identify the identity of the speaker, ensure the accuracy of voiceprint recognition, and improve the efficiency and reliability of recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 A schematic diagram of a voiceprint recognition method based on deep speech embedding provided by one embodiment of the present invention;

[0065] Figure 2 A schematic diagram of a voiceprint recognition system based on deep speech embedding provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0066] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0067] Example 1:

[0068] This embodiment provides a voiceprint recognition method based on deep speech embedding, including the following steps:

[0069] S1: Acquire voice data and obtain buffered waveform data based on the voice data;

[0070] S2: resampling based on the buffered waveform data and the preset sampling rate to obtain sampled voice data;

[0071] S3: Segment the sampled speech data based on the preset window time to obtain segmented speech data;

[0072] S4: Perform data enhancement processing based on the segmented speech data to obtain enhanced speech data;

[0073] S5: Extract features from the enhanced speech data based on the pre-trained model to obtain a high-dimensional feature vector;

[0074] S6: Calculate frame-level statistics based on the high-dimensional feature vector, and obtain the paragraph feature vector based on the attention mechanism and frame-level statistics;

[0075] S7: Perform feature vector matching based on the paragraph feature vector to perform identity mapping and realize voiceprint recognition.

[0076] The above scheme obtains high-dimensional feature vectors after noise reduction and data enhancement by resampling, paragraph segmentation, data enhancement and feature extraction operations on voice data, which can effectively suppress background noise and enhance the quality of voice data. It can extract fine high-dimensional feature vectors from voice data, thereby improving the accuracy and robustness of subsequent voiceprint recognition, and enhancing the recognition ability and adaptability of voiceprint recognition in complex environments; and based on the attention mechanism, it performs feature aggregation on frame-level statistical information, accurately aggregates frame-level features into paragraph-level representation, and finally performs voice identity mapping, which can quickly and accurately identify the identity of the speaker, ensure the accuracy of voiceprint recognition, and improve the efficiency and reliability of recognition.

[0077] Optionally, step S1 includes: obtaining initial voice information, obtaining voice data based on a preset format processing method and the initial voice information; and obtaining buffered waveform data based on a local buffering method and the voice data.

[0078] During the specific implementation process, by processing the voice data according to the preset format and locally buffering the voice data, it flexibly supports data testing and completes the preprocessing process of the voice data, so that voice data of different formats and sources can be received and subsequently recognized.

[0079] Optionally, the specific calculation process of step S2 satisfies the following formula:

[0080]

[0081] Where: x represents the buffered waveform data, y(n) represents the resampled speech data, f out represents the target sampling rate, f in Represents the original sampling rate, and n represents an integer.

[0082] During the specific implementation process, the buffered waveform is resampled by setting a certain sampling rate to ensure that sufficient sound information can be captured in each time interval and that the sampling rate of all input audio is consistent with the target voiceprint sampling rate stored in the database. The resampling sampling rate is set to 16000Hz. This sampling rate can reduce the complexity of voice data processing and storage while maintaining sufficient audio quality.

[0083] Optionally, the specific calculation process of step S3 satisfies the following formula:

[0084] L=T×SA

[0085] Where: L represents the window length for segmenting speech data, T represents the preset window time, and SA represents the sampling rate after resampling.

[0086] During the specific implementation process, the input sound segments are cut into fixed time windows (such as 25 milliseconds) to form sound segments of fixed length, ensuring that each segment contains sufficient voice information while avoiding data redundancy and increased processing complexity caused by overly long segments.

[0087] Optionally, step S4 includes:

[0088] Noise estimation is performed based on the segmented speech data and noise estimation algorithm to obtain the noise spectrum. The specific calculation process satisfies the following formula:

[0089]

[0090] Where X(t,f) represents the segmented speech data, Represents the estimated noise spectrum, EstimateNoise represents the minimum mean square error estimation function;

[0091] Filtering is performed based on the noise spectrum and segmented speech data to obtain filtered speech data. The specific calculation process satisfies the following formula:

[0092]

[0093] Where t represents a moment in the time domain, f represents a frequency in the frequency domain, Y(t,f) represents the filtered speech data, and S(t,f) represents the spectrum of the speech signal. represents the estimated noise spectrum;

[0094] Data enhancement and smoothing are performed based on the filtered speech data to obtain enhanced speech data.

[0095] In the specific implementation process, the segmented voice data is processed by a filter to suppress background noise and enhance the voice signal. First, the segmented voice data is subjected to noise estimation, that is, the characteristics of the background noise are estimated using a noise estimation algorithm, and the spectral characteristics of the background noise are identified by analyzing the silent segments or low-energy segments of the audio signal. Then, a noise suppression filter (such as a Wiener filter or an adaptive filter) is applied to process the audio signal. The filter adaptively adjusts the filter parameters according to the estimated noise characteristics to suppress background noise while retaining the effective components of the voice signal. Finally, the filtered signal is converted back to the time domain through the inverse transform of the time and frequency domains, and smoothed to enhance the clarity and recognizability of the voice signal to obtain enhanced voice data, which effectively suppresses background noise and enhances the quality of the voice data, thereby improving the accuracy and robustness of subsequent voiceprint recognition.

[0096] Optionally, step S5 includes: introducing skip connections into the deep learning network model, and pre-training the deep learning network model using a large-scale speech data set to obtain a pre-trained model; performing multi-layer convolution and pooling operations based on the pre-trained model and enhanced speech data to obtain a high-dimensional feature vector.

[0097] In the specific implementation process, the voice data is feature extracted through the pre-trained deep learning network model to generate a high-dimensional feature vector. Specifically, the residual network (ResNet) can be used to perform the following steps: (1) Use the ResNet pre-training model: Introducing jump connections in the residual network can alleviate the gradient vanishing problem in the deep network, so that it has more advantages in processing complex voice data and can better capture the subtle features in the voice signal; at the same time, a large-scale voice data set is selected to pre-train the deep learning network model, so that the system can learn rich voice features from big data, which helps to improve the accuracy of feature extraction; (2) The enhanced voice data is input into the pre-training model and subjected to multi-layer convolution and pooling operations to extract a fine high-dimensional voice feature vector, that is, a high-dimensional feature vector containing the key information of the voice signal, which is helpful for subsequent voiceprint recognition.

[0098] Optionally, skip connections are introduced into the deep learning network model, and a large-scale speech dataset is used to pre-train the deep learning network model to obtain a pre-trained model, including: randomly deforming, cropping or rotating the speech data based on the large-scale speech dataset to obtain a complex speech dataset in a complex noise environment; introducing skip connections into the deep learning network model, and using a complex speech dataset to pre-train the deep learning network model to obtain a pre-trained model.

[0099] During the specific implementation process, by randomly deforming, cropping or rotating large-scale speech data sets, the simulation of complex noise environments is achieved on the basis of diversified data, which increases the diversity of data and enables the deep learning network model to learn a wider range of voiceprint features, thereby realizing speech data recognition and processing in complex noise environments, thereby improving the recognition ability and adaptability in complex actual scenarios.

[0100] Optionally, step S6 includes:

[0101] Statistical information is obtained based on the statistical pooling function and high-dimensional feature vectors. The statistical information includes the mean and standard deviation. The specific calculation process satisfies the following formula:

[0102]

[0103]

[0104] Where x i represents the high-dimensional feature vector of the i-th frame, N represents the number of frames, mean(x i ) represents the mean, std(x i ) represents the standard deviation;

[0105] The attention weight of high-dimensional feature vectors is obtained based on deep neural networks. The specific calculation process satisfies the following formula:

[0106]

[0107] Where: W and b represent the preset learning parameters, A(x i ) represents the attention weight of the high-dimensional feature vector of the i-th frame, and T represents the period;

[0108] The weighted average feature vector is calculated based on the attention pooling mechanism and attention weight to obtain the paragraph feature vector. The specific calculation process satisfies the following formula:

[0109]

[0110] Where: P(x) represents the paragraph feature vector.

[0111] In the specific implementation process, a pooling function based on statistics and attention mechanism is used to aggregate frame-level features into paragraph-level representation. The statistical pooling function calculates the statistical information of each feature (such as the mean and standard deviation), while the attention mechanism emphasizes important features by learning weights and ignores irrelevant noise. The specific steps are as follows: (1) Statistical pooling: Calculate the statistical information of the feature vector at the frame level, including the mean and standard deviation. (2) Attention pooling: Use the attention mechanism to perform weighted averaging on the feature vector by learning the weights of different frame-level features. The attention weights are learned through deep neural network. For a given series of frame-level features {x1, x2, ..., x N}Calculate the attention weight A(x) for each frame i ), and calculate the weighted average feature vector based on these weights. By combining statistics and attention pooling functions, frame-level features can be more accurately aggregated into paragraph-level representations, capturing important information in the speech signal.

[0112] Optionally, step S7 includes:

[0113] Based on the paragraph feature vector and the feature vector database, feature vector matching is performed to obtain matching similarity. The specific calculation process satisfies the following formula:

[0114]

[0115] Where x represents the paragraph feature vector, y represents the feature vector in the feature vector database, and similarity(x,y) represents the matching similarity;

[0116] Matching mapping is performed based on preset matching principles and matching similarity to obtain the corresponding identity mapping label and realize voiceprint recognition.

[0117] During the specific implementation process, feature vector matching is performed by calculating the similarity of feature vectors, and the paragraph-level feature vectors are compared with the speaker label feature vectors in the database. The speaker identity is identified by calculating the similarity between the vectors; and the feature vector with the highest similarity is mapped to the corresponding speaker label to realize voiceprint recognition, which enables the speaker's identity to be identified quickly and accurately, improving the efficiency and reliability of the system.

[0118] In the specific implementation process, compared with the conventional mapping and matching of preset tag libraries, this embodiment achieves more accurate feature vector matching: (1) Frame-level feature aggregation: Before performing feature vector matching, the frame-level features are first aggregated through a pooling function based on statistics and attention mechanisms to obtain a paragraph-level feature vector that can more accurately reflect the important information of the speech signal. Compared with the conventional method of matching using only a single feature vector, the aggregated feature vector contains richer information, making the matching more accurate. (2) Paragraph-level audio feature vector similarity calculation: A specific similarity calculation formula is used to calculate the similarity of paragraph-level audio feature vectors, taking into account the inner product and modulus of the vector, which can more scientifically measure the similarity between two paragraph-level audio feature vectors. Compared with the conventional simple distance calculation or other similarity calculation methods, this formula can better adapt to the characteristics of voiceprint recognition and improve the accuracy of recognition. It also achieves more efficient identity mapping. Based on the principle of maximum similarity, this embodiment performs identity mapping based on the feature vector with the highest similarity, which can quickly and accurately determine the identity of the speaker. When processing large amounts of data and complex scenarios, this method can efficiently screen out the most likely speaker labels and improve the response speed and efficiency of the system.

[0119] The above scheme obtains high-dimensional feature vectors after noise reduction and data enhancement by resampling, paragraph segmentation, data enhancement and feature extraction operations on voice data, which can effectively suppress background noise and enhance the quality of voice data, so that fine high-dimensional feature vectors can be extracted from the voice data; and uses large-scale voice data for processing and simulation of more complex voice data to train deep learning network models, enhance noise resistance and adaptability, so that it can maintain a high recognition rate in complex environments, thereby improving the accuracy and robustness of subsequent voiceprint recognition, and enhancing the recognition ability and adaptability of voiceprint recognition in complex environments; and based on the attention mechanism, it performs feature aggregation on frame-level statistical information, accurately aggregates frame-level features into paragraph-level representation, and finally performs voice identity mapping, which can quickly and accurately identify the identity of the speaker, ensure the accuracy of voiceprint recognition, and greatly improve the recognition ability and adaptability of the system in complex environments.

[0120] This embodiment also provides a voiceprint recognition system based on deep speech embedding, including:

[0121] A voice acquisition module, used to acquire voice data and obtain buffered waveform data based on the voice data;

[0122] A data sampling module is used to resample the buffered waveform data based on a preset sampling rate to obtain sampled voice data;

[0123] A paragraph segmentation module is used to perform paragraph segmentation based on the sampled speech data and a preset window time to obtain segmented speech data;

[0124] A data enhancement module is used to perform data enhancement processing based on the segmented speech data to obtain enhanced speech data;

[0125] A feature extraction module is used to extract features from the enhanced speech data based on the pre-trained model to obtain a high-dimensional feature vector;

[0126] Feature aggregation module, which is used to calculate frame-level statistics based on high-dimensional feature vectors and obtain paragraph feature vectors based on the attention mechanism and frame-level statistics;

[0127] The label mapping module is used to perform feature vector matching based on paragraph feature vectors, thereby performing identity mapping and realizing voiceprint recognition.

[0128] The above scheme provides a voiceprint recognition system, which resamples, segments, enhances and extracts features on voice data through a voice acquisition module, a data sampling module, a paragraph segmentation module, a data enhancement module and a feature extraction module, and obtains a high-dimensional feature vector after noise reduction and data enhancement. It can effectively suppress background noise and enhance the quality of voice data. It can extract fine high-dimensional feature vectors from voice data, thereby improving the accuracy and robustness of subsequent voiceprint recognition and enhancing the recognition ability and adaptability of voiceprint recognition in complex environments. The feature aggregation module adopts an attention mechanism to perform feature aggregation on frame-level statistical information, accurately aggregates frame-level features to paragraph-level representation, and finally performs voice identity mapping through a label mapping module, which can quickly and accurately identify the identity of the speaker, ensure the accuracy of voiceprint recognition, and improve the efficiency and reliability of recognition.

[0129] During the specific implementation process, a universal data input interface can be constructed in the voice acquisition module to receive voice data in different formats and from different sources, including real-time recording input, pre-recorded audio file upload, and streaming audio input. Whether it is pre-recorded audio read from a local storage device or audio collected in real time through a microphone, it can be easily input into the voiceprint recognition system of this embodiment for test recognition, and real-time audio processing can be achieved; and a buffer area is established locally to load the original waveform data of the received voice data into the buffer for buffering to obtain buffered waveform data.

[0130] In the specific implementation process, it also includes continuous monitoring and optimization of system performance, regular updating and retraining of models to cope with the ever-changing voice environment and new sound input, ensuring the long-term stable and efficient operation of the system, and continuously improving the recognition ability and adaptability of the system by introducing new data and improving the model structure. Specifically, the system will be optimized from the following four aspects: (1) Diversified data collection, actively collecting voice data in different scenarios, including various noise environments (such as factories, streets, restaurants, etc.), voices with different accents and language habits, and voices of speakers of different age groups and genders. This will enable the model to learn a wider range of voiceprint features and improve its adaptability in complex real-world scenarios. (2) Data enhancement technology expansion, in addition to existing data enhancement methods such as filtering processing, time and frequency domain transformation, etc., explore new data enhancement technologies. For example, add different types of synthetic noise to the original voice data to simulate a more complex noise environment; perform random deformation of the voice signal, such as stretching, compression, distortion, etc., to increase the diversity of the data. Combine data enhancement methods in other fields, such as the idea of ​​flipping, rotating, cropping, etc. in the image field, and apply them to voice data. For example, random cropping or rotation operations are performed on the speech spectrum, which can be achieved by randomly adjusting the phase of the spectrum. (3) Multi-model fusion, combining different types of deep learning models, such as convolutional neural networks (CNN), recurrent neural networks (RNN), long short-term memory networks (LSTM), etc., and fusing them with the existing residual network (ResNet). Combining the different advantages of different types of models, CNN is good at extracting local features, RNN and LSTM are good at processing sequence data, and by fusion, the advantages of various models are fully utilized to improve the accuracy of voiceprint recognition. Design a hybrid model structure to fuse the intermediate layer outputs of different models, or use a layered fusion method to gradually fuse the feature representations of different models. For example, CNN can be used to extract the local spectral features of speech, and then input them into LSTM to process time series information, and finally fused with the deep features extracted by ResNet. (4) Attention mechanism optimization, further improve the existing attention mechanism to make it pay more attention to the key information in the speech. For example, introduce a multi-scale attention mechanism to pay attention to important features at different time scales and frequency scales of speech at the same time; or design an adaptive attention mechanism to automatically adjust the distribution of attention weights according to the noise level and complexity of the input speech. The attention mechanism is more closely integrated with other modules, such as feature extraction modules and pooling modules for joint optimization. For example, during feature extraction, the attention mechanism guides more refined extraction of important speech features; during pooling, the attention mechanism is used to dynamically adjust the size and position of the pooling area.

[0131] This embodiment proposes a voiceprint recognition system based on deep speech embedding. This system extracts high-dimensional features through deep speech embedding technology and combines statistical pooling with attention pooling techniques to achieve highly accurate voiceprint recognition. This method not only improves recognition performance but also supports online feature extraction, enabling the system to process audio data in real time, enhancing the user experience and making it suitable for a variety of application scenarios such as smart assistants and real-time monitoring. Furthermore, data augmentation technology enhances noise immunity and adaptability, enabling it to maintain high recognition rates in complex environments. The system design features adaptive learning capabilities for continuous performance optimization and supports various voiceprint data inputs. It supports online feature preparation, performing speech segmentation, data augmentation, and feature extraction in an online manner. It also supports statistical and attention-based pooling functions, aggregating frame-level features into segment-level representations and ultimately mapping them to speaker labels, enabling accurate recognition of individual voiceprints and increasing flexibility. Finally, the system continuously monitors and optimizes its performance, ensuring long-term stable and efficient operation. Therefore, the present invention has broad application prospects in areas such as financial security, smart homes, mobile device unlocking, and online identity authentication, providing users with a convenient and secure identity authentication solution.

[0132] In practical applications, deep learning-based voiceprint recognition technology extracts high-dimensional features of speech, greatly improving the system's recognition ability and adaptability in complex environments: deep neural networks can capture more speech details and complex features, significantly improving the accuracy of voiceprint recognition; deep learning models can maintain high recognition rates under complex conditions such as noisy environments and voice changes. With the help of efficient deep learning algorithms, voiceprint recognition systems can process audio data in real time. Real-time audio processing can enhance user experience and can also be widely used in scenarios such as smart assistants and real-time monitoring to ensure that the system can respond and recognize quickly; deep speech embedding technology has strong noise and interference resistance capabilities. By adding various environmental noise and interference factors during the training process, the model can operate stably in a variety of complex scenarios, ensuring the reliability of recognition results; deep learning models have adaptive learning capabilities and can continuously optimize and improve their own recognition performance. Through continuous learning and training, the system can adapt to different users, environments, and devices, further improving recognition results; deep learning-based voiceprint recognition systems can be integrated with other biometric recognition technologies (such as facial recognition and fingerprint recognition) to achieve multimodal authentication and provide higher security and reliability.

[0133] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A voiceprint recognition method based on deep speech embedding, characterized in that: The following steps are involved: Acquire voice data, and obtain buffered waveform data based on the voice data; Resampling is performed based on the buffered waveform data and a preset sampling rate to obtain sampled voice data; Perform segmentation based on the sampled speech data and the preset window time to obtain segmented speech data; Perform data enhancement processing based on the segmented speech data to obtain enhanced speech data; Perform feature extraction on the enhanced speech data based on the pre-trained model to obtain a high-dimensional feature vector; Calculate frame-level statistics based on high-dimensional feature vectors, and obtain paragraph feature vectors based on the attention mechanism and frame-level statistics; Feature vector matching is performed based on paragraph feature vectors to perform identity mapping and realize voiceprint recognition.

2. The voiceprint recognition method based on deep speech embedding according to claim 1, characterized in that: The acquiring of voice data and obtaining buffered waveform data based on the voice data includes: Acquire initial voice information, and acquire voice data based on a preset format processing method and the initial voice information; Buffered waveform data is obtained based on a local buffering method and voice data.

3. The voiceprint recognition method based on deep speech embedding according to claim 1, characterized in that: The resampling is performed based on the buffered waveform data and the preset sampling rate to obtain the sampled voice data. The specific calculation process satisfies the following formula: Where: x represents the buffered waveform data, y(n) represents the resampled speech data, f out represents the target sampling rate, f in represents the original sampling rate, and n represents an integer.

4. The voiceprint recognition method based on deep speech embedding according to claim 1, characterized in that: The segmentation is performed based on the sampled speech data and the preset window time to obtain segmented speech data. The specific calculation process satisfies the following formula: L=T×SA Wherein: L represents the window length of the segmented speech data, T represents the preset window time, and SA represents the sampling rate after resampling.

5. The voiceprint recognition method based on deep speech embedding according to claim 1, characterized in that: The step of performing data enhancement processing based on the segmented speech data to obtain enhanced speech data includes: Noise estimation is performed based on the segmented speech data and noise estimation algorithm to obtain the noise spectrum. The specific calculation process satisfies the following formula: Where X(t,f) represents the segmented speech data, Represents the estimated noise spectrum, EstimateNoise represents the minimum mean square error estimation function; Filtering is performed based on the noise spectrum and segmented speech data to obtain filtered speech data. The specific calculation process satisfies the following formula: Where t represents a moment in the time domain, f represents a frequency in the frequency domain, Y(t,f) represents the filtered speech data, and S(t,f) represents the spectrum of the speech signal. represents the estimated noise spectrum; Data enhancement and smoothing are performed based on the filtered speech data to obtain enhanced speech data.

6. The voiceprint recognition method based on deep speech embedding according to claim 1, characterized in that: The method of extracting features from the enhanced speech data based on the pre-trained model to obtain a high-dimensional feature vector includes: Introducing skip connections into the deep learning network model and pre-training the deep learning network model using a large-scale speech dataset to obtain a pre-trained model; Perform multi-layer convolution and pooling operations based on the pre-trained model and enhanced speech data to obtain high-dimensional feature vectors.

7. The voiceprint recognition method based on deep speech embedding according to claim 6, characterized in that: The method introduces skip connections into the deep learning network model and uses a large-scale speech dataset to pre-train the deep learning network model to obtain a pre-trained model, including: Randomly deform, crop, or rotate speech data based on large-scale speech datasets to obtain complex speech datasets in complex noisy environments; Skip connections are introduced into the deep learning network model, and a complex speech dataset is used to pre-train the deep learning network model to obtain a pre-trained model.

8. The voiceprint recognition method based on deep speech embedding according to claim 1, characterized in that: The calculation of frame-level statistical information based on the high-dimensional feature vector and obtaining a paragraph feature vector based on the attention mechanism and the frame-level statistical information include: Statistical information is obtained based on the statistical pooling function and the high-dimensional feature vector. The statistical information includes the mean and standard deviation. The specific calculation process satisfies the following formula: Where x i represents the high-dimensional feature vector of the i-th frame, N represents the number of frames, mean(x i ) represents the mean, std(x i ) represents the standard deviation; The attention weight of the high-dimensional feature vector is obtained based on the deep neural network. The specific calculation process satisfies the following formula: Where: W and b represent the preset learning parameters, A(x i ) represents the attention weight of the high-dimensional feature vector of the i-th frame, and T represents the period; The weighted average feature vector is calculated based on the attention pooling mechanism and attention weight to obtain the paragraph feature vector. The specific calculation process satisfies the following formula: Where: P(x) represents the paragraph feature vector.

9. The voiceprint recognition method based on deep speech embedding according to claim 1, characterized in that: The feature vector matching based on the paragraph feature vector is performed to perform identity mapping and realize voiceprint recognition, including: Based on the paragraph feature vector and the feature vector database, feature vector matching is performed to obtain matching similarity. The specific calculation process satisfies the following formula: In the formula, x represents the paragraph feature vector, y represents the feature vector in the feature vector database, similarity(x,y) indicates matching similarity; Matching mapping is performed based on preset matching principles and matching similarity to obtain the corresponding identity mapping label and realize voiceprint recognition.

10. A voiceprint recognition system based on deep speech embedding, characterized in that: include: A voice acquisition module, used to acquire voice data and obtain buffered waveform data based on the voice data; A data sampling module is used to resample the buffered waveform data based on a preset sampling rate to obtain sampled voice data; A paragraph segmentation module is used to perform paragraph segmentation based on the sampled speech data and a preset window time to obtain segmented speech data; A data enhancement module is used to perform data enhancement processing based on the segmented speech data to obtain enhanced speech data; A feature extraction module is used to extract features from the enhanced speech data based on the pre-trained model to obtain a high-dimensional feature vector; Feature aggregation module, which is used to calculate frame-level statistics based on high-dimensional feature vectors and obtain paragraph feature vectors based on the attention mechanism and frame-level statistics; The label mapping module is used to perform feature vector matching based on paragraph feature vectors, thereby performing identity mapping and realizing voiceprint recognition.

Citation Information

Patent Citations

  • Voiceprint recognition method and device, electronic equipment and computer readable storage medium

    CN116825114A

  • Voiceprint recognition method and system based on unsteady audio enhancement and multi-scale attention

    CN116863944A

  • Voiceprint recognition method

    WO2023070874A1

Cited By

  • Speaker confirmation method based on SASFV aggregation model

    CN120766685A

  • Voiceprint registration method and device based on deep voice embedding

    CN121662052A