Voiceprint recognition method and device, electronic equipment and storage medium
By simulating encoding and decoding distortion on benchmark speech samples, distorted speech samples of different channel types are generated and a voiceprint recognition model is trained. This solves the problems of feature mismatch and decreased noise immunity in cross-channel voiceprint recognition, and improves the accuracy and stability of recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING YUANJIAN INFORMATION TECH CO LTD
- Filing Date
- 2026-02-13
- Publication Date
- 2026-05-08
AI Technical Summary
In cross-channel voiceprint recognition, existing technologies suffer from problems such as decreased model noise resistance and failure of discrimination thresholds. In particular, feature mismatch is severe under different recording devices, codecs, sampling rates or acoustic environments, affecting recognition performance.
By simulating encoding and decoding distortion of benchmark speech samples, and using frequency domain spectral feature reconstruction and noise mixing or LPC domain parameter perturbation, distorted speech samples of different channel types are generated. Based on these distorted speech samples, a speaker recognition model is trained to enhance the diversity of training data.
It effectively improves the performance of cross-channel voiceprint recognition, solves the feature mismatch problem, maintains the stability of model noise resistance, avoids the failure of the discrimination threshold, and improves the accuracy and robustness of recognition.
Smart Images

Figure CN121999786A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voiceprint recognition, and more specifically, to a voiceprint recognition method, apparatus, electronic device, and storage medium. Background Technology
[0002] Voiceprint recognition is a biometric identification technology that identifies a speaker by analyzing the unique acoustic features of their voice. Deep learning-based approaches are currently the mainstream technology, but there is still room for improvement in its cross-channel recognition performance.
[0003] Cross-channel speaker recognition refers to the task of speaker identification when the registered and verified voices originate from different recording devices, codecs, sampling rates, or acoustic environments. In this scenario, different channels will produce varying degrees and types of nonlinear distortions to the audio signal, which not only undermines the stability of the speaker model but also causes a mismatch between inference features and training features in the deep learning framework, leading to a decrease in the model's noise resistance and the failure of the discrimination threshold. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a voiceprint recognition method, device, electronic device and storage medium that can achieve data enhancement by simulating encoding and decoding distortion of reference speech samples, effectively improve cross-channel voiceprint recognition performance, solve feature mismatch problem, maintain the stability of model noise resistance and avoid discrimination threshold.
[0005] In a first aspect, embodiments of this application provide a voiceprint recognition method, the method comprising: Obtain at least one reference speech sample corresponding to each channel type; For each reference speech sample, the frequency domain spectral feature reconstruction method and the noise mixing method are used to simulate the encoding and decoding distortion of each reference speech sample to obtain the first distorted speech sample of the corresponding channel type; and / or, the LPC domain parameter perturbation method and the signal reconstruction method are used to simulate the encoding and decoding distortion of each reference speech sample to obtain the second distorted speech sample of the corresponding channel type. The voiceprint recognition model is trained based on all first-distortion speech samples and / or all second-distortion speech samples to perform voiceprint recognition.
[0006] In one possible implementation, the step of using a frequency domain spectral feature reconstruction method and a noise mixing method to simulate encoding and decoding distortion for each reference speech sample to obtain a first distorted speech sample corresponding to the channel type includes: The reference speech sample is subjected to a short-time Fourier transform to obtain the amplitude spectrum and phase spectrum; The amplitude spectrum is then normalized by energy. Generate a random phase matrix with the same size as the phase spectrum; The signal is reconstructed based on the normalized amplitude spectrum and the random phase matrix to obtain the analog noise signal; The reference speech sample and the analog noise signal are mixed to obtain a first distorted speech sample of the channel type corresponding to the reference speech sample.
[0007] In one possible implementation, the energy normalization of the amplitude spectrum includes: Based on the amplitude spectrum, calculate the energy spectrum corresponding to the reference speech sample; The amplitude spectrum is normalized based on the energy spectrum after exponential compression.
[0008] In one possible implementation, normalizing the amplitude spectrum based on the exponentially compressed energy spectrum includes: The amplitude spectrum is normalized by substituting the energy spectrum into the following formula; ; in, The normalized amplitude spectrum of the t-th speech frame in the reference speech sample. Let be the amplitude spectrum of the t-th speech frame in the reference speech sample before normalization. Let be the energy value of the t-th speech frame in the reference speech sample. This is the preset exponential compression factor.
[0009] In one possible implementation, mixing the reference speech sample and the analog noise signal to obtain a first distorted speech sample of a channel type corresponding to the reference speech sample includes: Substituting the reference speech sample and the analog noise signal into the following formula, a first distorted speech sample of the channel type corresponding to the reference speech sample is obtained; ; in, This is the first distorted speech sample. As a baseline speech sample, This is the noise scaling factor. To simulate noise signals.
[0010] In one possible implementation, the step of using LPC domain parameter perturbation and signal reconstruction methods to simulate encoding and decoding distortion for each reference speech sample to obtain a second distorted speech sample corresponding to the channel type includes: For each speech frame in the reference speech sample, calculate the initial LPC domain parameters corresponding to the speech frame; Noise is added to the LPC domain parameters corresponding to the speech frame to obtain the target LPC domain parameters; Signal reconstruction is performed based on the target LPC domain parameters corresponding to each speech frame to obtain the second distorted speech sample of the corresponding channel type.
[0011] Secondly, embodiments of this application also provide a voiceprint recognition device, the device comprising: The acquisition module is used to acquire at least one reference speech sample corresponding to each channel type; The encoding / decoding distortion simulation module is used to simulate the encoding / decoding distortion of each reference speech sample by using frequency domain spectral feature reconstruction and noise mixing methods respectively, to obtain the first distorted speech sample of the corresponding channel type; and / or, to simulate the encoding / decoding distortion of each reference speech sample by using LPC domain parameter perturbation and signal reconstruction methods respectively, to obtain the second distorted speech sample of the corresponding channel type. The training module is used to train the voiceprint recognition model based on all first-distorted speech samples and / or all second-distorted speech samples for voiceprint recognition.
[0012] In one possible implementation, the encoding / decoding distortion simulation module is specifically used to perform a short-time Fourier transform on the reference speech sample to obtain an amplitude spectrum and a phase spectrum; to normalize the energy of the amplitude spectrum; to generate a random phase matrix with the same size as the phase spectrum; to reconstruct the signal based on the normalized amplitude spectrum and the random phase matrix to obtain a simulated noise signal; and to mix the reference speech sample and the simulated noise signal to obtain a first distorted speech sample of the channel type corresponding to the reference speech sample.
[0013] In one possible implementation, the encoding / decoding distortion simulation module is further configured to: Based on the amplitude spectrum, calculate the energy spectrum corresponding to the reference speech sample; The amplitude spectrum is normalized based on the energy spectrum after exponential compression.
[0014] In one possible implementation, the encoding / decoding distortion simulation module is further configured to: The amplitude spectrum is normalized by substituting the energy spectrum into the following formula; ; in, The normalized amplitude spectrum of the t-th speech frame in the reference speech sample. Let be the amplitude spectrum of the t-th speech frame in the reference speech sample before normalization. Let be the energy value of the t-th speech frame in the reference speech sample. This is the preset exponential compression factor.
[0015] In one possible implementation, the encoding / decoding distortion simulation module is further configured to: Substituting the reference speech sample and the analog noise signal into the following formula, a first distorted speech sample of the channel type corresponding to the reference speech sample is obtained; ; in, This is the first distorted speech sample. As a baseline speech sample, This is the noise scaling factor. To simulate noise signals.
[0016] In one possible implementation, the encoding / decoding distortion simulation module is specifically used to calculate the initial LPC domain parameters corresponding to each speech frame in the reference speech sample; add noise to the LPC domain parameters corresponding to the speech frame to obtain the target LPC domain parameters; and reconstruct the signal based on the target LPC domain parameters corresponding to each speech frame to obtain the second distorted speech sample of the corresponding channel type.
[0017] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the voiceprint recognition method as described in any of the first aspects.
[0018] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the voiceprint recognition method as described in any of the first aspects.
[0019] This application provides a voiceprint recognition method, apparatus, electronic device, and storage medium. The method includes: acquiring at least one reference speech sample corresponding to each channel type; performing encoding / decoding distortion simulation on each reference speech sample using a frequency domain spectral feature reconstruction method and a noise mixing method, and / or using an LPC domain parameter perturbation method and a signal reconstruction method; and training a voiceprint recognition model based on all first-distorted speech samples and / or all second-distorted speech samples obtained after the encoding / decoding distortion simulation for voiceprint recognition. This application can achieve data augmentation by performing encoding / decoding distortion simulation on the reference speech samples, effectively improving cross-channel voiceprint recognition performance, solving the feature mismatch problem, maintaining the stability of the model's noise resistance, and avoiding the discrimination threshold. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 A flowchart of deep learning-based voiceprint recognition provided in an embodiment of this application is shown; Figure 2 A flowchart illustrating a voiceprint recognition method provided in an embodiment of this application is shown; Figure 3 The spectrogram of the reference speech sample provided in the embodiment of this application is shown; Figure 4 The spectrogram of the second distorted speech sample provided in an embodiment of this application is shown; Figure 5 This document illustrates a flowchart illustrating encoding / decoding distortion using a frequency domain spectral feature reconstruction method combined with a noise mixing method, as provided in an embodiment of this application. Figure 6 The spectrogram of the first distorted speech sample provided in an embodiment of this application is shown; Figure 7 The flowchart of the encoding / decoding distortion model using LPC domain parameter perturbation and signal reconstruction methods provided in the embodiments of this application is shown. Figure 8 This paper shows a schematic diagram of the structure of a voiceprint recognition device provided in an embodiment of this application; Figure 9 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.
[0023] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0024] To enable those skilled in the art to utilize the content of this application, and in conjunction with the specific application scenario of "voiceprint recognition," the following implementation methods are provided. For those skilled in the art, the general principles defined herein can be applied to other embodiments and application scenarios without departing from the spirit and scope of this application. Although this application primarily describes the "voiceprint recognition field," it should be understood that this is merely an exemplary embodiment.
[0025] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.
[0026] Voiceprint recognition is a biometric identification technology that identifies a speaker by analyzing the unique acoustic features of their voice. This technology is widely used in various fields such as security authentication, telephone banking, smart assistants, voice control devices, and personalized services, providing a convenient and non-intrusive method for user verification and enhancing system security and user experience.
[0027] Deep learning-based voiceprint recognition is currently the mainstream method, referring to... Figure 1 The diagram shows a flowchart of a deep learning-based voiceprint recognition method provided in an embodiment of this application. After data preprocessing, the original speech undergoes feature extraction. The extracted features are then input into a trained voiceprint recognition model to obtain a speaker embedding vector. Similarity scores are assigned to the speaker embedding vectors obtained from two different audio recordings to determine whether they belong to the same person.
[0028] The data preprocessing stage typically includes steps such as Voice Activity Detection (VAD), speech pre-emphasis, DC component removal, audio energy normalization, and audio length segmentation. In the feature extraction stage, depending on the speaker recognition model used, different input methods may be employed, such as time-domain waveform input, short-time Fourier amplitude spectrum input, or filter bank input (Fbank, i.e., Mel-scale logarithmic energy spectrum). Speaker recognition models may include RawNet, Ecapa-TDNN, ResNet, and various improved network forms that have emerged in recent years. The speaker recognition model outputs a speaker embedding vector, which commonly has dimensions of 192, 256, 512, and 1024. For two audio clips to be identified as belonging to the same person, after extracting the speaker embedding vectors for each, a similarity score needs to be calculated. The most common method for similarity scoring is to directly calculate the cosine similarity between two vectors. For certain models and applications, methods such as Probabilistic Linear Discriminant Analysis (PLDA) may also be used for similarity calculation. Ultimately, the similarity result is used to determine whether the speakers in the two audio clips are the same person.
[0029] Deep learning-based voiceprint recognition systems have far surpassed the performance of the previous generation of i-vector-based voiceprint recognition systems, but their cross-channel recognition performance still needs improvement. Cross-channel voiceprint recognition refers to the task of identifying the speaker when the registration voice and verification voice come from different recording devices, codecs, sampling rates, or acoustic environments. Typical scenarios include "mobile phone registration → landline verification" and "high-definition microphone acquisition → VoIP network call verification." In these scenarios, different channels introduce varying degrees and types of nonlinear distortion into the audio signal, compromising the stability of the voiceprint model. In deep learning scenarios, the distortion introduced by different channels can cause a mismatch between inference features and training features, leading to decreased performance of noise models and failure of the discrimination threshold.
[0030] Looking back at the development history of voiceprint recognition, the main solutions to cross-channel voiceprint recognition problems have been constantly evolving with the advancement of voiceprint recognition technology. They can be broadly categorized into four main types: feature normalization, joint factor analysis, i-vector / PLDA, and data augmentation.
[0031] Feature regularization and joint factor analysis attempt to decouple speaker information and channel information during the feature extraction stage, removing interference from channel information and using only speaker information for voiceprint recognition. The I-vector / PLDA method improves the model's adaptability to cross-channel scenarios by adding an additional trained similarity discrimination model in the post-processing stage. Data augmentation methods focus on the training stage, aiming to improve the cross-channel performance of the voiceprint model by enriching the training data as much as possible. Among the above methods, the feature regularization method is similar to the method proposed in the embodiments of this application, and the data augmentation method belongs to the field of data augmentation, that is, improving the model's generalization performance by increasing the diversity of training data. The feature regularization method and the data augmentation method will be briefly introduced below.
[0032] Feature normalization is an early cross-channel solution used in voiceprint recognition. It treats channel differences as linear multiplicative distortions of the original spectrum, smoothing them out through a one-step linear transformation so that subsequent models retain only speaker information. Computationally, it first performs frame-level energy extraction, then calculates the mean energy at each frequency point across the entire speech segment as a channel characteristic estimate. Finally, it removes channel energy from the energy spectrum (e.g., by subtracting the mean from the logarithmic energy spectrum, or dividing the linear energy spectrum by the gain) to eliminate channel differences.
[0033] Data augmentation is a common and effective method in deep learning to improve the generalization performance of models. Its principle is to include as many voice samples as possible from the same speaker across different channels in the training data, allowing the model to learn cross-channel voiceprint feature extraction during training and improving its cross-channel performance. Data augmentation methods can be divided into two types: augmentation using real-world data and augmentation using simulated data. Augmentation using real-world data requires constructing a multi-speaker dataset for voiceprint model training. Each speaker ID in this dataset should have voice data from multiple different channels. Under these conditions, a voiceprint recognition model with high cross-channel performance can be trained. Augmentation using simulated data is known as data enhancement. Mainstream data enhancement methods include adding noise, reverberation, and spectral masking to the speech. Adding reverberation, for example, involves convolving the original audio with a pre-recorded or simulated room impulse response, introducing channel information corresponding to the recorded room impulse response, which can also improve cross-channel performance to some extent. Some researchers have also proposed applying the actual audio codec algorithms used in the communication link during the training process to simulate the distortion during audio signal propagation, thereby improving the adaptability of the voiceprint model to the audio codec algorithm.
[0034] However, traditional data augmentation methods all have certain drawbacks, which are detailed below: (1) Feature regularization method. The problem with this method is that it generally estimates channel characteristics based on only a single speech, resulting in a large estimation error and leakage or loss of speaker-related information. Current deep learning methods commonly use cepstral mean and variance normalization (CMVN) when training speaker models, which is a feature regularization method. Because it requires processing features similar to Fbank, this method is generally used after adding noise and reverberation during training, and is affected by noise and reverberation. Its improvement on cross-channel problems is limited, and current speaker models cannot effectively solve cross-channel problems after using this method.
[0035] (2) Data augmentation through real-time recording, i.e., constructing a multi-speaker, multi-channel training set, is an effective method to solve the cross-channel problem of voiceprint recognition. Its disadvantage lies in the difficulty of implementation. Voiceprint model training requires a large amount of speech signals with speaker annotations. If the same speaker's audio is to be collected in different channels, the difficulty and cost of collection will increase significantly. For the training data of hundreds of thousands or even millions of people required for voiceprint model training, this cost is often unacceptable. At the same time, existing training data often does not contain information about the speaker's real identity, and it is difficult to solve the problem of aligning the same speaker in different channels when supplementing the recording of other channel data, so the feasibility is poor.
[0036] (3) While traditional data augmentation methods can theoretically simulate scenarios with different channels and noise levels, several problems exist. Firstly, there is a lack of open-source resources related to impulse response. Although many datasets exist for room impulse response, most focus on the room's reverberation characteristics, and the acquisition devices are mostly microphone array audio acquisition boards, which are relatively simple and cannot meet the requirements of channel diversity. Secondly, recording high-quality room impulse responses is difficult. Common impulse signal generation methods, such as hitting a wooden board or popping a balloon, produce signals with uneven frequency energy distribution, leading to a shift in the channel simulation of the impulse response. This results in the simulated data differing from the real data, potentially affecting the performance of the trained model in real-world scenarios. Furthermore, even if an impulse response that accurately reflects channel information can be recorded, it cannot simulate channel differences in real-world scenarios. This is because the sound source in a real-world scenario is a natural human voice, while the sound source in the training scenario is a pre-recorded speech signal that has already been acquired by a microphone at least once. The microphone's frequency response characteristics have been added to the speech signal, and simulation based on this, with the superposition of other channel characteristics, results in significant differences from the actual scenario. Finally, while additive noise can simulate different types of noise in the environment, it cannot simulate the distortion introduced by audio encoding and decoding algorithms.
[0037] (4) While data augmentation methods using audio codec algorithms can simulate the distortion introduced by these algorithms during voice signal transmission, they also have limitations. First, in real-world communication environments, there are various audio codec methods, and some algorithms lack open-source Python implementations (such as the EVS method commonly used in 4G / 5G communication), making implementation work quite laborious. Second, practical audio codec methods often involve complex processes to ensure voice transmission quality at low bit rates. Although they can achieve real-time processing in communication scenarios, their processing time remains a bottleneck in neural network training, severely impacting network training speed. Finally, data augmentation using codec methods only considers the distortion introduced by audio codecs during transmission; it cannot simulate channel differences caused by audio acquisition devices and front-end audio processing algorithms, limiting the degree of cross-channel performance improvement.
[0038] In summary, among current related methods, feature regularization aims to decouple and eliminate channel differences in speech signals, rather than improve channel diversity during training; data augmentation methods based on recorded data are effective, but their feasibility is limited by implementation costs; data augmentation methods such as adding reverberation do not aim to increase the channel diversity of data, have limited impact on the channel characteristics of training samples, and differ significantly from actual channel conditions; data augmentation methods based on audio codec algorithms are affected by factors such as implementation difficulties, severe impact on training speed, and inability to compensate for differences in audio acquisition equipment.
[0039] In view of this, embodiments of this application provide a voiceprint recognition method, which includes: acquiring at least one reference speech sample corresponding to each channel type; for each reference speech sample, performing encoding and decoding distortion simulation on each reference speech sample using a frequency domain spectral feature reconstruction method and a noise mixing method respectively, to obtain a first distorted speech sample corresponding to the channel type; and / or, performing encoding and decoding distortion simulation on each reference speech sample using an LPC domain parameter perturbation method and a signal reconstruction method respectively, to obtain a second distorted speech sample corresponding to the channel type; and training a voiceprint recognition model based on all first distorted speech samples and / or all second distorted speech samples to perform voiceprint recognition. This method can effectively simulate speech signals collected by audio recording devices with different frequency response characteristics during the voiceprint model training stage, and uses low-complexity algorithms to simulate the distortion introduced by common audio encoding and decoding algorithms, improving the diversity of training data, and is fully compatible with current mainstream voiceprint models, thereby improving its cross-channel voiceprint recognition performance.
[0040] The following is a detailed description of a voiceprint recognition method provided in the embodiments of this application.
[0041] Reference Figure 2The diagram shown is a flowchart illustrating a voiceprint recognition method provided in an embodiment of this application. The exemplary steps of this embodiment are described below: S201. Obtain at least one reference speech sample corresponding to each channel type.
[0042] In the embodiments of this application, each reference speech sample is a distortion-free / low-distortion speech signal, which can be recorded in a recording scenario of the corresponding channel type, or obtained by channel conversion of reference speech samples of other channel types.
[0043] The channel type is determined based on information such as the equipment used to collect the voice signal, the environment, and the encoding method. It refers to the type of channel. A channel is the entire transmission and transformation path that a voice signal takes from the "sound source" to the "receiving end," as well as the sum of all factors along the path that may alter the signal.
[0044] Here, due to the large number of recording scenarios of different channel types (such as different codecs, recording equipment, sampling rates, transmission links, and acoustic environments), the high deployment and acquisition costs, and the difficulty in large-scale reproduction of some channels (such as specific low-bit-rate codecs and dedicated communication channels) in real-world scenarios, it is impossible to collect a large number of speech samples in each channel type's recording scenario. This leads to unstable performance of the voiceprint recognition model trained solely on recorded speech samples in cross-channel voiceprint recognition scenarios, causing a mismatch between inference features and training features, which in turn leads to a decrease in model noise resistance and failure of the discrimination threshold. Therefore, this application provides a method for channel conversion of reference speech samples to achieve data augmentation, thereby enriching the number of reference speech samples for each channel type. This solves the problem of performance instability of the voiceprint recognition model in cross-channel voiceprint recognition scenarios, causing a mismatch between inference features and training features, ensuring the stability of model noise resistance, and avoiding failure of the discrimination threshold. Specifically, the reference speech samples corresponding to the first channel type are channel converted through the following steps to obtain the reference speech samples corresponding to the second channel type: Step 1: For each frequency point, calculate the channel difference between the channel characteristic value of that frequency point in the second channel characteristics corresponding to the second channel type and the channel characteristic value of that frequency point in the first channel characteristics corresponding to the first channel type.
[0045] In this application's implementation, channel differences are measured by the amplitude gain between corresponding frequency points. A frequency point (or frequency grid) is simply a "frequency grid" or "frequency sampling point" on the frequency spectrum; it is a discrete frequency position obtained by discretizing the continuous frequency axis after a Short-Time Fourier Transform (STFT) / FFT. Channel characteristics refer to the comprehensive characteristics of the transformations and interferences exerted on the speech signal by the transmission medium, processing system, and acoustic environment. Simply put, it's the "pattern of change" the channel produces on the speech signal—it determines what kind of distortion, noise, delay, attenuation, etc., the signal will acquire from the speaker to the acquisition end. Channel characteristics include the channel characteristic values corresponding to each frequency point.
[0046] Assume the channel characteristic value at this frequency point in the first channel characteristic is The channel characteristic value for this frequency point in the second channel characteristic is Taking the channel characteristic estimation result based on Fbank as an example, its physical meaning is logarithmic energy. The channel difference value d is then calculated using the following formula: ; in, It is the natural base.
[0047] Taking the channel characteristic estimation result based on the short-time Fourier energy spectrum as an example, the channel difference value d is calculated using the following formula: .
[0048] Step 2: Calculate the filter based on the channel difference value and frequency of each frequency point.
[0049] In the embodiments of this application, the channel difference value of each frequency point is first used as the amplitude compensation gain corresponding to each frequency point; then, with the frequency of each frequency point as the abscissa and the corresponding amplitude compensation gain as the target response, the linear phase finite length unit impulse response (FIR) filter is designed and calculated using the frequency sampling method.
[0050] The frequency sampling method is a classic finite-length unit impulse response (FIR) filter design approach. Its core idea is to sample the target frequency response at equal intervals in the frequency domain, then perform an inverse Fourier transform (IFFT) on the sampled points to obtain the time-domain impulse response of the filter, thereby designing a linear-phase FIR filter that satisfies the target frequency response. Furthermore, during filter calculation, the fitting accuracy can be controlled by setting the filter order: a higher filter order results in a smaller fitting error between the frequency response and the target compensation gain, but also increases computational complexity and storage overhead. Therefore, in scenarios where computational resources permit, it is recommended to use a higher filter order (e.g., 2048th order) to obtain more accurate channel compensation results.
[0051] Furthermore, this application embodiment also provides a process for calculating the frequency of a frequency point, specifically including: For channel characteristics estimated based on linear frequency scale features such as the short-time Fourier amplitude spectrum, the linear frequency corresponding to the i-th frequency point can be calculated based on the signal sampling rate s and the transform length n used in the short-time Fourier transform. : ; For channel characteristics estimated based on nonlinear frequency-scale features similar to Fbank, it is necessary to first calculate the nonlinear frequency corresponding to its frequency point, and then convert it into the corresponding linear frequency. Taking Fbank features as an example, let... Let be the Mel frequency corresponding to the i-th Mel frequency point, and its conversion formula is as follows: .
[0052] Step 3: Convolve the reference speech sample corresponding to the first channel type with the calculated filter to obtain the reference speech sample corresponding to the second channel type.
[0053] In this embodiment, a convolution is performed between the reference speech sample corresponding to the first channel type and the calculated filter (the formula can be expressed as conv(reference speech sample corresponding to the first channel type, filter impulse response)). This achieves filtering of the reference speech sample corresponding to the first channel type, aligning its frequency response characteristics with the characteristics of the second channel type, thereby eliminating frequency response differences between different channels. Specifically, the frequency response of the filter is the amplitude gain (i.e., channel difference value) calculated in step one. The convolution operation is equivalent to applying a corresponding compensation gain to the energy of each frequency point of the reference speech sample corresponding to the first channel type in the frequency domain, thus canceling out channel differences.
[0054] Furthermore, the channel characteristics corresponding to each channel type are predetermined and stored in a channel characteristic database. Specifically, the channel characteristics corresponding to any channel type are determined according to the following steps: Step 1: Obtain the recorded voice sample corresponding to this channel type.
[0055] When acquiring speech samples, the recording conditions for the speech samples are the recording conditions under this channel type. The number of speakers should be as large as possible, while ensuring a balanced male-female ratio. Optionally, the recorded speech samples can also be preprocessed to improve the quality of the recorded speech samples, such as: (1) detecting the speech activity of the recorded speech samples to remove non-speech segments in the recorded speech samples; (2) standardizing the volume of the recorded speech samples to eliminate the influence of volume differences between speech samples; (3) resampling the recorded speech samples according to actual needs to ensure that the sampling rate of the recorded speech samples is the same as the sampling rate of the input speech features required by the voiceprint recognition model.
[0056] Among them, the voiceprint recognition model is the model used in voiceprint recognition to extract the speaker embedding vector. The speaker embedding vector is a high-dimensional feature vector that can represent the uniqueness of a speaker's identity. This vector can effectively distinguish different speakers and is the core basis for identity comparison and discrimination in voiceprint recognition.
[0057] Step 2: Calculate the energy spectrum of the recorded speech samples.
[0058] In the embodiments of this application, the energy spectrum type and calculation parameters should be consistent with the corresponding parameters of the input speech features required by the voiceprint recognition model. For example, if the voiceprint recognition model uses 80-dimensional Fbank features (Mel-Filter Bank Features) calculated with a frame length of 25ms, a frame shift of 10ms, and a Hamming window as input, then the energy spectrum should also be calculated using a frame length of 25ms, a frame shift of 10ms, and a Hamming window to calculate the 80-dimensional Fbank. Similarly, if the voiceprint recognition model uses the short-time Fourier energy spectrum as input, then only the short-time Fourier energy spectrum needs to be calculated according to the corresponding parameters; the Fbank does not need to be calculated (the calculation of the Fbank requires obtaining the short-time Fourier energy spectrum first).
[0059] The energy spectrum includes the energy values of each speech frame at various frequency points in the recorded speech samples. The energy value of each speech frame is used to characterize the degree to which the energy distribution of the speech frame is affected by noise.
[0060] S302. Based on the energy spectrum, select key speech frames from the recorded speech samples that meet the speech frame screening rules.
[0061] In this application's implementation, the speech frame selection rules are formulated by combining speech frame energy effectiveness and low-frequency channel sensitivity. Speech frame energy effectiveness refers to the degree to which the energy value of a speech frame represents the actual speech signal; the higher the energy value of a speech frame, the less its energy distribution is affected by noise and other interference, and the more accurately it reflects the energy distribution characteristics of the initial speech. Low-frequency channel sensitivity refers to the sensitivity of the low-frequency components of the speech signal to channel differences; the lower the frequency, the more significant the differences in frequency response characteristics between different channels, making it more suitable for characterizing channel characteristics. The highest energy value of the selected key speech frame is higher than that of non-key speech frames, or the highest energy value of the key speech frame in the low-frequency band is higher than that of non-key speech frames.
[0062] For example, the speech frame filtering rules can be: (1) Select the first preset number or preset proportion (e.g., 30%) of speech frames with the highest energy values. (2) Select speech frames with energy values higher than a specific energy threshold. (3) Select the first preset number or preset proportion (e.g., 30%) of speech frames with the highest energy values among the second preset number Y frequency points with the lowest frequencies. (4) Select speech frames with energy values higher than a specific energy threshold among the second preset number Y frequency points with the lowest frequencies.
[0063] Here, the keyframe filtering method provided in this application avoids the mutual interference between unvoiced sounds (energy mainly concentrated in the high-frequency band) and voiced sounds (energy is greater in the low-frequency band, decreasing towards the high-frequency band, but there is also energy distribution in the high-frequency band). By using one of the above rules, the accuracy of channel characteristic estimation can be effectively improved compared to using all speech frames.
[0064] S303. Calculate the channel characteristics corresponding to this channel type based on the energy values of key speech frames in the energy spectrum.
[0065] In the embodiments of this application, for each frequency point, the average linear energy of all key speech frames corresponding to that frequency point is calculated in the linear domain based on the energy spectrum, so as to obtain the channel characteristic value corresponding to that frequency point in the channel characteristics corresponding to that channel type.
[0066] Here, for the short-time Fourier energy spectrum, the energy values in the spectrum are the energy values in the linear domain. For Fbank, the energy values in the energy spectrum represent logarithmic energy values, which need to be restored to linear energy values through exponential operations before averaging.
[0067] Therefore, based on the energy spectrum, the average linear energy of all key speech frames corresponding to the corresponding frequency point is calculated in the linear domain to obtain the channel characteristic value corresponding to that frequency point in the channel characteristics corresponding to that channel type. This includes: if the energy spectrum is a short-time Fourier energy spectrum, calculating the average energy value of all key speech frames corresponding to the corresponding frequency point in the energy spectrum to obtain the channel characteristic value corresponding to that frequency point in the channel characteristics corresponding to that channel type. If the energy spectrum is an Fbank, the energy value of each key speech frame corresponding to the corresponding frequency point in the energy spectrum is exponentially calculated according to a preset exponent value to obtain the linear energy value of each key speech frame corresponding to the corresponding frequency point; the average linear energy value of all key speech frames corresponding to the corresponding frequency point is calculated to obtain the channel characteristic value corresponding to that frequency point in the channel characteristics corresponding to that channel type.
[0068] In addition, experiments have shown that direct averaging of logarithmic energy can also serve the purpose of channel estimation and channel compensation, but its effect is somewhat worse than averaging in the case of linear energy.
[0069] While the aforementioned channel characteristic estimation process requires multiple speaker voice signals of the same channel type, it is relatively easy to collect because it does not require speaker information annotation and can utilize abundant open-source voice data resources. When constructing a channel characteristic database, besides using the above estimation process to estimate from the recorded speech set, there are two alternative solutions. One is to refer to the frequency response parameters of the recording equipment. Different recording equipment generally publishes its own frequency response parameters, which can be used as a representation of the frequency response characteristics of the corresponding channel. The other is to select frequency response characteristics of different channel types and perform weighted fusion to generate channel characteristics not present in the collected recorded speech sample data. This alternative solution is generally used when the number of collected channel frequency response characteristics is small.
[0070] S202. For each reference speech sample, the frequency domain spectral feature reconstruction method and the noise mixing method are used to simulate the encoding and decoding distortion of each reference speech sample to obtain the first distorted speech sample of the corresponding channel type; and / or, the LPC domain parameter perturbation method and the signal reconstruction method are used to simulate the encoding and decoding distortion of each reference speech sample to obtain the second distorted speech sample of the corresponding channel type.
[0071] In the embodiments of this application, the frequency domain spectral feature reconstruction method refers to the process of reconstructing an analog noise signal based on the frequency domain spectral features (including amplitude spectrum and phase spectrum) of a reference speech sample. The noise mixing method refers to the process of mixing the reference speech sample with the reconstructed analog noise signal. The LPC domain parameter perturbation method refers to the process of perturbing (e.g., adding noise) the LPC domain parameters (including LPC coefficients and residual signals) of the reference speech sample. The signal reconstruction method refers to the process of reconstructing the signal based on the perturbed LPC domain parameters.
[0072] LPC stands for Linear Predictive Coding, which is the most core and classic parametric coding method in speech signal processing, especially widely used in low bit-rate speech encoding and decoding (such as telephone, VoIP, and voiceprint feature extraction).
[0073] Specifically, embodiments of this application may employ any of the following methods to simulate encoding / decoding distortion for each reference speech sample: Scenario 1: Using frequency domain spectral feature reconstruction and noise mixing methods, respectively, the encoding and decoding distortion simulation of each reference speech sample is performed to obtain the first distorted speech sample of the corresponding channel type.
[0074] Scenario 2: Using LPC domain parameter perturbation and signal reconstruction methods, respectively, encoding and decoding distortion simulations are performed on each reference speech sample to obtain the second distorted speech sample corresponding to the channel type.
[0075] Scenario 3: The frequency domain spectral feature reconstruction method and the noise mixing method are used to simulate the encoding and decoding distortion of each reference speech sample to obtain the first distorted speech sample of the corresponding channel type; and the LPC domain parameter perturbation method and the signal reconstruction method are used to simulate the encoding and decoding distortion of each reference speech sample to obtain the second distorted speech sample of the corresponding channel type.
[0076] Furthermore, this application proposes a simplified method for simulating encoding and decoding distortion. By employing short-time Fourier transform and its inverse transform, a secondary simulation is performed on the analog quantization error signal obtained using LPC domain parameter perturbation and signal reconstruction methods. This eliminates LPC-related calculation steps, further simplifying the computational complexity. Figure 3 The spectrogram of the reference speech sample provided in the embodiments of this application. Figure 4 The spectrogram of the second distorted speech sample provided in this embodiment of the application (the signal energy strength shown in the two graphs is not comparable). It can be seen that the energy of the second distorted speech sample is manifested as the diffusion of harmonic energy within the reference speech sample and the overlap of harmonic energies. Therefore, referring to... Figure 5 The diagram shows a flowchart of encoding / decoding distortion using a frequency domain spectral feature reconstruction method and a noise mixing method provided in this application embodiment. The simplified encoding / decoding distortion simulation method steps are as follows; S501. Perform a short-time Fourier transform on the reference speech sample to obtain the amplitude spectrum and phase spectrum.
[0077] In this embodiment, the Short-Time Fourier Transform (STFT) divides a continuous time-domain speech signal into many short frames with a fixed frame length and frame shift. A Fourier Transform (FFT) is then performed on each frame to obtain the frequency domain distribution (amplitude spectrum + phase spectrum). The amplitude spectrum, after STFT, represents the signal "intensity / energy" for each frame and each frequency point; it is the modulus of the complex spectrum and only reflects the strength of frequency components, not time delay / phase relationships. The phase spectrum, after STFT, represents the signal "time position / delay" for each frame and each frequency point; it is the argument of the complex spectrum and reflects the time alignment relationship between different frequency components.
[0078] S502, Perform energy normalization on the amplitude spectrum.
[0079] i. Calculate the energy spectrum corresponding to the reference speech sample based on the amplitude spectrum.
[0080] In this embodiment, the amplitude spectrum is substituted into the following formula to obtain the energy spectrum corresponding to the reference speech sample: ; in, Let be the energy value of the t-th speech frame in the reference speech sample. The amplitude spectrum of the i-th frequency point in the t-th speech frame of the reference speech sample before normalization. This represents the number of frequency points.
[0081] ii. Normalize the amplitude spectrum based on the energy spectrum after exponential compression.
[0082] In this embodiment of the application, the energy spectrum is substituted into the following formula to normalize the amplitude spectrum; ; in, The normalized amplitude spectrum of the t-th speech frame in the reference speech sample. Let be the amplitude spectrum of the t-th speech frame in the reference speech sample before normalization. Let be the energy value of the t-th speech frame in the reference speech sample. This is the preset exponential compression factor.
[0083] S503. Generate a random phase matrix with the same size as the phase spectrum.
[0084] In this embodiment of the application, a random phase matrix with the same size as the phase spectrum is generated, and each value in the random phase matrix is a random number between -π and π.
[0085] S504. Reconstruct the signal based on the normalized amplitude spectrum and random phase matrix to obtain the analog noise signal.
[0086] S505. Mix the reference speech sample and the analog noise signal to obtain the first distorted speech sample of the channel type corresponding to the reference speech sample.
[0087] In this embodiment of the application, the reference speech sample and the analog noise signal are substituted into the following formula to obtain a first distorted speech sample of the channel type corresponding to the reference speech sample; ; ; in, This is the first distorted speech sample. As a baseline speech sample, This is the noise scaling factor. To simulate noise signals, The i-th frame signal of the reference speech sample The i-th frame signal is used to simulate noise. The signal-to-noise ratio is randomly specified, in decibels.
[0088] Reference Figure 6 The image shown is a spectrogram of the first distorted speech sample provided in an embodiment of this application.
[0089] Furthermore, a wide variety of encoding and decoding technologies are used in communication systems. Through analysis, we have found that in the fields of mobile communication and network communication, common encoding and decoding methods, such as the AMR method used in 2G / 3G mobile communication, the G.723 and G.729 methods used in early narrowband VoIP and video conferencing, the EVS method used in current 4G / 5G mobile communication, and the Opus method used in real-time Internet voice communication, all base their voice signal encoding and decoding on Linear Predictive Coding (LPC).
[0090] The general process of LPC encoding and decoding is as follows: the audio signal is divided into frames, LPC coefficients and residual signals are calculated for each frame, and then the LPC coefficients and residual signals are transmitted separately using quantization compression. At the decoding end, the obtained LPC coefficients and residual signals are used to reconstruct the speech signal. The main source of error is the quantization error of the LPC coefficients and residual signals. A major difference between the different encoding and decoding methods mentioned above is the different quantization transmission methods for the LPC coefficients and residual signals. This application proposes an LPC-based encoding and decoding distortion simulation method. Specifically, refer to... Figure 7The diagram shown is a flowchart of the encoding / decoding distortion model using LPC domain parameter perturbation and signal reconstruction methods provided in an embodiment of this application. S701. For each speech frame in the reference speech sample, calculate the initial LPC domain parameters corresponding to the speech frame.
[0091] In this embodiment, the autocorrelation is calculated for each speech frame, and the initial LPC coefficients in the initial LPC domain parameters are obtained using the Levinson-Durbin recursive algorithm. The initial residual signal is calculated based on the obtained initial LPC coefficients. : ; in, These are speech frames from the baseline speech samples. For the first in this speech frame One sampling point. For the first in this speech frame Initial residual signal at each sampling point . The order calculated for LPC varies depending on the audio codec algorithm and its bitrate settings. It can be randomly set to an integer between 8 and 16.
[0092] It should be noted that the calculation process of autocorrelation and initial residual signal are standard steps in traditional linear predictive analysis. They are obtained by solving the autocorrelation function and recursively calculating the prediction error from the time-domain sampled sequence of speech frames. The core is to use the short-time stationarity of speech signals to decompose speech into LPC coefficients that represent the vocal tract spectrum envelope and residual signals that represent excitation information, thus laying the foundation for subsequent LPC domain parameter perturbation and signal reconstruction.
[0093] Furthermore, each speech frame in the reference speech sample is obtained by windowing the reference speech sample into frames. Specifically, the reference speech sample is divided into short-time frame signals, typically 20ms in length, corresponding to 160 points / 320 points at 8kHz / 16kHz sampling rates. Optionally, windowing can be applied to the framed signals, typically using common window functions such as the Hanning window or Hamming window. Different frame shift lengths can be set during frame division, such as 1 / 4 frame length, 1 / 2 frame length, or the frame shift equal to the frame length. During data augmentation, these selections can be set according to the settings in the actual encoding / decoding algorithm, or randomly for each training sample.
[0094] S702. Add noise to the LPC domain parameters corresponding to the speech frame to obtain the target LPC domain parameters.
[0095] In this embodiment, the quantization error of the initial LPC coefficients is simulated. Optionally, random white noise is used to simulate the quantization error of the initial LPC coefficients to obtain the target LPC coefficients with quantization error in the target LPC domain parameters. (Supplement: Arrows indicate vectors, and those without arrows indicate scalars. Since LPC coefficients are a set of numbers, they can be understood as...) ...) subscript The coefficient represents a quantization error.
[0096] ; in, It is random white noise, and the speech frame They have the same length; This is the scaling factor. The scaling factor can be obtained according to the calculation method of signal to noise ratio (SNR), and is generally set to 10~25dB: .
[0097] In addition, random white noise is used to simulate the initial residual signal. The quantization error is calculated to obtain the target residual signal Res in the target LPC domain parameters, which contains the quantization error. The process can be referenced from the method for determining the target LPC coefficients.
[0098] S703. Reconstruct the signal based on the target LPC domain parameters corresponding to each speech frame to obtain the second distorted speech sample of the corresponding channel type.
[0099] In this embodiment of the application, the target LPC domain parameters corresponding to each speech frame are substituted into the following formula to obtain the second distorted speech sample of the corresponding channel type.
[0100] ; Among them, when hour . The first speech frame in the second distorted speech sample One sampling point, The first speech frame in the second distorted speech sample Each sampling point. The overlap-add method is applied to the reconstructed audio frame signal, and the complete simulated speech signal with quantization error (i.e., the second distorted speech sample) is restored according to the parameter settings used when windowing the frames.
[0101] S203. Train the voiceprint recognition model based on all first-distorted speech samples and / or all second-distorted speech samples to perform voiceprint recognition.
[0102] In the embodiments of this application, the voiceprint recognition model can be trained using any of the following methods to perform voiceprint recognition: Scenario 1: Train the voiceprint recognition model based on all first-distortion speech samples to perform voiceprint recognition.
[0103] Scenario 2: Train the voiceprint recognition model based on all second-distorted speech samples to perform voiceprint recognition.
[0104] Scenario 3: Train the voiceprint recognition model based on all first-distortion speech samples and all second-distortion speech samples.
[0105] In addition, when using all first-distorted speech samples and / or all second-distorted speech samples as sample data to train the voiceprint recognition model, any traditional model training method can be used, and then voiceprint recognition can be performed using the trained voiceprint recognition model.
[0106] Based on the same inventive concept, this application also provides a voiceprint recognition device corresponding to the voiceprint recognition method. Since the principle of the device in this application is similar to the voiceprint recognition method described above in this application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0107] Reference Figure 8 The diagram shown is a structural schematic of a voiceprint recognition device provided in an embodiment of this application. The device includes: The acquisition module 801 is used to acquire at least one reference speech sample corresponding to each channel type; The encoding / decoding distortion simulation module 802 is used to simulate the encoding / decoding distortion of each reference speech sample by using a frequency domain spectral feature reconstruction method and a noise mixing method to obtain a first distorted speech sample of the corresponding channel type; and / or, to simulate the encoding / decoding distortion of each reference speech sample by using an LPC domain parameter perturbation method and a signal reconstruction method to obtain a second distorted speech sample of the corresponding channel type. Training module 803 is used to train the voiceprint recognition model based on all first-distorted speech samples and / or all second-distorted speech samples for voiceprint recognition.
[0108] The voiceprint recognition device provided in this application embodiment can achieve data enhancement by simulating encoding and decoding distortion of the reference speech sample, effectively improving the cross-channel voiceprint recognition performance, solving the feature mismatch problem, maintaining the stability of the model's noise resistance, and avoiding the discrimination threshold.
[0109] like Figure 9As shown in the embodiment of this application, an electronic device 900 includes a processor 901, a memory 902, and a bus. The memory 902 stores machine-readable instructions executable by the processor 901. When the electronic device is running, the processor 901 communicates with the memory 902 via the bus, and the processor 901 executes the machine-readable instructions to perform the steps of the voiceprint recognition method described above.
[0110] Specifically, the memory 902 and processor 901 mentioned above can be general-purpose memory and processor, without any specific limitations. When the processor 901 runs the computer program stored in the memory 902, it can execute the above-mentioned voiceprint recognition method.
[0111] Corresponding to the above-described voiceprint recognition method, this application embodiment also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the above-described voiceprint recognition method.
[0112] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces; the indirect coupling or communication connection of devices or modules can be electrical, mechanical, or other forms.
[0113] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0114] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0115] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0116] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A voiceprint recognition method, characterized in that, The method includes: Obtain at least one reference speech sample corresponding to each channel type; For each reference speech sample, the frequency domain spectral feature reconstruction method and the noise mixing method are used to simulate the encoding and decoding distortion of each reference speech sample to obtain the first distorted speech sample of the corresponding channel type; and / or, the LPC domain parameter perturbation method and the signal reconstruction method are used to simulate the encoding and decoding distortion of each reference speech sample to obtain the second distorted speech sample of the corresponding channel type. The voiceprint recognition model is trained based on all first-distortion speech samples and / or all second-distortion speech samples to perform voiceprint recognition.
2. The voiceprint recognition method according to claim 1, characterized in that, The method employs both frequency domain spectral feature reconstruction and noise mixing to simulate encoding and decoding distortion for each reference speech sample, resulting in a first distorted speech sample corresponding to the channel type, including: The reference speech sample is subjected to a short-time Fourier transform to obtain the amplitude spectrum and phase spectrum; The amplitude spectrum is then normalized by energy. Generate a random phase matrix with the same size as the phase spectrum; The signal is reconstructed based on the normalized amplitude spectrum and the random phase matrix to obtain the analog noise signal; The reference speech sample and the analog noise signal are mixed to obtain a first distorted speech sample of the channel type corresponding to the reference speech sample.
3. The voiceprint recognition method according to claim 2, characterized in that, The energy normalization of the amplitude spectrum includes: Based on the amplitude spectrum, calculate the energy spectrum corresponding to the reference speech sample; The amplitude spectrum is normalized based on the energy spectrum after exponential compression.
4. The voiceprint recognition method according to claim 3, characterized in that, The normalization of the amplitude spectrum based on the exponentially compressed energy spectrum includes: The amplitude spectrum is normalized by substituting the energy spectrum into the following formula; ; in, The normalized amplitude spectrum of the t-th speech frame in the reference speech sample. Let be the amplitude spectrum of the t-th speech frame in the reference speech sample before normalization. Let be the energy value of the t-th speech frame in the reference speech sample. This is the preset exponential compression factor.
5. The voiceprint recognition method according to claim 2, characterized in that, The step of mixing the reference speech sample and the analog noise signal to obtain a first distorted speech sample of the channel type corresponding to the reference speech sample includes: Substituting the reference speech sample and the analog noise signal into the following formula, a first distorted speech sample of the channel type corresponding to the reference speech sample is obtained; ; in, This is the first distorted speech sample. As a baseline speech sample, This is the noise scaling factor. To simulate noise signals.
6. The voiceprint recognition method according to claim 1, characterized in that, The LPC domain parameter perturbation method and signal reconstruction method are used to simulate encoding and decoding distortion for each reference speech sample to obtain a second distorted speech sample corresponding to the channel type, including: For each speech frame in the reference speech sample, calculate the initial LPC domain parameters corresponding to the speech frame; Noise is added to the LPC domain parameters corresponding to the speech frame to obtain the target LPC domain parameters; Signal reconstruction is performed based on the target LPC domain parameters corresponding to each speech frame to obtain the second distorted speech sample of the corresponding channel type.
7. A voiceprint recognition device, characterized in that, The device includes: The acquisition module is used to acquire at least one reference speech sample corresponding to each channel type; The encoding / decoding distortion simulation module is used to simulate the encoding / decoding distortion of each reference speech sample by using frequency domain spectral feature reconstruction and noise mixing methods respectively, to obtain the first distorted speech sample of the corresponding channel type; and / or, to simulate the encoding / decoding distortion of each reference speech sample by using LPC domain parameter perturbation and signal reconstruction methods respectively, to obtain the second distorted speech sample of the corresponding channel type. The training module is used to train the voiceprint recognition model based on all first-distorted speech samples and / or all second-distorted speech samples for voiceprint recognition.
8. The apparatus according to claim 7, characterized in that, The encoding / decoding distortion simulation module is specifically used for: The reference speech sample is subjected to a short-time Fourier transform to obtain the amplitude spectrum and phase spectrum; The amplitude spectrum is then normalized by energy. Generate a random phase matrix with the same size as the phase spectrum; The signal is reconstructed based on the normalized amplitude spectrum and the random phase matrix to obtain the analog noise signal; The reference speech sample and the analog noise signal are mixed to obtain a first distorted speech sample of the channel type corresponding to the reference speech sample.
9. An electronic device, characterized in that, include: The device includes a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is in operation, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the voiceprint recognition method as described in any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the voiceprint recognition method as described in any one of claims 1 to 6.
Citation Information
Cited By
A voiceprint recognition processing method for scheduling a phone call
CN122177124A