Voiceprint recognition method and device, electronic equipment and storage medium

By using the generative adversarial network to generate synthetic speech samples and adjust the generated sample ratio, the problem of insufficient performance and generalization capabilities of voiceprint recognition technology in low-resource scenarios is solved, and efficient voiceprint recognition effect is achieved.

CN120220693APending Publication Date: 2025-06-27SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510225301.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the case of limited data sets, insufficient diversity and low resource scenarios, the model performance is difficult to meet actual needs, resulting in insufficient generalization capabilities.

Method used

Through a method based on the Generative Adversarial Network (GAN), the generator trains the generator to generate synthetic speech samples, gradually adjusts the proportion of generated samples, iteratively updates the Generative Adversarial Network, and obtains the optimal Generative Sample Scale ratio, which is used to train the Voiceprint Recognition Model.

Benefits of technology

Effectively expand the training data set, improve the performance and generalization capabilities of the model, and ensure that high-precision recognition results can be maintained in low-resource scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220693A_ABST
    Figure CN120220693A_ABST
Patent Text Reader

Abstract

The invention provides a voiceprint recognition method and device, electronic equipment and a storage medium, and belongs to the technical field of voice recognition, and the method comprises the steps: training a generative adversarial network based on a training data set; iteratively updating the generation sample proportion of a generator in the generative adversarial network based on the performance of a discriminator in the generative adversarial network on the verification data set to obtain an optimal generation sample proportion; training a voiceprint recognition model based on the optimal generated sample proportion to obtain a target voiceprint recognition model; and inputting the voice to be recognized into the target voiceprint recognition model to obtain a voiceprint recognition result output by the target voiceprint recognition model. According to the method, the generative adversarial network is applied to generate and synthesize the voice sample, so that the problems of limited data set and insufficient diversity in voiceprint recognition are solved; in the model training process, the number of the generated samples is gradually adjusted, the optimal proportion of the generated samples to the real samples is found, it is ensured that the model keeps high precision, and the good generalization ability is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech recognition technology, and in particular, to a voiceprint recognition method, device, electronic device and storage medium. Background Art

[0002] As an identity authentication method based on voice signals, voiceprint recognition technology has been widely used in fields such as security monitoring, intelligent devices, and financial payment in recent years. Its core task is to identify the speaker's identity by extracting unique features in the voice signal.

[0003] However, despite the significant progress made in voiceprint recognition technology, existing methods still face many challenges in practical applications. Especially in scenarios with limited datasets, insufficient diversity, and low resources, the performance of the model often fails to meet the actual requirements. The performance of the voiceprint recognition model highly depends on the scale and quality of the training data. In practical applications, especially in dialects, minority languages, or specific scenarios, it is very difficult to obtain large-scale and diverse labeled data. Data scarcity not only limits the training effect of the model but also leads to insufficient generalization ability of the model, making it difficult to handle complex voice environments. Although traditional data augmentation methods can expand the dataset to a certain extent, the generated samples often lack diversity and cannot effectively simulate the complex distribution of real voices. In addition, the data distribution in low-resource scenarios usually exhibits long-tail characteristics, that is, the number of samples of certain categories is extremely small, further exacerbating the difficulty of model training and resulting in low model performance.

[0004] Therefore, there is an urgent need for a voiceprint recognition method to solve the problem of low model performance caused by data scarcity in the prior art. Summary of the Invention

[0005] The present invention provides a voiceprint recognition method, device, electronic device and storage medium to solve the defect of low performance of the voiceprint recognition model caused by data scarcity in the prior art.

[0006] The present invention provides a voiceprint recognition method, including the following steps: Based on a training dataset, train a generative adversarial network; Based on the performance of the discriminator in the generative adversarial network on a validation dataset, iteratively update the generation sample ratio of the generator in the generative adversarial network to obtain an optimal generation sample ratio; Based on the optimal generation sample ratio, train a voiceprint recognition model to obtain a target voiceprint recognition model; Input the voice to be recognized into the target voiceprint recognition model to obtain the voiceprint recognition result output by the target voiceprint recognition model.

[0007] A voiceprint recognition method provided by the present invention, iteratively updating the generation sample ratio of a generator in the generative adversarial network according to the performance of a discriminator in the verification dataset to obtain an optimal generation sample ratio, includes: During the training process of the generative adversarial network, gradually reduce the ratio of generated samples; If the performance of the discriminator on the verification dataset decreases, reduce the ratio of generated samples; If the performance of the discriminator on the verification dataset increases, maintain or increase the ratio of generated samples; Iteratively update the ratio of generated samples until the optimal generation sample ratio is obtained.

[0008] A voiceprint recognition method provided by the present invention, the generator includes an input layer, a fully connected layer, a convolutional layer, a LeakyReLU activation function, and a Tanh activation function; The fully connected layer is used to learn the complex representation of the input random noise vector; The convolutional layer is used to extract local dependencies in the speech signal; The Tanh activation function is used to scale the output signal to the range of [-1, 1].

[0009] A voiceprint recognition method provided by the present invention, the discriminator includes a multi-layer convolutional layer, a fully connected layer, and a Sigmoid activation function; The multi-layer convolutional layer is used to extract features from the input data; The fully connected layer is used to generate a classification result based on the extracted features; The Sigmoid activation function is used to compress the output data to the range of (0, 1).

[0010] A voiceprint recognition method provided by the present invention, the target voiceprint recognition model extracts features from the speech to be recognized based on the following steps: Perform frame splitting on the input signal to obtain multiple frames of speech signals; Process each frame of the speech signals in the multiple frames of speech signals using a Hamming window to obtain multiple frames of windowed speech signals; Perform a fast Fourier transform on each frame of the windowed speech signals in the multiple frames of windowed speech signals to obtain multiple frames of short-time spectra; Square each frame of the short-time spectra in the multiple frames of short-time spectra to obtain multiple frames of power spectra; Input the multiple frames of power spectra into a Mel filter bank to obtain a filtered signal output by the Mel filter bank; Take the logarithm of the filtered signal to obtain a logarithmic filtered signal; Perform discrete cosine transform on the logarithmic filter signal to obtain Mel frequency cepstral coefficients; Calculate the first-order difference and second-order difference of the Mel frequency cepstral coefficients to obtain Delta coefficients and Delta-Delta coefficients respectively; Concatenate the Mel frequency cepstral coefficients, the Delta coefficients and the Delta-Delta coefficients to obtain a feature vector.

[0011] According to a voiceprint recognition method provided by the present invention, the target voiceprint recognition model extracts features from the voice to be recognized based on the following steps: Extract features from the voice to be recognized through convolutional networks of different scales connected in sequence to obtain feature maps of different scales. Each convolutional network includes a convolutional layer and a max pooling layer; Use an attention module for each of the feature maps to obtain multiple attention feature maps correspondingly; Perform fusion processing on each of the attention feature maps to obtain a feature vector; Among them, the attention module generates an attention feature map based on the following steps: Perform global average pooling on each feature channel to obtain the global feature representation of each feature channel; Input the global feature representation into a fully connected layer to obtain the importance weight of each feature channel; Update the feature map based on the importance weight to obtain a channel attention result; Compress the weight feature map in the channel dimension to obtain a spatial attention map; Multiply the spatial attention map by the feature map to obtain a spatial attention result; Perform weighted processing on the channel attention result and the spatial attention result to obtain an attention feature map.

[0012] The present invention also provides a voiceprint recognition device, including the following modules: A first training module for training a generative adversarial network based on a training data set; A sample update module for iteratively updating the generation sample ratio of the generator in the generative adversarial network based on the performance of the discriminator in the generative adversarial network on a validation data set to obtain an optimal generation sample ratio; A second training module for training a voiceprint recognition model based on the optimal generation sample ratio to obtain a target voiceprint recognition model; A target recognition module for inputting the voice to be recognized into the target voiceprint recognition model to obtain a voiceprint recognition result output by the target voiceprint recognition model.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the voiceprint recognition method described in any one of the above is implemented.

[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the voiceprint recognition method described in any one of the above is implemented.

[0015] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the voiceprint recognition method described in any one of the above is implemented.

[0016] The voiceprint recognition method, device, electronic device, and storage medium provided by the present invention train a generative adversarial network based on a training data set; iteratively update the generation sample ratio of the generator in the generative adversarial network based on the performance of the discriminator in the generative adversarial network on a validation data set to obtain an optimal generation sample ratio; train a voiceprint recognition model based on the optimal generation sample ratio to obtain a target voiceprint recognition model; input the voice to be recognized into the target voiceprint recognition model to obtain the voiceprint recognition result output by the target voiceprint recognition model. The present invention applies a generative adversarial network to generate synthetic voice samples to solve the problems of limited data set and insufficient diversity in voiceprint recognition; gradually adjusts the number of generated samples during the model training process, and through multiple iterations, finds the optimal ratio of generated samples to real samples, so as to train a voiceprint recognition model with better performance based on sample data, ensuring that the model maintains high accuracy and has good generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is one of the flow diagrams of the voiceprint recognition method provided by the present invention; Figure 2 is another flow diagram of the voiceprint recognition method provided by the present invention; Figure 3 is the structural diagram of the voiceprint recognition device provided by the present invention; Figure 4 is the structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts fall within the protection scope of the present invention.

[0020] It should be noted that in the description of the embodiments of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element. The orientation or positional relationship indicated by terms such as "upper", "lower", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation of the present invention. Unless otherwise expressly specified and defined, the terms "mounted", "connected" and "coupled" shall be construed broadly, for example, it may be a fixed connection, a detachable connection or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0021] The terms "first", "second", etc. in this application are used to distinguish similar objects and are not used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category and do not limit the number of objects. For example, the first object may be one or more. In addition, "and / or" means at least one of the connected objects, and the character " / " generally means that the related objects before and after are in an "or" relationship.

[0022] Figure 1 is one of the flow diagrams of the voiceprint recognition method provided by the present invention, as Figure 1 shown, the method includes the following: S110, training a generative adversarial network based on a training data set; S120. Iteratively update the generation sample ratio of the generator in the generative adversarial network based on the performance of the discriminator in the generative adversarial network on the validation dataset to obtain the optimal generation sample ratio. S130. Train the voiceprint recognition model based on the optimal generation sample ratio to obtain the target voiceprint recognition model. S140. Input the voice to be recognized into the target voiceprint recognition model to obtain the voiceprint recognition result output by the target voiceprint recognition model.

[0023] It should be noted that the execution subject of the task construction method provided in the embodiments of the present application can be a server or a computer device, such as a mobile phone, a tablet computer, a notebook computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc.

[0024] In the embodiments of the present invention, a generative adversarial network (GAN) consists of a generator and a discriminator. The generator generates synthetic voice samples through random noise, and the discriminator is responsible for distinguishing real samples and generated samples. Through adversarial training, the generator gradually generates more realistic voice samples, thereby expanding the training dataset.

[0025] The voiceprint recognition method provided in the embodiments of the present invention trains the generative adversarial network based on the training dataset; iteratively updates the generation sample ratio of the generator in the generative adversarial network based on the performance of the discriminator in the generative adversarial network on the validation dataset to obtain the optimal generation sample ratio; trains the voiceprint recognition model based on the optimal generation sample ratio to obtain the target voiceprint recognition model; and inputs the voice to be recognized into the target voiceprint recognition model to obtain the voiceprint recognition result output by the target voiceprint recognition model. The present invention applies a generative adversarial network to generate synthetic voice samples to solve the problems of limited dataset and insufficient diversity in voiceprint recognition; gradually adjusts the number of generated samples during the model training process, and through multiple iterations, finds the optimal ratio between the generated samples and the real samples, thereby training a voiceprint recognition model with better performance based on the sample data, ensuring that the model maintains high accuracy and has good generalization ability.

[0026] In an alternative embodiment, the iteratively updating the generation sample ratio of the generator in the generative adversarial network based on the performance of the discriminator in the generative adversarial network on the validation dataset to obtain the optimal generation sample ratio includes: During the training of the generative adversarial network, gradually reduce the proportion of generated samples; If the performance of the discriminator on the validation dataset deteriorates, reduce the proportion of generated samples; If the performance of the discriminator on the validation dataset improves, maintain or increase the proportion of generated samples; Iteratively update the proportion of generated samples until the optimal proportion of generated samples is obtained.

[0027] To avoid overfitting caused by excessive generated samples, the present invention proposes a dynamic sample balancing mechanism. This mechanism iteratively reduces the number of generated samples, gradually adjusts the ratio of generated samples to real samples, and finds the optimal sample ratio. Specifically, during the training process of the model, the number of generated samples is gradually reduced, and after each reduction, the model is retrained and the model performance is recorded. Through multiple iterations, the optimal ratio of generated samples to real samples is found to ensure that the model has good generalization ability while maintaining high accuracy.

[0028] In the embodiments of the present invention, during the training process, the number of samples generated by the GAN is gradually reduced, and after each reduction, the model is retrained and the model performance is recorded. Specifically, a larger number of generated samples are used during the initial training, and as the training progresses, the number of generated samples is gradually reduced until the optimal proportion of generated samples is found.

[0029] To further optimize the proportion of generated samples, the embodiments of the present invention introduce a dynamic adjustment mechanism, which dynamically adjusts the number of generated samples by monitoring the performance of the model on the validation set. If the performance of the model on the validation set deteriorates, the number of generated samples is reduced; if the performance improves, the number of generated samples is maintained or increased.

[0030] The voiceprint recognition method provided by the embodiments of the present invention effectively avoids the overfitting problem caused by the model's excessive dependence on generated data through the optimization strategy of iteratively reducing generated samples, and significantly improves the generalization ability of the model.

[0031] In an alternative embodiment, the generator includes an input layer, a fully connected layer, a convolutional layer, a LeakyReLU activation function, and a Tanh activation function; The fully connected layer is used to learn the complex representation of the input random noise vector; The convolutional layer is used to extract local dependencies in the voice signal; The Tanh activation function is used to scale the output signal to the range of [-1, 1].

[0032] Specifically, the generator adopts a multi-layer neural network structure. The input is a random noise vector, and the output is a synthesized speech sample. The generator maps the random noise into a high-quality speech signal through a series of fully connected layers and convolutional layers. To ensure the diversity of the generated samples, the generator introduces the LeakyReLU activation function and uses the Tanh activation function in the last layer to limit the output within the range of [-1, 1], simulating the amplitude range of real speech signals.

[0033] The voiceprint recognition method provided by the embodiment of the present invention ensures the diversity of the generated samples through the LeakyReLU activation function and uses the Tanh activation function to limit the output signal range, simulating real speech signals.

[0034] In an optional embodiment, the discriminator includes multiple convolutional layers, fully connected layers, and a Sigmoid activation function; The multiple convolutional layers are used to extract features from the input data; The fully connected layers are used to generate classification results based on the extracted features; The Sigmoid activation function is used to compress the output data into the range of (0, 1).

[0035] Specifically, the discriminator is a binary classification neural network. The input is a real speech sample or a generated sample, and the output is the authenticity probability of the sample. The discriminator extracts speech features through multiple convolutional layers and fully connected layers and uses the Sigmoid activation function to output the authenticity probability of the sample. The goal of the discriminator is to maximize the ability to distinguish between real samples and generated samples, thereby promoting the generator to generate more realistic speech samples.

[0036] Optionally, to ensure the quality of the generated samples, the following quality control mechanisms are introduced: Wasserstein GAN with Gradient Penalty (WGAN-GP): Adopt the WGAN-GP architecture to stabilize the training process through the gradient penalty term and avoid the problem of Mode Collapse; Diversity monitoring: Monitor the diversity of the generated samples by visualizing the spectrograms and statistical distributions of the generated samples to ensure that the generated samples cover the statistical distribution of real speech.

[0037] In an optional embodiment, the target voiceprint recognition model extracts features from the speech to be recognized based on the following steps: Perform frame splitting on the input signal to obtain multiple frames of speech signals; Process each frame of the speech signals in the multiple frames of speech signals with a Hamming window to obtain multiple frames of windowed speech signals; Perform a fast Fourier transform on each frame of the windowed speech signals of the multi-frame windowed speech signals to obtain multi-frame short-time spectra; Square each frame of the short-time spectra of the multi-frame short-time spectra to obtain multi-frame power spectra; Input the multi-frame power spectra into a Mel filter bank to obtain a filtered signal output by the Mel filter bank; Take the logarithm of the filtered signal to obtain a logarithmic filtered signal; Perform a discrete cosine transform on the logarithmic filtered signal to obtain Mel frequency cepstral coefficients; Calculate the first-order difference and second-order difference of the Mel frequency cepstral coefficients to correspondingly obtain Delta coefficients and Delta-Delta coefficients; Concatenate the Mel frequency cepstral coefficients, the Delta coefficients, and the Delta-Delta coefficients to obtain a feature vector.

[0038] Specifically, first, frame the input speech signal, splitting the continuous speech signal into multiple short-time frames. The length of each frame is usually 20 - 30 milliseconds, and there is a certain overlap (such as 10 milliseconds) between frames. The purpose of framing is to capture the short-time characteristics of the speech signal. Apply a Hamming Window to each frame of the speech signal to reduce the spectral leakage effect at the frame edges and ensure the accuracy of spectral analysis. Perform a Fast Fourier Transform (FFT) on each windowed speech signal frame to convert the time-domain signal into a frequency-domain signal, obtaining a short-time spectrum. Square the FFT result to get the power spectrum of each frame for subsequent Mel filter bank processing. Pass the power spectrum through a set of Mel filters to simulate the non-linear perception of frequency by the human ear. The design of the Mel filter bank is based on the Mel scale, with narrower filters in the low-frequency region and wider filters in the high-frequency region, which can better capture the low-frequency characteristics of the speech signal. Take the logarithm of the output of the Mel filter bank to compress the dynamic range and enhance the details in the low-frequency part. Perform a Discrete Cosine Transform (DCT) on the logarithmically compressed output of the Mel filter bank to obtain Mel Frequency Cepstral Coefficients (MFCC). The DCT transform can concentrate the spectral energy on a few coefficients, reduce the feature dimension, and at the same time retain the key spectral features of the speech signal. Usually, the first 13 MFCC coefficients are extracted as static features to describe the spectral characteristics of the speech signal. Calculate the first-order difference (Delta coefficient) of the MFCC to capture the short-time changes of the speech signal, and calculate the second-order difference (Delta-Delta coefficient) of the MFCC to capture the long-time changes of the speech signal. Concatenate the static MFCC coefficients, Delta coefficients, and Delta-Delta coefficients to form the final feature vector for subsequent model training and recognition.

[0039] In an alternative embodiment, the target voiceprint recognition model extracts features from the speech to be recognized based on the following steps: Extract features from the speech to be recognized through a convolutional network of different scales connected in sequence, obtaining feature maps of different scales. Each convolutional network includes a convolutional layer and a max-pooling layer; Use an attention module for each of the feature maps to correspondingly obtain multiple attention feature maps; Perform a fusion process on each of the attention feature maps to obtain a feature vector; Among them, the attention module generates an attention feature map based on the following steps: Perform global average pooling on each feature channel to obtain the global feature representation of each feature channel; Input the global feature representation into a fully connected layer to obtain the importance weights of each feature channel; Update the feature map based on the importance weights to obtain the channel attention result; Compress the weight feature map in the channel dimension to obtain the spatial attention map; Multiply the spatial attention map by the feature map to obtain the spatial attention result; Perform weighted processing on the channel attention result and the spatial attention result to obtain the attention feature map.

[0040] Traditional convolutional neural networks usually can only capture features at a single scale when processing speech signals, and it is difficult to effectively handle the time-varying characteristics of speech signals. To solve this problem, the present invention introduces multi-scale convolutional kernels and an attention mechanism, enabling the model to adaptively focus on the key speaker features in speech signals. Specifically: Extract multi-level features of speech signals through convolutional kernels of different scales: The network includes multiple convolutional layers, and each convolutional layer uses convolutional kernels of different sizes to capture the short-term and long-term features of speech signals. Small-scale convolutional kernels (such as 3x1) can capture the short-term features of speech signals (such as the rapid changes of phonemes), while large-scale convolutional kernels (such as 7x1) can capture the long-term features of speech signals (such as the slow changes of intonation). After each convolutional layer, a max-pooling layer is connected to reduce the feature dimension and retain key features; To further enhance the model's ability to capture key features, a multi-scale attention mechanism is introduced. This mechanism dynamically adjusts the weight distribution of the feature map by calculating the importance weights of each feature channel, enabling the model to adaptively focus on the key speaker features in the speech signal. Specifically, global average pooling is performed on each feature channel to obtain the global feature representation of each channel. Global average pooling can capture the global information of the feature channel and reduce the influence of local noise. The result of global average pooling is input into a fully connected layer to calculate the importance weights of each channel. The fully connected layer learns the dependency relationships between feature channels through non-linear transformations. The calculated weights are applied to the original feature map to enhance the response of important feature channels and suppress the response of irrelevant feature channels. The feature map is compressed in the channel dimension to obtain a spatial attention map, which can capture the importance of different positions in the feature map. The spatial attention map is multiplied by the original feature map to enhance the response of important spatial positions and suppress the response of irrelevant spatial positions. The results of channel attention and spatial attention are weighted and fused to form the final feature representation. Through the multi-scale attention mechanism, the model can adaptively focus on the key speaker features in the speech signal, reducing the influence of background noise and speech variability on the recognition result.

[0041] The voiceprint recognition method provided by the embodiments of the present invention enables the model to adaptively focus on the key speaker features in the speech signal through a multi-scale convolutional network and an attention mechanism, reducing the influence of background noise and speech variability on the recognition result.

[0042] For ease of understanding, the preferred voiceprint recognition embodiments of the present invention will be described below in conjunction with Figure 2 For ease of understanding, the preferred voiceprint recognition embodiments of the present invention will be described below in conjunction with

[0043] Specifically, the voiceprint recognition embodiment includes the following six steps: 1) Data collection and preprocessing: Collect diverse speech samples and perform preprocessing; 2) GAN data generation: Generate high-quality synthetic speech samples through GAN to expand the training data set; 3) MFCC feature extraction: Perform MFCC feature extraction on real samples and generated samples; 4) Multi-scale feature fusion and attention mechanism: Extract multi-level features of the speech signal through a multi-scale convolutional network and an attention mechanism; 5) Model training and optimization: Train the voiceprint recognition model using the enhanced data set and optimize the model performance through a dynamic sample balancing mechanism; 6) Model evaluation and application: Evaluate the model performance on the test set and apply the model to actual scenarios.

[0044] Specifically, data collection and preprocessing consists of two parts: data collection and preprocessing. For data collection, diverse speech samples are collected from public speech databases (such as Common Voice) to ensure that the samples cover different genders, languages, and speech characteristics. For preprocessing, the speech samples are converted to the WAV format, the speech signals are framed and windowed, with each frame being 20 - 30 milliseconds long and overlapping by 10 milliseconds between frames. The FFT transform is performed on each frame of the speech signal to obtain the short-time spectrum.

[0045] GAN data generation includes: GAN model training, where a random noise vector is used as the input to the generator to generate synthetic speech samples. The discriminator classifies the generated samples and real samples, and the generator and discriminator continuously improve their performance through adversarial training. The WGAN-GP architecture is adopted to stabilize the training process through a gradient penalty term; quality control of generated samples, by visualizing the spectrograms and statistical distributions of the generated samples to monitor the quality and diversity of the generated samples; dynamic sample balancing, where the generated samples and real samples are mixed in a certain proportion to form an enhanced training dataset, and the number of generated samples is reduced through iteration to find the optimal proportion of generated samples and avoid overfitting.

[0046] MFCC feature extraction: perform Mel filter bank processing on each frame of the speech signal to obtain the Mel spectrum, perform logarithmic compression and DCT transform on the Mel spectrum to extract 13-dimensional MFCC coefficients, calculate the first-order difference (Delta coefficient) and second-order difference (Delta-Delta coefficient) of the MFCC to form the final feature vector; feature standardization and zero-padding, standardize the MFCC features by subtracting the mean and dividing by the standard deviation, and perform zero-padding on shorter speech samples to ensure the consistency of the length of the input data.

[0047] Multi-scale feature fusion and attention mechanism includes: multi-scale convolutional network (MSA-CNN), which uses convolutional kernels of different scales (such as 3x3, 5x5, 7x7) to extract short-time and long-time features of the speech signal. A max-pooling layer is connected after each convolutional layer to reduce the dimension of the feature map and retain key features; attention mechanism, through channel attention mechanism and spatial attention mechanism, dynamically adjusts the feature weights to focus on the key speaker features in the speech signal, and the results of channel attention and spatial attention are weighted and fused to form the final feature representation.

[0048] Model training and optimization include: model training, inputting the enhanced features into a multi-scale convolutional network for training, using the Adam optimizer and the sparse categorical cross-entropy loss function for model training, and adjusting the training cycle according to the size of the dataset, usually 30 - 50 epochs; dynamic sample balancing: gradually reducing the number of generated samples during training, retraining the model and recording the performance, and through multiple iterations, finding the optimal ratio of generated samples to real samples to ensure that the model has good generalization ability while maintaining high accuracy.

[0049] Model evaluation and application include: model evaluation, evaluating the recognition accuracy, precision, recall, and F1 score of the model on the test set, and analyzing the classification performance of the model through a confusion matrix; model application: applying the trained model to actual scenarios, such as multi-language voiceprint recognition, biometric authentication, and voice assistants. In low-resource scenarios, augmenting the training dataset by generating samples can significantly improve the model performance.

[0050] In summary, the voiceprint recognition method provided by the present invention generates high-quality synthetic speech samples through GAN, combines a dynamic sample balancing mechanism, and adaptively expands the training data in long-tail scenarios; adopts MFCC feature extraction technology, fuses static spectrum and dynamic time-varying features, and enhances the voice representation ability; designs a multi-scale convolutional network and an attention mechanism, captures short-term and long-term features of the voice signal through convolutional kernels of different scales, and adaptively focuses on key speaker information; introduces an optimization strategy of iteratively reducing generated samples to avoid model overfitting and improve generalization performance. The present invention solves problems such as data scarcity, insufficient feature representation ability, and poor model robustness in the prior art, and provides an efficient and reliable solution for high-precision voiceprint recognition in low-resource scenarios.

[0051] Next, the voiceprint recognition device provided by the embodiments of the present application will be described. The voiceprint recognition device described below can be correspondingly referred to the voiceprint recognition method described above.

[0052] Figure 3 is a schematic structural diagram of the voiceprint recognition device provided by the present invention, as Figure 3 shown, the voiceprint recognition device may include but is not limited to; The first training module 310 is used to: train the generative adversarial network based on the training dataset; The sample update module 320 is used to: iteratively update the generation sample ratio of the generator in the generative adversarial network based on the performance of the discriminator in the validation dataset in the generative adversarial network to obtain the optimal generation sample ratio; The second training module 330 is used to: train the voiceprint recognition model based on the optimal generation sample ratio to obtain the target voiceprint recognition model; A target recognition module 340, configured to: input the speech to be recognized into the target voiceprint recognition model, and obtain a voiceprint recognition result output by the target voiceprint recognition model.

[0053] It should be noted that the voiceprint recognition device provided in the embodiments of the present invention can execute the voiceprint recognition method described in any of the above embodiments during specific operation, and details thereof will not be elaborated in this embodiment.

[0054] Figure 4 An example of a schematic physical structure diagram of an electronic device is shown as Figure 4 shown. The electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communication interface 420, and the memory 430 complete mutual communication through the communication bus 440. The processor 410 can call logic instructions in the memory 430 to execute a voiceprint recognition method, which includes: training a generative adversarial network based on a training data set; Based on the performance of the discriminator in the generative adversarial network on a validation data set, iteratively update the generation sample ratio of the generator in the generative adversarial network to obtain an optimal generation sample ratio; Based on the optimal generation sample ratio, train a voiceprint recognition model to obtain a target voiceprint recognition model; Input the speech to be recognized into the target voiceprint recognition model, and obtain a voiceprint recognition result output by the target voiceprint recognition model.

[0055] In addition, when the logic instructions in the above-mentioned memory 430 are implemented in the form of software function units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0056] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the voiceprint recognition method provided by each of the above methods, and the method includes: training a generative adversarial network based on a training data set; Based on the performance of the discriminator in the generative adversarial network on a validation data set, iteratively update the generation sample ratio of the generator in the generative adversarial network to obtain an optimal generation sample ratio; Based on the optimal generation sample ratio, train a voiceprint recognition model to obtain a target voiceprint recognition model; Input the voice to be recognized into the target voiceprint recognition model to obtain the voiceprint recognition result output by the target voiceprint recognition model.

[0057] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the voiceprint recognition method provided by each of the above methods, and the method includes: training a generative adversarial network based on a training data set; Based on the performance of the discriminator in the generative adversarial network on a validation data set, iteratively update the generation sample ratio of the generator in the generative adversarial network to obtain an optimal generation sample ratio; Based on the optimal generation sample ratio, train a voiceprint recognition model to obtain a target voiceprint recognition model; Input the voice to be recognized into the target voiceprint recognition model to obtain the voiceprint recognition result output by the target voiceprint recognition model.

[0058] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0059] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A voiceprint recognition method, characterized in that: include: Based on the training data set, the generative adversarial network is trained; Iteratively updating the ratio of generated samples of the generator in the generative adversarial network based on the performance of the discriminator in the generative adversarial network on the verification data set to obtain an optimal ratio of generated samples; Based on the optimal generated sample ratio, the voiceprint recognition model is trained to obtain a target voiceprint recognition model; The speech to be recognized is input into the target voiceprint recognition model to obtain the voiceprint recognition result output by the target voiceprint recognition model.

2. The voiceprint recognition method according to claim 1, characterized in that: The method of iteratively updating the ratio of generated samples of the generator in the generative adversarial network based on the performance of the discriminator in the generative adversarial network on the verification data set to obtain the optimal ratio of generated samples includes: In the process of training the generative adversarial network, gradually reducing the proportion of generated samples; If the performance of the discriminator on the validation data set decreases, reducing the proportion of generated samples; If the performance of the discriminator on the validation data set improves, maintaining or increasing the proportion of generated samples; Iteratively update the ratio of generated samples until the optimal ratio of generated samples is obtained.

3. The voiceprint recognition method according to claim 1 or 2, characterized in that: The generator includes an input layer, a fully connected layer, a convolutional layer, a LeakyReLU activation function and a Tanh activation function; The fully connected layer is used to learn a complex representation of the input random noise vector; The convolutional layer is used to extract local dependencies in the speech signal; The Tanh activation function is used to scale the output signal to the range of [-1, 1].

4. The voiceprint recognition method according to claim 1 or 2, characterized in that: The discriminator includes multiple convolutional layers, fully connected layers and Sigmoid activation functions; The multi-layer convolutional layer is used to extract features from input data; The fully connected layer is used to generate classification results based on the extracted features; The Sigmoid activation function is used to compress the output data into the range of (0, 1).

5. The voiceprint recognition method according to claim 1, characterized in that: The target voiceprint recognition model extracts features of the speech to be recognized based on the following steps: Perform frame processing on the input signal to obtain multiple frames of speech signals; Using a Hamming window to process each frame of speech signal in the multi-frame speech signal to obtain a multi-frame windowed speech signal; Performing a fast Fourier transform on each frame of the multi-frame windowed speech signal to obtain a multi-frame short-time spectrum; Squaring each frame of the short-time spectrum in the multiple frames of short-time spectrum to obtain a multiple-frame power spectrum; Inputting the multi-frame power spectrum into a Mel filter bank to obtain a filtered signal output by the Mel filter bank; Taking the logarithm of the filtered signal to obtain a logarithmic filtered signal; Performing discrete cosine transform on the logarithmic filter signal to obtain Mel-frequency cepstrum coefficients; Calculate the first-order difference and the second-order difference of the Mel-frequency cepstral coefficients, and obtain the Delta coefficient and the Delta-Delta coefficient accordingly; The Mel frequency cepstral coefficient, the Delta coefficient and the Delta-Delta coefficient are concatenated to obtain a feature vector.

6. The voiceprint recognition method according to claim 1, characterized in that: The target voiceprint recognition model extracts features of the speech to be recognized based on the following steps: Extracting features of the speech to be recognized by sequentially connecting convolutional networks of different scales to obtain feature maps of different scales, each of the convolutional networks including a convolutional layer and a maximum pooling layer; Using an attention module on each of the feature maps respectively, and obtaining a plurality of attention feature maps accordingly; Performing fusion processing on the attention feature maps to obtain a feature vector; The attention module generates an attention feature map based on the following steps: Perform global average pooling on each feature channel to obtain the global feature representation of each feature channel; Input the global feature representation into a fully connected layer to obtain the importance weight of each feature channel; Based on the importance weight, updating the feature map to obtain a channel attention result; Compressing the weight feature map in the channel dimension to obtain a spatial attention map; Multiplying the spatial attention map by the feature map to obtain a spatial attention result; The channel attention result and the spatial attention result are weighted to obtain an attention feature map.

7. A voiceprint recognition device, characterized in that: include: The first training module is used to: train the generative adversarial network based on the training data set; A sample updating module, used to iteratively update the generated sample ratio of the generator in the generative adversarial network based on the performance of the discriminator in the generative adversarial network on the verification data set to obtain an optimal generated sample ratio; A second training module is used to: train the voiceprint recognition model based on the optimal generated sample ratio to obtain a target voiceprint recognition model; The target recognition module is used to: input the speech to be recognized into the target voiceprint recognition model to obtain the voiceprint recognition result output by the target voiceprint recognition model.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the voiceprint recognition method according to any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the voiceprint recognition method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the voiceprint recognition method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Training method and system for optimizing machine readable area recognition model

    CN120976670A