Model training method based on speech multi-task and speech multi-task processing method
By sharing feature information in the voice multitasking model and adopting a staged training strategy, the problem that the existing technology is difficult to detect forged sounds and recognize voiceprints at the same time is solved, and efficient recognition and verification of forged voiceprints and voiceprints is achieved, which improves the overall performance of the model.
Patent Information
- Application Number
- CN202510089613.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The prior art is difficult to effectively detect the authenticity of forged sounds and identify the voiceprint information of a specific person at the same time, and cannot fully deal with the security threats brought by synthetic voice.
Through a model training method based on voice multitasking, the pre-trained large model is used as a base to enable the forged detection tasks and voiceprint recognition tasks to share information at the feature representation layer, and combined with phased training strategies, the overall performance of the model is improved.
Accurate identification of forged voice and verification of speaker identity, taking into account information authenticity and individual identity security, and improving the overall performance of the model.
Smart Images

Figure CN119580745B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a model training method based on speech multi-task and a speech multi-task processing method. Background Art
[0002] Generative AI technology can generate new content, such as images, sounds, text, etc., by learning from large-scale data. The rapid development of generative AI technology has brought convenience to content production and human-computer interaction, but it has also brought many security risks and social problems.
[0003] Taking the generation of voice content as an example, the technology of forged voice has become one of the areas of concern due to its highly realistic generation effect. Among them, the voice synthesized by the deep learning model can not only imitate the voice characteristics of a specific person, but also achieve highly realistic conversation capabilities, which greatly increases the risk of forged voice being maliciously used in certain fields. Therefore, in practical applications, it is necessary to identify whether the voice is forged, and to accurately extract and verify the voiceprint information of a specific person in the forged voice, in order to comprehensively deal with the security threats brought by synthetic voice.
[0004] At present, there is an urgent need to provide a new solution to effectively and simultaneously complete the authenticity detection of forged voices and the voiceprint recognition task of specific people's voices, so as to cope with the multiple challenges brought about by the rapid development of synthetic speech technology. Summary of the invention
[0005] The embodiment of this specification provides a model training method based on speech multi-task, which uses a large speech pre-trained model as a base, so that the forgery detection task and the voiceprint recognition task can share information at the feature representation layer, improve the efficiency of feature utilization, and combine the phased training strategy to obtain the potential connection between the two downstream tasks, further improving the overall performance of the model. The method includes:
[0006] Acquire a target speech signal, and extract a first speech feature in the target speech signal using a pre-trained feature extraction model;
[0007] Processing the first speech feature by using a forgery detection model to obtain a first speech forgery probability;
[0008] training the forgery detection model based on the first speech forgery probability, and freezing the parameters of the forgery detection model after the training is completed;
[0009] The first speech feature is processed by a voiceprint recognition model to obtain first voiceprint information, and the voiceprint recognition model is trained based on the first voiceprint information.
[0010] Further, in some embodiments, the voiceprint recognition model includes a first convolutional layer, a first pooling layer, and a first fully connected layer;
[0011] The processing of the first speech feature by a voiceprint recognition model to obtain first voiceprint information includes:
[0012] Performing enhancement processing on the first speech feature through the first convolutional layer to obtain an enhanced feature representation;
[0013] Compressing the enhanced feature representation through the first pooling layer to obtain a first feature vector;
[0014] The first fully connected layer is used to perform feature transformation on the first feature vector to obtain the first voiceprint information.
[0015] Further, in some embodiments, the first convolutional layer includes a grouped convolutional network, a feature integration network, a channel enhancement network, and a residual connection network;
[0016] The step of performing enhancement processing on the first speech feature through the first convolutional layer to obtain an enhanced feature representation includes:
[0017] Performing a group convolution operation on the first speech feature according to a preset channel grouping strategy through the group convolution network to obtain a plurality of initial features;
[0018] The plurality of initial features are integrated through a step-by-step connection mechanism of the feature integration network to obtain a multi-scale feature representation;
[0019] Performing weighted processing on each feature channel represented by the multi-scale feature through the channel enhancement network to obtain a weighted feature;
[0020] The weighted feature is superimposed on the first speech feature through the residual connection network to obtain the enhanced feature representation.
[0021] Further, in some embodiments, the channel enhancement network includes a feature compression unit, a channel weight calculation unit and a channel enhancement unit;
[0022] The step of performing weighted processing on each feature channel represented by the multi-scale feature through the channel enhancement network to obtain a weighted feature includes:
[0023] Performing a global average pooling operation on the multi-scale feature representation by the feature compression unit to generate a global description vector;
[0024] Performing a nonlinear transformation on the global description vector by the channel weight calculation unit to generate a channel weight vector;
[0025] The channel enhancement unit performs weighted processing on each feature channel represented by the multi-scale feature according to the channel weight vector to obtain a weighted feature.
[0026] Further, in some embodiments, the forgery detection model includes a plurality of second convolutional layers, a second pooling layer, a second fully connected layer, and an output layer;
[0027] The step of processing the first voice feature by using a forgery detection model to obtain a first voice forgery probability includes:
[0028] Extracting original forged features from the first speech features using the plurality of second convolutional layers;
[0029] Downsampling the original forged features through the second pooling layer to obtain intermediate forged features;
[0030] Mapping the intermediate forged feature into a second feature vector through the second fully connected layer;
[0031] The second feature vector is input into the output layer to obtain the first speech forgery probability.
[0032] Further, in some implementations, the training of the forgery detection model based on the first speech forgery probability and, after the training is completed, freezing the parameters of the forgery detection model include:
[0033] Calculating a loss value of the forgery detection model according to the first speech forgery probability and the corresponding true label, and iteratively training parameters of the forgery detection model using the loss value;
[0034] If the iteration termination condition is met, the training is completed and the parameters of the forgery detection model are fixed.
[0035] Further, in some embodiments, the feature extraction model includes a plurality of Transformer layers;
[0036] The extracting the first speech feature in the target speech signal by using a pre-trained feature extraction model includes:
[0037] The target speech signal is encoded through a plurality of pre-trained Transformer layers to obtain the first speech feature.
[0038] Further, in some embodiments, each of the Transformer layers includes a multi-head attention sublayer and a feed-forward neural network sublayer;
[0039] The encoding of the target speech signal through the pre-trained multiple Transformer layers to obtain the first speech feature includes:
[0040] Inputting the target speech signal into the multi-head attention sublayer, using the multi-head attention sublayer to determine the attention score between each speech frame in the target speech signal, and obtaining an intermediate feature representation according to each of the attention scores;
[0041] The intermediate feature representation is transformed layer by layer using the feedforward neural network sublayer to generate the first speech feature.
[0042] Furthermore, in some implementations, the first speech feature is a feature representation of four dimensions, where the four dimensions include a batch size, a number of feature layers, a length of a speech sample, and a feature dimension of each feature layer.
[0043] The embodiment of this specification also proposes a voice multi-task processing method, the method comprising:
[0044] Acquire a speech signal to be processed, and extract a second speech feature from the speech signal to be processed using a pre-trained feature extraction model;
[0045] Using the trained forgery detection model to detect the second voice feature, and obtain a voice forgery detection result;
[0046] extracting second voiceprint information from the second speech feature using a trained voiceprint recognition model, so as to perform speaker verification according to the second voiceprint information;
[0047] The forgery detection model and the voiceprint recognition model are trained by any one of the speech multi-task based model training methods.
[0048] Furthermore, in some implementations, the voice forgery detection result includes a second voice forgery probability, and the method further includes:
[0049] If the second voice forgery probability is less than or equal to a preset voice forgery probability, the second voice feature is input into the trained voiceprint recognition model.
[0050] The embodiment of this specification also proposes a model training device based on speech multi-task, the device comprising:
[0051] A first feature extraction module, used to obtain a target speech signal and extract a first speech feature in the target speech signal through a pre-trained feature extraction model;
[0052] a forgery probability determination module, configured to process the first speech feature through a forgery detection model to obtain a first speech forgery probability;
[0053] A first model training module, configured to train the forgery detection model based on the first speech forgery probability, and freeze the parameters of the forgery detection model after the training is completed;
[0054] The second model training module is used to process the first speech feature through a voiceprint recognition model to obtain first voiceprint information, and train the voiceprint recognition model based on the first voiceprint information.
[0055] The embodiment of the present specification also provides a voice multi-task processing device, the device comprising:
[0056] A second feature extraction module is used to obtain a speech signal to be processed, and extract a second speech feature in the speech signal to be processed through a pre-trained feature extraction model;
[0057] a voice forgery detection module, configured to detect the second voice feature using a trained forgery detection model to obtain a voice forgery detection result;
[0058] a voiceprint information acquisition module, configured to extract the second voiceprint information from the second speech feature by using a trained voiceprint recognition model, so as to perform speaker verification according to the second voiceprint information;
[0059] The forgery detection model and the voiceprint recognition model are trained by any one of the speech multi-task based model training methods.
[0060] The embodiments of the present specification also provide a storage medium, wherein the storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the steps of the above method.
[0061] An embodiment of the present specification also provides an electronic device, comprising: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the above method.
[0062] The embodiments of the present specification also provide a computer program product, wherein the computer program product stores at least one instruction, and the at least one instruction is suitable for being loaded by a processor and executing the above method steps.
[0063] In the embodiments of the present specification, first, a target voice signal is obtained, and a first voice feature in the target voice signal is extracted through a pre-trained feature extraction model. Then, the first voice feature is processed through a forgery detection model to obtain a first voice forgery probability. Next, the forgery detection model is trained based on the first voice forgery probability, and after the training is completed, the parameters of the forgery detection model are frozen. Furthermore, the first voice feature is processed through a voiceprint recognition model to obtain first voiceprint information, and the voiceprint recognition model is trained based on the first voiceprint information. Finally, the trained forgery detection model and voiceprint recognition model are used to perform forgery detection and speaker verification on the processed voice signal. On the one hand, the forgery detection model can accurately capture subtle forgery features in the voice signal through training optimization, thereby achieving effective judgment of the authenticity of the voice. After freezing the forgery detection model, the voiceprint recognition model focuses on extracting individual features and improving the ability to distinguish different speakers through targeted training. It can not only obtain the potential connection between the two downstream tasks, but also avoid interference between tasks. On the other hand, through the feature sharing mechanism, a high-quality feature foundation is provided for forgery detection and voiceprint recognition. Combined with the phased training strategy, the voiceprint recognition model has the ability to perform accurate voiceprint analysis on specific individuals based on the optimized features. Using the trained forgery detection model and voiceprint recognition model, it is possible to simultaneously recognize forged voices and verify the identity of the speaker, thereby taking into account both information authenticity and individual identity security and improving the overall performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 A schematic diagram of a system architecture for applying the speech multi-task based model training method provided in the embodiments of this specification.
[0065] Figure 2 A flowchart of a model training method based on speech multi-task provided in an embodiment of this specification.
[0066] Figure 3 A flowchart of another model training method based on speech multi-task provided in an embodiment of this specification.
[0067] Figure 4 A schematic diagram of a process for extracting voiceprint information from a speech signal provided in an embodiment of this specification.
[0068] Figure 5 A flowchart of enhancing speech features provided in an embodiment of this specification.
[0069] Figure 6 The present invention is a flowchart of a method for voice multi-tasking provided in an embodiment of the present invention.
[0070] Figure 7A schematic diagram of the principle of a voice multi-tasking processing method provided in an embodiment of this specification.
[0071] Figure 8 A schematic diagram of the structure of a model training device based on speech multi-task provided in an embodiment of this specification.
[0072] Fig. 9 A schematic diagram of the structure of a voice multi-task processing device provided in an embodiment of this specification.
[0073] Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION
[0074] In order to make the purpose, technical solutions and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.
[0075] Figure 1 A schematic diagram of an architecture of a model training method based on speech multi-tasks provided in an embodiment of this specification is shown.
[0076] like Figure 1 As shown, the system architecture 100 may include one or more of terminal devices such as a smart phone 101, a portable computer 102, a desktop computer 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal device and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or optical fiber cables, etc. The terminal device may be various electronic devices with data processing functions, and the electronic device may have a display screen, which is used to display the first voice forgery probability, the first voiceprint information, the voice forgery detection result, and the second voiceprint information, etc.
[0077] It is understandable that the terminal device can also be various electronic devices with sound collection components and sound playback components. For example, the target voice signal used in the model training process can be obtained by real-time collection of the sound collection component. Of course, the target voice signal can also be pre-recorded by other voice recording devices and stored in the terminal device or server 105, and then called when training the model. This specification does not limit this.
[0078] It should be understood that Figure 1The number of terminal devices, networks and servers in the embodiment is only for illustration. According to the implementation requirements, there may be any number of terminal devices, networks and servers. For example, the server 105 may be a server cluster composed of multiple servers.
[0079] See also Figure 2 , which is a flowchart of a model training method based on speech multi-tasks in the embodiment of this specification. In the embodiment of this specification, the model training method based on speech multi-tasks is applied to a model training device based on speech multi-tasks or an electronic device equipped with a model training device based on speech multi-tasks. Figure 2 The process shown is described in detail, and the model training method based on speech multi-task can specifically include the following steps:
[0080] S202, obtaining a target speech signal, and extracting a first speech feature in the target speech signal through a pre-trained feature extraction model;
[0081] In one or more embodiments of this specification, the target voice signal refers to the voice data used to train the forgery detection task and the voiceprint recognition task. The target voice signal can be a voice signal obtained by converting the text content into a voice synthesis engine based on the input text content by the generative large model, or it can be an audio signal containing voice content imported by a microphone, a recording device or an audio file, which is not limited in the embodiments of this specification.
[0082] The first speech feature in the target speech signal is extracted by a pre-trained feature extraction model. Among them, the pre-trained feature extraction model can be a Whisper model based on the Transformer architecture, and the Whisper model is a deep learning model for automatic speech recognition. Of course, the feature extraction model can also be other speech pre-training models. The embodiment of this specification does not limit the type and structure of the feature extraction model. Through the layer-by-layer processing of the speech pre-training model, a high-dimensional feature representation of the target speech signal can be generated. The feature extraction model can capture the time series pattern and frequency characteristics in the speech signal by training on large-scale speech data. Accordingly, the first speech feature can include information such as audio characteristics, time series patterns, and context dependencies in the target speech signal.
[0083] Optionally, the first speech feature in the target speech signal extracted by the Whisper model is a feature representation of four dimensions, and the four dimensions include batch size, number of feature layers, length of speech samples, and feature dimensions of each feature layer. For example, the first speech feature is recorded as (B, L, T, dim), where B represents the batch size, L represents the number of feature layers, such as L=4, T represents the length of the speech sample, which is a fixed value, and dim represents the feature dimension of each feature layer, such as dim=384.
[0084] It should be noted that the first speech feature is universal and applicable to different downstream tasks, such as forgery detection tasks and voiceprint recognition tasks.
[0085] In addition, before inputting the target speech signal into the pre-trained feature extraction model, the signal can be pre-processed, such as by noise reduction, normalization, etc., so that the target speech signal can be used as input for subsequent feature extraction in a digitized form while ensuring that it contains sufficient speech information for analysis and processing.
[0086] By extracting the first speech feature through the speech pre-training model, it can ensure that the target speech signal can go through an efficient feature extraction process and generate high-quality, multi-level speech feature representation, laying the feature foundation for subsequent forgery detection and voiceprint recognition tasks. This process can not only capture the global dependencies in the speech signal, but also extract high-level feature information, improving the robustness and accuracy of the entire model.
[0087] S204, processing the first voice feature by using a forgery detection model to obtain a first voice forgery probability;
[0088] The forgery detection model refers to a deep learning network used to determine the authenticity of a speech signal. The first speech forgery probability indicates the possibility that the target speech signal is a forged speech, which can be denoted as P, and its value is [0, 1]. The closer the probability value is to 1, the higher the possibility that the speech signal is a forged speech, and the closer the probability value is to 0, the higher the possibility that the speech signal is a real speech. For example, if the first speech forgery probability P is greater than 0.5, the target speech signal can be determined to be a forged speech, otherwise, the target speech signal can be determined to be a real speech.
[0089] Optionally, the forgery detection model may include a convolution layer, a pooling layer, a fully connected layer, and an output layer. Specifically, the first voice feature is used as input and passed to the convolution layer of the forgery detection model, and the first voice feature is scanned by a fixed-size filter to extract local patterns that may exist in the first voice feature, such as frequency changes and time dynamic characteristics, so as to capture local features that may exist in the forged voice, such as artificial traces of the voice generation algorithm. The convolution feature map is downsampled by the pooling layer to extract the maximum value or average value of the local area, and the redundant data is reduced while retaining the key features. The pooled features are flattened and input into the fully connected layer for high-dimensional feature combination and mapping, so as to capture the complex features of the forged voice and further enhance the recognition ability of the forged voice pattern. Finally, the output of the fully connected layer is passed to the output layer, and the Sigmoid activation function is used to convert the calculation result of the model into the first voice forgery probability.
[0090] The forgery detection model can effectively capture subtle forgery patterns in speech signals and map these features into probability values, providing a reliable basis for judging the authenticity of the target speech signal.
[0091] S206, training the forgery detection model based on the first voice forgery probability, and freezing the parameters of the forgery detection model after the training is completed;
[0092] By optimizing the parameters of the forgery detection model, the forgery detection model can more accurately judge the authenticity of the target speech signal. Exemplarily, the loss value of the forgery detection model can be calculated based on the first speech forgery probability corresponding to the target speech signal and the corresponding true label, and the parameters of the forgery detection model are iteratively trained according to the loss value. If the iteration termination condition is met, the training is completed. Among them, the true label corresponding to the target speech signal indicates whether the sample is a forged speech, such as a true label of 1 for forged speech, and a true label of 0 for real speech. The iteration termination conditions include reaching the maximum number of iterations, the convergence of the loss function, etc.
[0093] For example, the parameters of the forgery detection model can be updated using a stochastic gradient descent algorithm. According to the principle of back propagation, the loss function is continuously calculated, and the parameters of the forgery detection model are updated according to the calculated loss value. The loss function can be a binary cross entropy loss function, which is used to measure the difference between the probability of forgery of the first voice and the corresponding true label. When the loss function converges to the minimum value, the training of the model parameters is completed. The model parameters can also be updated in a reverse iterative manner, and the training of the model parameters is completed when the preset number of iterations is met. After the iteration is completed, the optimized model parameters can be obtained. In other examples, the alternating least squares method, the Adam optimization algorithm, etc. can be used to minimize the loss function, and the model parameters can be updated from back to front to optimize the model parameters.
[0094] After the training of the forgery detection model is completed, the parameters of the model can be well adapted to the forged voice detection task. In order to avoid the subsequent tasks (such as voiceprint recognition tasks) interfering with the parameters of the model, the parameters of the forgery detection model need to be fixed to maintain task independence. Optionally, the gradients of all parameters in the forgery detection model can be set to zero, or the update of the forgery detection model parameters can be removed in the optimizer to prevent them from being updated in the subsequent training process.
[0095] The forgery detection model and the voiceprint recognition model are two independent task modules. By freezing the parameters of the forgery detection model, interference with the forgery detection capability can be avoided when training the voiceprint recognition model, while ensuring that the two tasks focus on their respective optimization goals at a specific stage. Moreover, the frozen forgery detection model can stably output the probability of voice forgery in the subsequent reasoning stage without the need to readjust the parameters, thereby ensuring robustness in complex scenarios. After freezing the parameters, the forgery detection model does not need to perform gradient calculations during reasoning, thereby reducing the use of computing resources and improving system efficiency.
[0096] S208: Process the first speech feature through a voiceprint recognition model to obtain first voiceprint information, and train the voiceprint recognition model based on the first voiceprint information.
[0097] In one or more embodiments of this specification, the first voice feature contains individual features in the voice signal, which can reflect the identity information of the target speaker. Therefore, the voiceprint recognition model can extract and characterize the embedding vector of the speaker's identity from the first voice feature, that is, the first voiceprint information. The embodiments of this specification do not limit the type of voiceprint recognition model, such as the ECAPA-TDNN model, the Transformer model, etc., wherein the ECAPA-TDNN model is a voiceprint recognition model based on time delay neural network optimization.
[0098] By training and optimizing the parameters of the voiceprint recognition model, the model has higher accuracy and generalization ability in voiceprint feature extraction and speaker classification. Exemplarily, the first speech feature is input into the voiceprint recognition model, and the first voiceprint information is generated through steps such as feature extraction, aggregation and compression. Then, the loss value of the voiceprint recognition model is calculated based on the first voiceprint information and the real speaker label, and the parameters of the voiceprint recognition model are iteratively trained according to the loss value. If the iteration termination condition is met, the training is completed. Among them, it is assumed that there are N speaker categories, and each category corresponds to a real speaker label. The iteration termination conditions include reaching the maximum number of iterations, the convergence of the loss function, etc. It should be noted that when calculating the loss value of the voiceprint recognition model based on the first voiceprint information and the real speaker label, it is first necessary to map the first voiceprint information to the corresponding speaker category, and calculate the loss value based on the speaker category and the real speaker label.
[0099] For example, the parameters of the voiceprint recognition model can be updated using the stochastic gradient descent algorithm. According to the principle of back propagation, the loss function is continuously calculated, and the parameters of the voiceprint recognition model are updated according to the calculated loss value. The loss function can be a binary cross entropy loss function, a center loss function, etc., which is used to measure the difference between the first voiceprint information and the corresponding real speaker label. When the loss function converges to the minimum value, the training of the model parameters is completed. The model parameters can also be updated in a reverse iterative manner, and the training of the model parameters is completed when the preset number of iterations is met. After the iteration is completed, the optimized model parameters can be obtained. In other examples, the alternating least squares method, the Adam optimization algorithm, etc. can be used to minimize the loss function, and the model parameters can be updated from back to front to optimize the model parameters.
[0100] By enhancing the distinguishing ability of voiceprint embedding, the generated voiceprint embedding can better characterize the identity characteristics of the speaker and provide high-quality feature support for subsequent speaker verification tasks.
[0101] See also Figure 3 , provides a flow chart of another model training method based on speech multi-task for the embodiment of this specification. The model training method based on speech multi-task can specifically include the following steps:
[0102] S302, obtaining a target speech signal, and encoding the target speech signal through a plurality of pre-trained Transformer layers to obtain a first speech feature;
[0103] Among them, multiple Transformer layers have the same structure and are stacked. Each Transformer layer is used to extract different levels of features of the target speech signal, including from low-level frequency information to high-level semantic representation.
[0104] Optionally, each Transformer layer includes a multi-head attention sublayer and a feedforward neural network sublayer. Accordingly, the target speech signal can be input into the multi-head attention sublayer, and the multi-head attention sublayer is used to determine the attention scores between each speech frame in the target speech signal, and the intermediate feature representation is obtained according to each attention score. Specifically, the target speech signal can be mapped into a query vector, a key vector and a value vector corresponding to each speech frame through linear mapping, and the attention scores between each speech frame are calculated according to the dot product of the query vector and the key vector corresponding to each speech frame. Then, an attention matrix is generated according to each attention score, and the value vector corresponding to each speech frame is weighted using the attention matrix to obtain the intermediate feature representation. Finally, the intermediate feature representation is transformed layer by layer using the fully connected network and the nonlinear activation function in the feedforward neural network sublayer to generate the first speech feature.
[0105] It should be noted that the parameters of the pre-trained multiple Transformer layers during the entire model training process are fixed and do not need to be updated. The embodiments of this specification can make full use of the deep characteristics learned by the pre-trained model on a wide range of speech data, improve the inter-task feature representation capability, and thus improve the overall performance of the model.
[0106] S304, processing the first voice feature by using a forgery detection model to obtain a first voice forgery probability;
[0107] In one or more embodiments of the present specification, the forgery detection model may include multiple second convolutional layers, second pooling layers, second fully connected layers, and output layers. For example, four second convolutional layers may be included, each of which is connected to batch normalization and ReLU (Rectified Linear Unit) activation functions to introduce nonlinearity and increase the expressive power of the model. The purpose of the convolutional layer is to extract forgery features, and the size and step size of its filter are designed to capture tiny traces of forgery. The second pooling layer is connected to the back of the second convolutional layer to reduce the data dimension while retaining important feature information. After passing through multiple second convolutional layers and second pooling layers, the features are further processed by the fully connected layer and the final classification decision is made. The fully connected layer will be reduced to only one output node, and the Sigmoid activation function is used to output a probability value between 0 and 1, that is, the first speech forgery probability is obtained, indicating the possibility that the target speech signal is a forged speech.
[0108] Exemplarily, multiple second convolutional layers can be used to extract the original forged features in the first speech features, wherein each second convolutional layer performs a convolution operation on the input first speech features, and applies batch normalization and ReLU activation function after the convolution operation to enhance the feature expression capability. The original forged features are downsampled by the second pooling layer to obtain intermediate forged features. The intermediate forged features are then mapped to a second feature vector by a second fully connected layer, such as performing a weighted sum operation on the intermediate forged features to generate a high-dimensional feature representation, that is, obtaining a second feature vector, and the second feature vector is input into the output layer to obtain the first speech forgery probability.
[0109] S306, training the forgery detection model based on the first voice forgery probability, and freezing the parameters of the forgery detection model after the training is completed;
[0110] For step S306, please refer to the detailed description of step S206 in another embodiment of this specification, which will not be repeated here.
[0111] S308, processing the first speech feature through a voiceprint recognition model to obtain first voiceprint information, wherein the voiceprint recognition model includes a first convolutional layer, a first pooling layer, and a first fully connected layer;
[0112] In one or more embodiments of this specification, step S308 may specifically include the following steps:
[0113] S402, performing enhancement processing on the first speech feature through the first convolutional layer to obtain an enhanced feature representation;
[0114] Optionally, the first convolutional layer includes a grouped convolutional network, a feature integration network, a channel enhancement network, and a residual connection network. Based on the network structure of the first convolutional layer, step S402 may further include the following steps:
[0115] S502, performing a group convolution operation on the first speech feature according to a preset channel grouping strategy through the group convolution network to obtain a plurality of initial features;
[0116] Among them, the grouped convolutional network groups multiple channels of the input first speech feature through a preset channel grouping strategy. For example, the number of channels is C and the number of groups is G. If G=C, it means that each channel is processed independently, which is equivalent to a depth-wise separable convolution. If 1<G<C, it indicates uniform division by the number of channels. The appropriate number of groups and channel allocation method can be selected according to the distribution of speech features, such as grouping by frequency range or specific speech pattern. When the input first speech feature is divided into G groups, each group contains C / G channels, and convolution operations are performed on each group of features to obtain the corresponding initial features.
[0117] The grouped convolution decomposes the computational task of full-channel convolution, improving processing efficiency. Moreover, the grouping operation can independently process the features of different groups and capture the short-term and long-term dependent information in the speech signal. The obtained initial features retain diverse time and frequency context information, providing a rich feature basis for subsequent feature integration.
[0118] S504, integrating the multiple initial features through the step-by-step connection mechanism of the feature integration network to obtain a multi-scale feature representation;
[0119] The step-by-step connection mechanism is a recursive feature integration method. At each step, the initial features of the current group are cumulatively connected with the previously integrated features, thereby gradually forming a global multi-scale feature representation.
[0120] For example:
[0121] (1)
[0122] in, is the current integrated feature, which contains the feature information of the current group and all previous groups. For the previous integration features, is the initial feature of the current group, and Concat() represents the channel dimension concatenation operation of the tensor. After the G initial feature groups are integrated, the final multi-scale feature representation is obtained.
[0123] The step-by-step connection integrates the features from different groups in sequence, which not only retains the independence of each group's features, but also reflects the multi-scale characteristics through cumulative connections, enhancing the ability to express complex speech patterns. Importantly, the step-by-step connection mechanism accumulates feature information from multiple resolutions, providing a more comprehensive feature representation for downstream tasks.
[0124] S506, performing weighted processing on each feature channel represented by the multi-scale feature through the channel enhancement network to obtain a weighted feature;
[0125] Among them, the channel enhancement network includes a feature compression unit, a channel weight calculation unit and a channel enhancement unit. The weighted feature is the optimized feature after channel enhancement, which can adaptively highlight the task-related feature channels.
[0126] Exemplarily, the feature compression unit may perform a global average pooling operation on the multi-scale feature representation to generate a global description vector. Then, the channel weight calculation unit may perform a nonlinear transformation on the global description vector to generate a channel weight vector.
[0127] For example, for the global description vector L, a two-layer fully connected network can be used to perform a nonlinear transformation on L. The first layer of the fully connected network is used to reduce the dimension and generate a low-dimensional representation ,have:
[0128] (2)
[0129] in, Can be a ReLU activation function, , are the weights and biases of the first layer of the fully connected network;
[0130] The second layer of fully connected network is used to restore the dimension and generate the channel weight vector ,have:
[0131] (3)
[0132] in, It can be a Sigmoid activation function, , are the weights and biases of the second layer of the fully connected network.
[0133] Finally, the channel enhancement unit performs weighted processing on each feature channel represented by the multi-scale feature according to the channel weight vector to obtain the weighted feature.
[0134] By adaptively adjusting the weights of feature channels, key information related to voice identity is enhanced. Lower weights are given to unimportant channels to reduce the model's sensitivity to noise features or irrelevant features. The channel weighting mechanism optimizes the quality of feature expression and enables the model to focus on effective information.
[0135] S508, superimposing the weighted feature and the first speech feature through the residual connection network to obtain the enhanced feature representation.
[0136] Optionally, the weighted feature and the first speech feature may be element-by-element superimposed to obtain an enhanced feature representation, which retains the basic information of the original speech while highlighting key features related to the speaker's identity, thereby improving the accuracy of voiceprint matching.
[0137] Among them, the first speech feature is directly introduced through the residual connection network, ensuring that the original feature information is not lost after multi-layer processing, and avoiding the weakening of key information due to over-optimization. The residual connection superimposes the optimized weighted features with the original features to form a more comprehensive feature expression. The residual connection alleviates the problem of gradient disappearance in deep networks by means of jump connections, ensuring that the gradient can be effectively propagated from downstream to upstream, thereby improving the training stability of the model and accelerating the convergence speed.
[0138] The embodiments of this specification effectively extract multi-scale features, dynamically optimize feature channels and retain original feature information through the coordinated processing of grouped convolution, step-by-step connection, channel weighting and residual connection, and finally generate enhanced feature representation, which not only improves the model's ability to model complex patterns of speech signals, but also significantly enhances the model's robustness, feature utilization efficiency and training stability, providing a solid foundation for multi-task speech processing.
[0139] S404, compressing the enhanced feature representation through the first pooling layer to obtain a first feature vector;
[0140] The first pooling layer is used to calculate the global average and standard deviation of the enhanced feature representation to generate a first feature vector of fixed dimension, which effectively reduces the dimension of the feature representation and reduces the computational complexity. Moreover, the feature vector can summarize the global information of the target speech signal while retaining the speaker's characteristics, which is helpful for the subsequent extraction of the speaker's voiceprint information.
[0141] S406: Use the first fully connected layer to perform feature transformation on the first feature vector to obtain the first voiceprint information.
[0142] The fully connected layer extracts deep patterns and complex relationships in the first feature vector through weighted combination, and can generate feature representations with higher expressiveness. The first voiceprint information output by the fully connected layer can effectively distinguish the speech characteristics of different speakers in the high-dimensional feature space. Moreover, the first voiceprint information after feature transformation has good versatility and distinguishing ability, and can be used for tasks such as voiceprint matching and identity authentication.
[0143] The embodiments of this specification gradually optimize the feature representation capability of the speech signal through feature enhancement of the first convolutional layer, feature compression of the first pooling layer, and high-dimensional mapping of the first fully connected layer. The first voiceprint information finally generated has richness, robustness, and distinctiveness, which not only improves the quality of the voiceprint information, but also provides more efficient and accurate feature input for subsequent voiceprint recognition tasks.
[0144] S310: Training the voiceprint recognition model based on the first voiceprint information.
[0145] For step S310, please refer to the detailed description of training the voiceprint recognition model in step S208 in another embodiment of this specification, which will not be repeated here.
[0146] The voiceprint recognition model focuses on the individual feature information in the target voice, and can specifically extract the speaker's voiceprint features and verify them. The voiceprint recognition model is trained with voiceprint information, so that the voiceprint recognition model has the ability to accurately distinguish different individuals in the high-dimensional voice feature space.
[0147] The anti-counterfeiting detection model in the embodiments of this specification can be optimized for the subtle differences between synthesized speech and real speech, and the voiceprint recognition model is optimized for the feature extraction and recognition of voiceprints, thereby improving the processing effect of each task. Moreover, by training the model in stages, it can be ensured that the model achieves good results in both anti-counterfeiting detection tasks and voiceprint recognition tasks. Because training the anti-counterfeiting detection model first can enhance the model's sensitivity to forged speech, providing cleaner and more targeted speech data for subsequent voiceprint recognition. By focusing on the forgery detection task first, the model can give priority to learning how to distinguish between real speech and forged speech in the early stages of training, thereby avoiding the interference of forged speech that may occur during voiceprint recognition. Therefore, it is possible to improve the accuracy of voiceprint recognition while ensuring the ability to detect forged speech. In addition, through the feature sharing mechanism and the staged training strategy, the voiceprint recognition model has the ability to perform accurate voiceprint analysis on a specific individual based on the optimized features, thereby avoiding the inefficiency caused by the separate processing of speech signal authenticity and speaker identity authentication tasks.
[0148] See also Figure 6 , is a flow chart of a voice multitasking processing method provided in an embodiment of this specification. In an embodiment of this specification, the voice multitasking processing method is applied to a voice multitasking processing device or an electronic device equipped with a voice multitasking processing device. In one or more embodiments of this specification, reference Figure 7 FIG. 1 is a schematic diagram showing the principle of a method for voice multitasking processing provided by an embodiment of this specification, combined with Figure 7 The network architecture and implementation principle shown are Figure 6 The steps shown in the following are explained:
[0149] S602, obtaining a speech signal to be processed, and extracting a second speech feature in the speech signal to be processed by using a pre-trained feature extraction model;
[0150] In one or more embodiments of the present specification, the speech signal to be processed 701 refers to an audio signal that needs to be subjected to forgery detection and speaker verification. The speech signal to be processed 701 may be the same as the target speech signal or may be different from the target speech signal, and the present specification does not limit this.
[0151] Furthermore, the speech signal to be processed 701 may be subjected to preprocessing such as noise reduction, normalization, framing and windowing, so as to adjust the speech signal to be processed into data that can be directly processed by the feature extraction model.
[0152] exist Figure 7In the example, the pre-trained feature extraction model is the Whisper model 702. For example, the Whisper model 702 is composed of four stacked blocks, each of which is a Transformer architecture. The speech signal to be processed 701 is encoded by multiple blocks to obtain a feature representation of four dimensions, that is, the second speech feature 703 is obtained. The four dimensions may include the batch size, the number of feature layers, the length of the speech sample, and the feature dimension of each feature layer. The encoding process of the speech signal to be processed 701 can refer to the encoding process of the target speech signal in step S302, which will not be repeated here.
[0153] S604, using the trained forgery detection model to detect the second voice feature to obtain a voice forgery detection result;
[0154] Among them, the voice forgery detection result includes the second voice forgery probability. The forgery detection model 704 includes multiple second convolutional layers, a second pooling layer, a second fully connected layer and an output layer. Specifically, the second voice feature 703 is input into the forgery detection model 704, and the original forgery features in the second voice feature 703 are extracted by using multiple second convolutional layers, and the original forgery features are downsampled by the second pooling layer to obtain intermediate forgery features. Then, the intermediate forgery features are mapped to a second feature vector by the second fully connected layer, such as performing a weighted sum operation on the intermediate forgery features to generate a high-dimensional feature representation, that is, obtaining a second feature vector, and the second feature vector is input into the output layer to obtain the second voice forgery probability.
[0155] If the second voice forgery probability is less than or equal to the preset voice forgery probability, the second voice feature 703 can be input into the trained voiceprint recognition model 705. For example, if the preset voice forgery probability is 0.5, if the second voice forgery probability P is less than or equal to 0.5, the voice signal to be processed can be determined to be a real voice 706, otherwise, the voice signal to be processed can be determined to be a forged voice 707. This step provides the first layer of protection for forged voice detection to avoid the potential risks brought by false voice information.
[0156] S606: Use the trained voiceprint recognition model to extract second voiceprint information from the second speech feature, so as to perform speaker verification according to the second voiceprint information.
[0157] The second voiceprint information is an embedded vector representing the speaker's identity. The voiceprint recognition model 705 includes a first convolutional layer, a first pooling layer, and a first fully connected layer, and of course also includes an output layer ( Figure 7). It should be noted that the first convolution layer includes a group convolution network, a feature integration network, a channel enhancement network and a residual connection network, and the channel enhancement network includes a feature compression unit, a channel weight calculation unit and a channel enhancement unit. The second speech feature 703 is input into the voiceprint recognition model 705 to extract the corresponding second voiceprint information. The process can refer to the process of enhancing the first speech feature in step S308, which will not be repeated here.
[0158] After obtaining the second voiceprint information, the second voiceprint information can be similarly calculated with the template voiceprint in the registered voiceprint library, such as cosine similarity. If the similarity value is higher than the preset similarity threshold, the verification is passed and the input voice is the target speaker 708. If the similarity is lower than the threshold, the verification fails and the input voice is not the target speaker.
[0159] The trained voiceprint recognition model 705 can accurately identify the identity of the speaker in the voice signal, and by combining with the forgery detection result, it provides double protection for the authenticity and individuality of the voice signal.
[0160] It should be noted that the forgery detection model 704 and the voiceprint recognition model 705 are optimized by the method described in detail in steps S202 to S208 in one embodiment of this specification, or trained by the method described in detail in steps S302 to S310 in another embodiment of this specification, which will not be repeated here.
[0161] In the embodiments of this specification, the speech features in the speech signal are extracted through a pre-trained feature extraction model, so that the feature representation contains rich speech information in the speech signal, providing a high-quality feature basis for subsequent processing. Forgery detection based on speech features can identify the authenticity of the speech from the subtle features of the speech signal, and at the same time extract individual features in the speech signal in the voiceprint recognition model. The combination of the two models can analyze and process the target speech signal from multiple dimensions, and then simultaneously complete the judgment of the authenticity of the speech and the verification of the identity of the specific speaker. It can not only deal with the problems of information fraud and privacy leakage that may be caused by forged speech, but also provide security authentication of the identity of a specific individual in complex application scenarios, and has high practical value.
[0162] See also Figure 8 , is a schematic diagram of a structure of a model training device based on speech multi-task provided in an embodiment of this specification. Figure 8As shown, the model training device 1 based on voice multi-task can be implemented as all or part of the electronic device through software, hardware or a combination of both. According to some embodiments, the model training device 1 based on voice multi-task includes a first feature extraction module 11, a forgery probability determination module 12, a first model training module 13 and a second model training module 14, specifically including:
[0163] A first feature extraction module 11 is used to obtain a target speech signal and extract a first speech feature in the target speech signal through a pre-trained feature extraction model;
[0164] a forgery probability determination module 12, configured to process the first speech feature through a forgery detection model to obtain a first speech forgery probability;
[0165] A first model training module 13 is used to train the forgery detection model based on the first speech forgery probability, and freeze the parameters of the forgery detection model after the training is completed;
[0166] The second model training module 14 is used to process the first speech feature through a voiceprint recognition model to obtain first voiceprint information, and train the voiceprint recognition model based on the first voiceprint information.
[0167] Optionally, the voiceprint recognition model includes a first convolutional layer, a first pooling layer and a first fully connected layer; when the second model training module 14 processes the first speech feature through the voiceprint recognition model to obtain the first voiceprint information, it is specifically used to:
[0168] Performing enhancement processing on the first speech feature through the first convolutional layer to obtain an enhanced feature representation;
[0169] Compressing the enhanced feature representation through the first pooling layer to obtain a first feature vector;
[0170] The first fully connected layer is used to perform feature transformation on the first feature vector to obtain the first voiceprint information.
[0171] Optionally, the first convolution layer includes a group convolution network, a feature integration network, a channel enhancement network and a residual connection network; when the second model training module 14 performs enhancement processing on the first speech feature through the first convolution layer to obtain an enhanced feature representation, it is specifically used to:
[0172] Performing a group convolution operation on the first speech feature according to a preset channel grouping strategy through the group convolution network to obtain a plurality of initial features;
[0173] The plurality of initial features are integrated through a step-by-step connection mechanism of the feature integration network to obtain a multi-scale feature representation;
[0174] Performing weighted processing on each feature channel represented by the multi-scale feature through the channel enhancement network to obtain a weighted feature;
[0175] The weighted feature is superimposed on the first speech feature through the residual connection network to obtain the enhanced feature representation.
[0176] Optionally, the channel enhancement network includes a feature compression unit, a channel weight calculation unit and a channel enhancement unit; when the second model training module 14 performs weighted processing on each feature channel represented by the multi-scale feature through the channel enhancement network to obtain a weighted feature, it is specifically used to:
[0177] Performing a global average pooling operation on the multi-scale feature representation by the feature compression unit to generate a global description vector;
[0178] Performing a nonlinear transformation on the global description vector by the channel weight calculation unit to generate a channel weight vector;
[0179] The channel enhancement unit performs weighted processing on each feature channel represented by the multi-scale feature according to the channel weight vector to obtain a weighted feature.
[0180] Optionally, the forgery detection model includes a plurality of second convolutional layers, a second pooling layer, a second fully connected layer and an output layer; when the forgery probability determination module 12 processes the first speech feature through the forgery detection model to obtain the first speech forgery probability, it is specifically used to:
[0181] Extracting original forged features from the first speech features using the plurality of second convolutional layers;
[0182] Downsampling the original forged features through the second pooling layer to obtain intermediate forged features;
[0183] Mapping the intermediate forged feature into a second feature vector through the second fully connected layer;
[0184] The second feature vector is input into the output layer to obtain the first speech forgery probability.
[0185] Optionally, when the first model training module 13 trains the forgery detection model based on the first speech forgery probability and freezes the parameters of the forgery detection model after the training is completed, it is specifically used to:
[0186] Calculating a loss value of the forgery detection model according to the first speech forgery probability and the corresponding true label, and iteratively training parameters of the forgery detection model using the loss value;
[0187] If the iteration termination condition is met, the training is completed and the parameters of the forgery detection model are fixed.
[0188] Optionally, the feature extraction model includes a plurality of Transformer layers; when the first feature extraction module 11 extracts the first speech feature in the target speech signal through the pre-trained feature extraction model, it is specifically used to:
[0189] The target speech signal is encoded through a plurality of pre-trained Transformer layers to obtain the first speech feature.
[0190] Optionally, each of the Transformer layers includes a multi-head attention sublayer and a feedforward neural network sublayer; when the first feature extraction module 11 encodes the target speech signal through the pre-trained multiple Transformer layers to obtain the first speech feature, it is specifically used to:
[0191] Inputting the target speech signal into the multi-head attention sublayer, using the multi-head attention sublayer to determine the attention score between each speech frame in the target speech signal, and obtaining an intermediate feature representation according to each of the attention scores;
[0192] The intermediate feature representation is transformed layer by layer using the feedforward neural network sublayer to generate the first speech feature.
[0193] Optionally, the first speech feature in the first feature extraction module 11 is a feature representation of four dimensions, and the four dimensions include batch size, number of feature layers, length of speech samples and feature dimensions of each feature layer.
[0194] The above device embodiments correspond to the method embodiments. For specific descriptions, please refer to the description of the method embodiments, which will not be repeated here. The device embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For specific descriptions, please refer to the corresponding method embodiments.
[0195] See also Fig. 9 , is a structural diagram of a voice multi-task processing device provided in an embodiment of this specification. Fig. 9As shown, the voice multi-task processing device 2 can be implemented as all or part of the electronic device through software, hardware or a combination of both. According to some embodiments, the voice multi-task processing device 2 includes a second feature extraction module 21, a voice forgery detection module 22 and a voiceprint information acquisition module 23, specifically including:
[0196] A second feature extraction module 21 is used to obtain a speech signal to be processed, and extract a second speech feature in the speech signal to be processed through a pre-trained feature extraction model;
[0197] A voice forgery detection module 22, configured to detect the second voice feature using a trained forgery detection model to obtain a voice forgery detection result;
[0198] A voiceprint information acquisition module 23, configured to extract the second voiceprint information from the second speech feature using a trained voiceprint recognition model, so as to perform speaker verification according to the second voiceprint information;
[0199] The forgery detection model and the voiceprint recognition model are trained by the method described in the model training device 1 based on speech multi-task.
[0200] Optionally, the voice forgery detection result includes a second voice forgery probability, and the voice multi-task processing device 2 further includes a forgery result evaluation module, which is specifically used to:
[0201] If the second voice forgery probability is less than or equal to a preset voice forgery probability, the second voice feature is input into the trained voiceprint recognition model.
[0202] The above device embodiments correspond to the method embodiments. For specific descriptions, please refer to the description of the method embodiments, which will not be repeated here. The device embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For specific descriptions, please refer to the corresponding method embodiments.
[0203] The present specification also provides a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded and executed by a processor as described above. Figures 2 to 7 The method of the embodiment shown in the figure can be specifically executed by referring to Figures 2 to 7 The specific description of the illustrated embodiment will not be repeated here.
[0204] The present specification also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor as described above. Figures 2 to 7 The method of the embodiment shown in the figure can be specifically executed by referring to Figures 2 to 7The specific description of the illustrated embodiment will not be repeated here.
[0205] The embodiments of this specification also provide Fig.10 The structural diagram of the electronic device shown in FIG. Fig.10 At the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above-mentioned voice activity detection method.
[0206] Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the executor of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0207] In the 1990s, it was very clear whether the improvement of a technology was hardware improvement (for example, improvement of the circuit structure of diodes, transistors, switches, etc.) or software improvement (improvement of the method flow). However, with the development of technology, many improvements of the method flow today can be regarded as direct improvements of the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented with hardware entity modules. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask chip manufacturers to design and make dedicated integrated circuit chips. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0208] The controller may be implemented in any suitable manner, for example, the controller may take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (e.g., software or firmware) executable by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320, and the memory controller may also be implemented as part of the control logic of the memory. It is also known to those skilled in the art that, in addition to implementing the controller in a purely computer-readable program code manner, the controller may be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller may be considered as a hardware component, and the devices for implementing various functions included therein may also be considered as structures within the hardware component. Or even, the devices for implementing various functions may be considered as both software modules for implementing the method and structures within the hardware component.
[0209] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0210] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0211] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0212] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0213] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0214] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0215] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0216] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0217] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0218] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0219] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0220] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0221] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0222] The above description is only an embodiment of the present specification and is not intended to limit the present specification. For those skilled in the art, the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims of the present specification.
Claims
1. A model training method based on speech multi-task, the method comprising: Acquire a target speech signal, and extract a first speech feature in the target speech signal using a pre-trained feature extraction model; Processing the first speech feature by using a forgery detection model to obtain a first speech forgery probability; training the forgery detection model based on the first speech forgery probability, and freezing the parameters of the forgery detection model after the training is completed; Processing the first speech feature through a voiceprint recognition model to obtain first voiceprint information, and training the voiceprint recognition model based on the first voiceprint information; The voiceprint recognition model includes a group convolutional network, a feature integration network, a channel enhancement network, a residual connection network, a first pooling layer and a first fully connected layer; processing the first speech feature by the voiceprint recognition model specifically includes: Performing a group convolution operation on the first speech feature according to a preset channel grouping strategy through the group convolution network to obtain a plurality of initial features; The plurality of initial features are integrated through a step-by-step connection mechanism of the feature integration network to obtain a multi-scale feature representation; Performing weighted processing on each feature channel represented by the multi-scale feature through the channel enhancement network to obtain a weighted feature; superimposing the weighted feature and the first speech feature through the residual connection network to obtain an enhanced feature representation; Compressing the enhanced feature representation through the first pooling layer to obtain a first feature vector; The first fully connected layer is used to perform feature transformation on the first feature vector to obtain the first voiceprint information.
2. According to the model training method based on speech multi-tasks according to claim 1, the channel enhancement network includes a feature compression unit, a channel weight calculation unit and a channel enhancement unit; The step of performing weighted processing on each feature channel represented by the multi-scale feature through the channel enhancement network to obtain a weighted feature includes: Performing a global average pooling operation on the multi-scale feature representation by the feature compression unit to generate a global description vector; Performing a nonlinear transformation on the global description vector by the channel weight calculation unit to generate a channel weight vector; The channel enhancement unit performs weighted processing on each feature channel represented by the multi-scale feature according to the channel weight vector to obtain a weighted feature.
3. According to the model training method based on speech multi-tasks according to claim 1, the forgery detection model comprises a plurality of second convolutional layers, a second pooling layer, a second fully connected layer and an output layer; The step of processing the first voice feature by using a forgery detection model to obtain a first voice forgery probability includes: Extracting original forged features from the first speech features using the plurality of second convolutional layers; Downsampling the original forged features through the second pooling layer to obtain intermediate forged features; Mapping the intermediate forged feature into a second feature vector through the second fully connected layer; The second feature vector is input into the output layer to obtain the first speech forgery probability.
4. The model training method based on speech multi-task according to claim 1, wherein the forgery detection model is trained based on the first speech forgery probability, and after the training is completed, the parameters of the forgery detection model are frozen, comprising: Calculating a loss value of the forgery detection model according to the first speech forgery probability and the corresponding true label, and iteratively training parameters of the forgery detection model using the loss value; If the iteration termination condition is met, the training is completed and the parameters of the forgery detection model are fixed.
5. According to the model training method based on speech multi-tasks according to claim 1, the feature extraction model includes multiple Transformer layers; The extracting the first speech feature in the target speech signal by using a pre-trained feature extraction model includes: The target speech signal is encoded through a plurality of pre-trained Transformer layers to obtain the first speech feature.
6. According to the model training method based on speech multi-tasks according to claim 5, each of the Transformer layers includes a multi-head attention sublayer and a feedforward neural network sublayer; The encoding of the target speech signal through the pre-trained multiple Transformer layers to obtain the first speech feature includes: Inputting the target speech signal into the multi-head attention sublayer, using the multi-head attention sublayer to determine the attention score between each speech frame in the target speech signal, and obtaining an intermediate feature representation according to each of the attention scores; The intermediate feature representation is transformed layer by layer using the feedforward neural network sublayer to generate the first speech feature.
7. According to the speech multi-task based model training method of claim 5, the first speech feature is a feature representation of four dimensions, and the four dimensions include batch size, number of feature layers, length of speech samples and feature dimensions of each feature layer.
8. A voice multitasking method, the method comprising: Acquire a speech signal to be processed, and extract a second speech feature in the speech signal to be processed by using a pre-trained feature extraction model; Using the trained forgery detection model to detect the second voice feature, and obtain a voice forgery detection result; extracting second voiceprint information from the second speech feature using a trained voiceprint recognition model, so as to perform speaker verification according to the second voiceprint information; Wherein, the forgery detection model and the voiceprint recognition model are trained by the method described in any one of claims 1-7.
9. The voice multitasking method according to claim 8, wherein the voice forgery detection result comprises a second voice forgery probability, and the method further comprises: If the second voice forgery probability is less than or equal to a preset voice forgery probability, the second voice feature is input into the trained voiceprint recognition model.
10. A model training device based on speech multi-task, the device comprising: A first feature extraction module, used to obtain a target speech signal and extract a first speech feature in the target speech signal through a pre-trained feature extraction model; a forgery probability determination module, configured to process the first speech feature through a forgery detection model to obtain a first speech forgery probability; A first model training module, configured to train the forgery detection model based on the first speech forgery probability, and freeze the parameters of the forgery detection model after the training is completed; a second model training module, configured to process the first speech feature through a voiceprint recognition model to obtain first voiceprint information, and train the voiceprint recognition model based on the first voiceprint information; The voiceprint recognition model includes a group convolutional network, a feature integration network, a channel enhancement network, a residual connection network, a first pooling layer and a first fully connected layer; the second model training module is specifically used for: Performing a group convolution operation on the first speech feature according to a preset channel grouping strategy through the group convolution network to obtain a plurality of initial features; The plurality of initial features are integrated through a step-by-step connection mechanism of the feature integration network to obtain a multi-scale feature representation; Performing weighted processing on each feature channel represented by the multi-scale feature through the channel enhancement network to obtain a weighted feature; superimposing the weighted feature and the first speech feature through the residual connection network to obtain an enhanced feature representation; Compressing the enhanced feature representation through the first pooling layer to obtain a first feature vector; The first fully connected layer is used to perform feature transformation on the first feature vector to obtain the first voiceprint information.
11. A voice multitasking processing device, the device comprising: A second feature extraction module is used to obtain a speech signal to be processed, and extract a second speech feature in the speech signal to be processed through a pre-trained feature extraction model; a voice forgery detection module, configured to detect the second voice feature using a trained forgery detection model to obtain a voice forgery detection result; a voiceprint information acquisition module, configured to extract the second voiceprint information from the second speech feature by using a trained voiceprint recognition model, so as to perform speaker verification according to the second voiceprint information; Wherein, the forgery detection model and the voiceprint recognition model are trained by the method described in any one of claims 1-7.
12. A storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the computer program implements the steps of the method described in any one of claims 1 to 7, or implements the steps of the method described in any one of claims 8 to 9.
13. An electronic device comprising: A processor and a memory; wherein the memory stores a computer program, the computer program being suitable for being loaded by the processor and executing the steps of the method as described in any one of claims 1 to 7, or implementing the steps of the method as described in any one of claims 8 to 9.
14. A computer program product having at least one instruction stored thereon, wherein when the at least one instruction is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented, or the steps of the method described in any one of claims 8 to 9 are implemented.
Citation Information
Patent Citations
Speech recognition method and device, model training method and device, medium and electronic equipment
CN115376498A
Speaker identification method and system based on adaptive class boundary interval, and storage medium
CN117877493A
Counterfeit voice detection method and device, storage medium and electronic equipment
CN119170020A