Adaptive Parameter Direction and Amplitude Correction Audio Discrimination Method, Device, Equipment and Medium Based on Continuous Learning

By adjusting the direction and amplitude correction of model parameters during continuous learning, the catastrophic forgetting problem of deep learning models when applied across data sets is solved, the recognition accuracy of new and old data sets is improved, and efficient identification of false speech is achieved.

CN118675509BActive Publication Date: 2025-07-22TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410726652.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-06
Publication Date
2025-07-22
Estimated Expiration
2044-06-06

AI Technical Summary

Technical Problem

The accuracy of existing deep learning models will be greatly reduced when facing speech generated by new algorithms, and they are prone to catastrophic forgetting when applied across data sets, losing the ability to recognize on old data sets.

Method used

By introducing adaptive parameter direction and amplitude correction in the continuous learning process, adjusting the update direction and amplitude of model parameters, using different speech generation methods to obtain the data set, constrain the update amplitude of important parameters, reduce the forgetting of old tasks, and enhance the recognition ability of new tasks.

Benefits of technology

The recognition accuracy of the model on the new data set is improved, while maintaining the recognition ability of the old data set, effectively identifying the old and new types of false speech, and improving the detection accuracy of deep synthetic speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118675509B_ABST
    Figure CN118675509B_ABST
Patent Text Reader

Abstract

The present invention provides an adaptive parameter direction and amplitude correction audio discrimination method, apparatus, device and medium based on continuous learning, which specifically relates to the technical field of audio discrimination. In the embodiments of the present invention, by introducing adaptive amplitude constraints and direction correction in the parameter update during continuous learning, the reduction of the accuracy of the model on the source dataset is reduced, and the accuracy of the model on the target dataset is improved. Thus, using the method proposed in the embodiments of the present invention for continuous learning can greatly improve the detection ability of the model for the generated audio in the new scenario, while hardly affecting the detection ability of the model for the speech types in the previous scenario. Thus, it is possible to effectively identify new and old types of fake speech by using continuous learning, which has the advantage of improving the detection accuracy of deep synthetic speech.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio authentication, and particularly to an audio authentication method, device, equipment and medium for adaptively correcting the direction and amplitude of parameters based on continuous learning. Background Art

[0002] In recent years, with the rapid development of deep learning, voice conversion and speech synthesis technologies have become increasingly mature, and the generated speech produced by deep models has reached a level comparable to that of real people, with wide applications in fields such as human-computer interaction, smart home, entertainment, and education. However, the abuse of generated speech also brings harm to people and society, and the corresponding technology for authenticating the authenticity of speech has also received extensive attention. With the advancement of related research, the detection of generated speech based on deep learning performs excellently on most datasets, but the accuracy rate will still decrease significantly in the scenario of facing speech generated by new algorithms and unknown algorithms.

[0003] In related technologies, in the face of the above problems, the current main solution method is based on the continuous learning method, fine-tuning the model on the speech generated by new and unknown algorithms, and at the same time, during the fine-tuning process, the model adaptively identifies the boundary between the speech generated by the old algorithm and the new algorithm, and adjusts the model parameters.

[0004] However, the above method faces a key problem when applied across datasets: catastrophic forgetting. When the model is fine-tuned with a new dataset to improve the recognition ability of new data, they tend to lose the accuracy rate on the old dataset. This phenomenon is particularly obvious when the feature distributions of different datasets do not match.

[0005] It can be seen that there is an urgent need for a new method for authenticating generated audio at present. Summary of the Invention

[0006] In an embodiment of the present invention, an audio authentication method, device, equipment and medium for adaptively correcting the direction and amplitude of parameters based on continuous learning are provided to effectively authenticate new and old types of false speech by using continuous learning, reduce the decrease in the accuracy rate of the model on the source dataset, and improve the accuracy rate of the model on the target dataset.

[0007] In a first aspect of an embodiment of the present invention, an audio authentication method for adaptively correcting the direction and amplitude of parameters based on continuous learning is proposed, and the method includes:

[0008] Obtain a second generated speech dataset and a speech authentication source model, where the speech authentication source model can authenticate the generated speech in the first generated speech dataset, and the first generated speech dataset and the second generated speech dataset are obtained based on different speech generation methods;

[0009] Using the second generated speech dataset, update the model parameters of the speech discrimination source model by correcting the parameter direction and amplitude, to obtain an updated speech discrimination model, where the updated speech discrimination model can discriminate the generated speech in the first generated speech dataset and the generated speech in the second generated speech dataset;

[0010] Among them, the update of the parameter direction and amplitude correction includes the following steps:

[0011] Adjust the update direction of each model parameter;

[0012] Based on the importance of each model parameter to the overall model, constrain the update amplitude of each model parameter.

[0013] In an alternative embodiment of the present invention, constraining the update amplitude of each model parameter based on the importance of each model parameter to the overall model includes:

[0014] Determine the importance of each model parameter to the overall model based on the following formula:

[0015]

[0016] Among them, Ω j The importance of each model parameter determined based on the j-th generated speech dataset to the overall model, g i,j x l (i, j) represents the influence on the entire model parameter set during the change of each model parameter when the input is x l (i, j); Among them, M(x l (i, j)) represents the output of the overall model M when the input is x l (i, j); ω represents the parameter value of each model parameter obtained after learning the j-th generated speech dataset;

[0017] Determine the update amplitude of the parameter based on the following formula:

[0018]

[0019] Among them, λ represents the constraint strength constant, ω * represents the parameter value of each model parameter obtained after learning the (j - 1)-th generated speech dataset, ω - ω * represents the change amount of each model parameter of the speech discrimination source model during the process of learning the j-th generated speech dataset.

[0020] In an alternative embodiment of the present invention, adjusting the update direction of each model parameter includes: determining the update direction of each model parameter based on the following formula:

[0021]

[0022] Among them, i represents the batch to which the sub-dataset belongs when updating the voice discrimination source model, j represents that the generated voice dataset where the input sub-dataset is located when updating the voice discrimination source model is the j-th generated voice dataset, and α i,j represents a preset constant, T represents transpose, represents the average value of the input voice in the i-th sub-dataset of the j-th generated voice dataset in the (l-1)-th layer; P l (i, j) represents the update direction of the parameters of the l-th layer when updating the voice discrimination source model based on the i-th sub-dataset of the j-th generated voice dataset.

[0023] In an alternative embodiment of the present invention, the parameter direction and amplitude correction updates of each model parameter of the voice discrimination source model are performed based on the following formula:

[0024] W l (i, j) = W l (i - 1, j) + γ(i, j)P l ΔW l BP (i, j) + R;

[0025] Among them, W l (i, j) represents each model parameter in the l-th layer updated based on the i-th sub-dataset of the j-th generated voice dataset, γ represents the learning rate, and P l represents the update direction of each model parameter of the l-th layer, and ΔW l BP (i, j) represents the gradient calculated by using the classical backpropagation algorithm based on the i-th sub-dataset of the j-th generated voice dataset.

[0026] In an alternative embodiment of the present invention, the voice discrimination source model is trained by the following steps:

[0027] Obtain the first generated voice dataset;

[0028] Use the first generated voice dataset to train the deep learning model to obtain the voice discrimination source model.

[0029] In an alternative embodiment of the present invention, using the first generated voice dataset to train the deep learning model includes: updating the parameters of the deep learning model based on the following formula:

[0030] W l (i, j) = W l (i - 1, j) + γ(i, j)ΔW l BP(i,j);

[0031] wherein, W l (i,j) represents each model parameter in the l-th layer updated based on the i-th sub-dataset of the j-th generated speech dataset, γ represents the learning rate, and ΔW l BP (i,j) represents the gradient calculated by using the classical backpropagation algorithm based on the i-th sub-dataset of the j-th generated speech dataset.

[0032] In the second aspect of the embodiments of the present invention, an audio discrimination device for adaptively correcting the parameter direction and amplitude based on continuous learning is proposed. The device includes:

[0033] An acquisition module, configured to acquire a second generated speech dataset and a speech discrimination source model, where the speech discrimination source model can discriminate the generated speech in the first generated speech dataset, and the first generated speech dataset and the second generated speech dataset are obtained based on different speech generation methods;

[0034] An update module, configured to use the second generated speech dataset to update the parameter direction and amplitude of each model parameter of the speech discrimination source model to obtain an updated speech discrimination model, where the updated speech discrimination model can discriminate the generated speech in the first generated speech dataset and the generated speech in the second generated speech dataset;

[0035] wherein, the update of the parameter direction and amplitude correction includes the following steps:

[0036] Adjust the update direction of each model parameter;

[0037] Based on the importance of each model parameter to the overall model, constrain the update amplitude of each model parameter.

[0038] In an optional embodiment of the present invention, constraining the update amplitude of each model parameter based on the importance of each model parameter to the overall model includes:

[0039] Determine the importance of each model parameter to the overall model based on the following formula:

[0040]

[0041] wherein, Ω j The importance of each model parameter determined based on the j-th generated speech dataset to the overall model, g i,j x l (i,j) represents the influence on the entire model parameter set during the change of each model parameter when the input is x l (i,j); wherein, M(x l(i,j)) represents the input x l When (i,j), it is the output of the overall model M; ω represents the parameter values of each model parameter obtained after learning the jth generated speech dataset;

[0042] Determine the update amplitude of this parameter based on the following formula:

[0043]

[0044] Among them, λ represents the constraint strength constant, ω * represents the parameter values of each model parameter obtained after learning the (j - 1)th generated speech dataset, ω - ω * represents the change amount of each model parameter of the speech discrimination source model during the process of learning the jth generated speech dataset.

[0045] In an alternative embodiment of the present invention, adjusting the update direction of each model parameter includes: determining the update direction of each model parameter based on the following formula:

[0046]

[0047] Among them, i represents the batch to which the sub - dataset belongs when updating the speech discrimination source model, j represents that the sub - dataset input when updating the speech discrimination source model is the jth generated speech dataset, α i,j represents a preset constant, T represents transpose, represents the average value of the input speech in the ith sub - dataset of the jth generated speech dataset in the (l - 1)th layer; P l (i,j) represents the update direction of the parameters in the lth layer when updating the speech discrimination source model based on the ith sub - dataset of the jth generated speech dataset.

[0048] In an alternative embodiment of the present invention, update the parameter direction and amplitude correction of each model parameter of the speech discrimination source model based on the following formula:

[0049] W l (i,j) = W l (i - 1,j)+γ(i,j)P l ΔW l BP (i,j)+R;

[0050] Among them, W l (i,j) represents each model parameter in the lth layer updated based on the ith sub - dataset of the jth generated speech dataset, γ represents the learning rate, P l represents the update direction of each model parameter in the lth layer, ΔW l BP(i,j) represents the gradient calculated using the classical backpropagation algorithm for the i-th subset based on the j-th generated speech dataset.

[0051] In an alternative embodiment of the present invention, the speech discrimination source model is trained using the following steps:

[0052] Obtain a first generated speech dataset;

[0053] Use the first generated speech dataset to train a deep learning model to obtain a speech discrimination source model.

[0054] In an alternative embodiment of the present invention, using the first generated speech dataset to train a deep learning model includes: updating the parameters of the deep learning model based on the following formula:

[0055] W l (i,j) = W l (i - 1,j) + γ(i,j)ΔW l BP (i,j);

[0056] where W l (i,j) represents each model parameter in the l-th layer updated based on the i-th subset of the j-th generated speech dataset, γ represents the learning rate, and ΔW l BP (i,j) represents the gradient calculated using the classical backpropagation algorithm for the i-th subset based on the j-th generated speech dataset.

[0057] In a third aspect of the embodiments of the present invention, an electronic device is proposed, including: a memory for storing one or more programs; a processor; when the one or more programs are executed by the processor, implementing the adaptive parameter direction and amplitude correction audio discrimination method based on continuous learning as described in any one of the first aspects above.

[0058] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is proposed, on which a computer program is stored, and when the computer program is executed by a processor, implementing the adaptive parameter direction and amplitude correction audio discrimination method based on continuous learning as described in any one of the first aspects above.

[0059] In the embodiments of the present invention, by introducing adaptive amplitude constraints and direction correction during the continuous learning process for parameter update, the reduction in the accuracy of the model on the source dataset is reduced, and the accuracy of the model on the target dataset is improved. Thus, using the method proposed in the embodiments of the present invention for continuous learning can greatly improve the detection ability of the model for the generated audio in new scenarios, while hardly affecting the detection ability of the model for the speech types in previous scenarios. Therefore, it is possible to effectively identify new and old types of spoofed speech using continuous learning, which has the advantage of improving the detection accuracy of deep synthetic speech. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0061] Figure 1 is a schematic flowchart of the steps of an adaptive parameter direction and amplitude correction audio discrimination method based on continuous learning provided by an embodiment of the present invention;

[0062] Figure 2 is an architecture diagram of an adaptive parameter direction and amplitude correction audio discrimination device based on continuous learning provided by an embodiment of the present invention;

[0063] Figure 3 is a schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0065] Currently, spoofed speech detection can achieve a high accuracy on most datasets. However, in a cross-dataset scenario, that is, when training a model on a source dataset to detect the speech types of a target dataset, the accuracy will be significantly reduced. If the model parameters are directly fine-tuned using the data of the target dataset, the model will "forget" the knowledge learned on the source dataset, manifested as a significant reduction in the recognition accuracy of the model on the source dataset. In related technologies, the continuous learning method for generated speech discrimination overcomes "forgetting" by correcting the direction of the gradient, but it still cannot fully solve the catastrophic forgetting problem.

[0066] Based on this, an embodiment of the present invention proposes an audio discrimination method for adaptively modifying the direction and amplitude of parameters based on continuous learning, enabling the generated speech discrimination model to continuously learn the speech data set generated by the newly generated algorithm, and at the same time adaptively adjusting the amplitude and direction of the model gradient.

[0067] Specifically, in the first invention of the embodiment of the present invention, an audio discrimination method for adaptively modifying the direction and amplitude of parameters based on continuous learning (Adaptive Direction and Amplitude Modification, ADAM) is proposed. Refer to Figure 1 , Figure 1 is a schematic flowchart of the steps of an audio discrimination method for adaptively modifying the direction and amplitude of parameters based on continuous learning provided by the embodiment of the present invention. The method includes the following steps:

[0068] S101: Obtain a second generated speech data set and a speech discrimination source model, where the speech discrimination source model can discriminate the generated speech in the first generated speech data set.

[0069] Specifically, the first generated speech data set and the second generated speech data set are obtained based on different speech generation methods.

[0070] In the embodiment of the present invention, the second generated speech data set may include generated speech obtained by one or more new and unknown speech generation methods.

[0071] In the embodiment of the present invention, the generated speech in the first generated speech data set is generated speech obtained based on one or more known speech generation methods, and the speech discrimination source model has been pre-trained based on the first generated speech data set. The purpose of the embodiment of the present invention is to continuously train the speech discrimination source model so that it learns the generated speech obtained by new and unknown speech generation methods and avoids "forgetting" the generated speech generated by known speech generation methods.

[0072] S102: Use the second generated speech data set to update the direction and amplitude of each model parameter of the speech discrimination source model to obtain an updated speech discrimination model, where the updated speech discrimination model can discriminate the generated speech in the first generated speech data set and the generated speech in the second generated speech data set.

[0073] In the embodiment of the present invention, not only the update direction of the model parameters on the second generated speech data set is considered, but also by introducing a constraint term, different update amplitudes are given to different model parameters.

[0074] Specifically, in step 102, the update of the direction and amplitude of the parameters includes the following steps:

[0075] S1021, Adjust the update directions of each model parameter.

[0076] In the embodiments of the present invention, when the model learns forged audios of new types and unknown types on a new task, by orthogonalizing the direction of the model gradient, the interference of the model to the knowledge learned from the old task is reduced. The corrected direction P of the gradient is as shown in the formula:

[0077]

[0078] where i represents the batch to which the sub-dataset belongs when updating the voice discrimination source model, j represents that the generated voice dataset where the sub-dataset input when updating the voice discrimination source model is the j-th generated voice dataset, α i,j represents a preset constant, T represents the transpose, represents the average value of the input voice in the i-th sub-dataset of the j-th generated voice dataset in the (l - 1)-th layer; P l (i, j) represents the update direction of the parameter in the l-th layer when updating the voice discrimination source model based on the i-th sub-dataset of the j-th generated voice dataset.

[0079] In the embodiments of the present invention, based on the above formula, when learning the input voice in any sub-dataset of any generated voice dataset for the model parameters of any layer of the model, the update method of the model parameters can be adjusted.

[0080] S1022, Constrain the update amplitude of each model parameter based on the importance of each model parameter to the overall model.

[0081] In the embodiments of the present invention, considering those parameters that are relatively important for the recognition of the old task, while orthogonalizing their update directions, their update degrees are constrained, so that the changes of these parameters are small, which helps to reduce the model's forgetting of the old task. For those parameters whose changes have little impact on the recognition of the old task after change, the constraint on their update amplitude can be relaxed, so that these parameters can better learn new knowledge of the new task, thereby improving the learning effect of the model on the new task.

[0082] Specifically, in the embodiments of the present invention, considering the model as a function M from the perspective of the function, the influence of the parameter on the overall recognition effect of the model can be measured by the derivative of the function at different parameters. Specifically, for the deep model M, during the training process, the change estimation amount of the model parameters is as follows: M(x l (i, j), ω + δ) - M(x l (i, j), ω) ≈ g i,j x l (i, j)δ;

[0083] where ω and δ respectively represent specific parameters of the model M at a certain moment and a small change thereof; g i,j x l (i,j) represents the impact on the entire model parameter set during the process of the parameter ω changing by δ when inputting x l (i,j). Specifically, when each dataset (e.g., the first generated speech dataset) is learned by the speech discrimination source model, a small change amount can be added to each model parameter in the speech discrimination source model based on this dataset. Then, the output value M(x l (i,j), ω) of the speech discrimination source model for each model parameter, and the output value M(x l (i,j), ω + δ) of the speech discrimination source model for each model parameter and the change amount are subtracted to obtain the output change amount of the speech discrimination source model. The output change amount of the speech discrimination source model for each model parameter is regarded as the importance of each model parameter determined based on this dataset for the current speech discrimination source model, that is, the importance of the parameter ω for the overall model M, and its calculation is as shown in the formula:

[0084]

[0085] From the above analysis, it can be seen that g i,j x l (i,j) represents the importance of the parameter ω for the model when inputting a part of the data x l (i,j). After the training of each generated speech dataset is completed, all the data used in this training can be input into the model obtained by the current training. According to the following formula, the importance matrix Ω of each parameter for the knowledge learned by the model from all input data in the j-th task can be calculated j .

[0086]

[0087] Based on the above formula, the importance matrix Ω of the knowledge learned by the model from all input data in the j-th generated speech data can be determined j , and based on Ω j , the update amplitude of each model parameter in the (j + 1)-th can be determined. Correspondingly, for the training of the j-th generated speech data, the update amplitude of each model parameter during the training of the j-th generated speech data can be determined based on the importance matrix Ω j-1 of the knowledge learned by the model from all input data in the (j - 1)-th generated speech data.

[0088] In the embodiments of the present invention, after each completion of the learning of a generated speech dataset, an importance matrix can be calculated based on one or more generated speech datasets that have completed learning, so that the importance matrix can be continuously updated, and then the adaptive constraint on the update amplitude of the model parameters can be realized based on the continuously updated importance matrix.

[0089] In the embodiments of the present invention, after completing the learning of a generated speech dataset (i.e., completing a training task in the associated learning process of the model), the importance matrix can be determined based on all input data of all generated speech datasets that have completed training, or random sampling can be performed on the input data of all generated speech datasets that have completed training to obtain partial input data, and the importance matrix can be determined using this partial input data, or the importance matrix can be determined only based on all input data of the most recently completed training of a generated speech dataset.

[0090] Specifically, in the embodiments of the present invention, the update amplitude of the parameter is determined based on the following formula:

[0091]

[0092] where λ represents the constraint strength constant, ω represents the parameter values of each model parameter obtained after learning the j-th generated speech dataset, ω * represents the parameter values of each model parameter obtained after learning the (j - 1)-th generated speech dataset, ω - ω * represents the change amount of each model parameter of the speech discrimination source model during the process of learning the j-th generated speech dataset.

[0093] In the embodiments of the present invention, while correcting the model parameter update direction of the model on the second generated speech dataset, by introducing a constraint term, different update amplitudes are given to different model parameters. Specifically, the following formula can be used to update the parameter direction and amplitude correction of each model parameter of the speech discrimination source model:

[0094] W l (i,j) = W l (i - 1,j) + γ(i,j)P l ΔW l BP (i,j) + R;

[0095]

[0096] where W l (i,j) represents each model parameter in the l-th layer updated based on the i-th sub-dataset of the j-th generated speech dataset, γ represents the learning rate, P l represents the update direction of each model parameter in the l-th layer, ΔWl BP (i, j) represents the gradient calculated by using the classical backpropagation algorithm for the i-th sub-dataset based on the j-th generated speech dataset.

[0097] In an embodiment of the present invention, j is greater than 1. That is to say, the above formula is applied to the subsequent training process of the speech discrimination source model obtained after the model completes the training of the first generated speech dataset.

[0098] In an alternative embodiment of the present invention, the speech discrimination source model is trained by the following steps:

[0099] S1. Obtain the first generated speech dataset.

[0100] S2. Use the first generated speech dataset to train the deep learning model to obtain the speech discrimination source model.

[0101] In an embodiment of the present invention, in step S2, the parameters of the deep learning model can be updated based on the following formula:

[0102] W l (i, j) = W l (i - 1, j) + γ(i, j)ΔW l BP (i, j);

[0103] where, W l (i, j) represents each model parameter in the l-th layer updated based on the i-th sub-dataset of the j-th generated speech dataset, γ represents the learning rate, and ΔW l BP (i, j) represents the gradient calculated by using the classical backpropagation algorithm for the i-th sub-dataset based on the j-th generated speech dataset.

[0104] In an embodiment of the present invention, j is equal to 1. That is to say, the above formula is applied to the training process of the model learning the first generated speech dataset.

[0105] In a second aspect of the embodiments of the present invention, an audio discrimination device for adaptively correcting the direction and amplitude of parameters based on continuous learning is proposed. Refer to Figure 2 , Figure 2 which is a schematic architecture diagram of an audio discrimination device for adaptively correcting the direction and amplitude of parameters based on continuous learning proposed in an embodiment of the present invention. The device includes:

[0106] An acquisition module 201, configured to acquire a second generated speech dataset and a speech discrimination source model, where the speech discrimination source model can discriminate the generated speech in a first generated speech dataset, and the first generated speech dataset and the second generated speech dataset are obtained based on different speech generation methods;

[0107] An update module 202, configured to use the second generated speech dataset to update each model parameter of the speech discrimination source model by correcting the parameter direction and amplitude, so as to obtain an updated speech discrimination model, where the updated speech discrimination model can discriminate the generated speech in the first generated speech dataset and the generated speech in the second generated speech dataset;

[0108] Wherein, the update by correcting the parameter direction and amplitude includes the following steps:

[0109] Adjust the update direction of each model parameter;

[0110] Based on the importance of each model parameter to the overall model, constrain the update amplitude of each model parameter.

[0111] In an optional embodiment of the present invention, constraining the update amplitude of each model parameter based on the importance of each model parameter to the overall model includes:

[0112] Determine the importance of each model parameter to the overall model based on the following formula:

[0113]

[0114] Where, W j The importance of each model parameter determined based on the jth generated speech dataset, g i,j x l (i, j) represents the influence on the entire model parameter set during the change of each model parameter when the input is x l (i, j); Where, M(x l (i, j)) represents the output of the overall model M when the input is x l (i, j); ω represents the parameter value trained based on the jth generated speech dataset;

[0115] Determine the update amplitude of the parameter based on the following formula:

[0116]

[0117] Where, λ represents a constraint strength constant, ω represents the parameter value before update, and ω * represents the parameter value after update.

[0118] In an alternative embodiment of the present invention, adjusting the update direction of each model parameter includes: determining the update direction of each model parameter based on the following formula:

[0119]

[0120] where i represents the batch to which the sub-dataset belongs when updating the voice discrimination source model, j represents that the sub-dataset input when updating the voice discrimination source model is the j-th generated voice dataset in the generated voice dataset, α i,j represents a preset constant, T represents transpose, represents the average value of the input voice in the i-th sub-dataset of the j-th generated voice dataset in the (l-1)-th layer; P l (i, j) represents the update direction of the parameters in the l-th layer when updating the voice discrimination source model based on the i-th sub-dataset of the j-th generated voice dataset.

[0121] In an alternative embodiment of the present invention, updating the parameter direction and amplitude correction of each model parameter of the voice discrimination source model based on the following formula:

[0122] W l (i, j) = W l (i - 1, j) + γ(i, j)P l ΔW l BP (i, j) + R;

[0123] where W l (i, j) represents each model parameter in the l-th layer updated based on the i-th sub-dataset of the j-th generated voice dataset, γ represents the learning rate, P l represents the update direction of each model parameter in the l-th layer, and ΔW l BP (i, j) represents the gradient calculated by using the classical backpropagation algorithm based on the i-th sub-dataset of the j-th generated voice dataset.

[0124] In an alternative embodiment of the present invention, the voice discrimination source model is trained by the following steps:

[0125] Obtain the first generated voice dataset;

[0126] Train the deep learning model by using the first generated voice dataset to obtain the voice discrimination source model.

[0127] In an alternative embodiment of the present invention, training the deep learning model by using the first generated voice dataset includes: updating the parameters of the deep learning model based on the following formula:

[0128] W l (i,j) = W l (i - 1,j) + γ(i,j)ΔW l BP (i,j);

[0129] where W l (i,j) represents each model parameter in the l-th layer updated based on the i-th sub-dataset of the j-th generated speech dataset, γ represents the learning rate, and ΔW l BP (i,j) represents the gradient calculated using the classical backpropagation algorithm based on the i-th sub-dataset of the j-th generated speech dataset.

[0130] Based on the same inventive concept, an embodiment of the present invention discloses an electronic device, Figure 3 which shows a schematic diagram of an electronic device disclosed in an embodiment of the present invention, as Figure 3 shown, the electronic device 100 includes: a memory 110 and a processor 120. The memory of the electronic device is not less than 12G, the main frequency of the processor is not lower than 2.4GHz. The memory 110 and the processor 120 are communicatively connected via a bus. A computer program is stored in the memory 110, and this computer program can run on the processor 120 to implement an audio discrimination method for adaptively correcting the direction and amplitude of parameters based on continuous learning disclosed in an embodiment of the present invention.

[0131] Based on the same inventive concept, an embodiment of the present invention discloses a computer-readable storage medium, on which a computer program / instructions are stored. When the computer program / instructions are executed by a processor, they implement an audio discrimination method for adaptively correcting the direction and amplitude of parameters based on continuous learning disclosed in an embodiment of the present invention.

[0132] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.

[0133] Embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, apparatuses, electronic devices, and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate for implementing in the process Figure 1 each process or multiple processes and / or blocks Figure 1means for the functions specified in one or more boxes.

[0134] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one Figure 1 process or more processes and / or boxes Figure 1 means for the functions specified in one or more boxes.

[0135] These computer program instructions may also be loaded onto a computer or other programmable data processing terminal device, such that a series of operational steps are performed on the computer or other programmable terminal device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one Figure 1 process or more processes and / or boxes Figure 1 means for the functions specified in one or more boxes.

[0136] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.

[0137] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.

[0138] The above has introduced in detail a method, apparatus, device, and medium for audio discrimination with adaptive parameter direction and amplitude correction based on continuous learning. In this article, specific examples are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. An adaptive parameter direction and amplitude correction audio discrimination method based on continuous learning, characterized in that, The method includes: Obtaining a second generated speech dataset and a speech discrimination source model, where the speech discrimination source model can discriminate the generated speech in the first generated speech dataset, and the first generated speech dataset and the second generated speech dataset are obtained based on different speech generation methods; Using the second generated speech dataset to update the respective model parameters of the speech discrimination source model by correcting the parameter direction and amplitude, to obtain an updated speech discrimination model, where the updated speech discrimination model can discriminate the generated speech in the first generated speech dataset and the generated speech in the second generated speech dataset; Among them, the update by correcting the parameter direction and amplitude includes the following steps: Adjusting the update direction of each model parameter; Constraining the update amplitude of each model parameter based on the importance of each model parameter to the overall model; The constraining the update amplitude of each model parameter based on the importance of each model parameter to the overall model includes: Determining the importance of each model parameter to the overall model based on the following formula: Among them, Ω j The importance of each model parameter determined based on the j-th generated speech dataset to the overall model, g i,j x l (i, j) represents the influence on the entire model parameter set during the change of each model parameter when the input is x l (i, j); Among them, M(x l (i, j)) represents the output of the overall model M when the input is x l (i, j); ω represents the parameter values of each model parameter obtained after learning the j-th generated speech dataset; Determining the update amplitude of this parameter based on the following formula: Among them, λ represents the constraint strength constant, ω * represents the parameter values of each model parameter obtained after learning the (j - 1)-th generated speech dataset, ω - ω * represents the change amount of each model parameter of the speech discrimination source model during the process of learning the j-th generated speech dataset.

2. The adaptive parameter direction and amplitude correction audio discrimination method based on continuous learning according to claim 1, wherein Adjusting the update direction of each model parameter includes: determining the update direction of each model parameter based on the following formula: Among them, i represents the batch to which the sub-dataset belongs when updating the voice discrimination source model, j represents that the generated voice dataset where the input sub-dataset is located when updating the voice discrimination source model is the j-th generated voice dataset, α i,j represents a preset constant, T represents transpose, represents the average value of the input voice in the i-th sub-dataset of the j-th generated voice dataset in the (l - 1)-th layer; P l (i, j) represents the update direction of the parameters of the l-th layer when updating the voice discrimination source model based on the i-th sub-dataset of the j-th generated voice dataset.

3. The audio discrimination method for adaptively correcting the direction and amplitude of parameters based on continuous learning according to claim 2, wherein Updating the respective model parameters of the speech discrimination source model by correcting the parameter direction and amplitude based on the following formula: W l (i, j) = W l (i - 1, j) + γ(i, j)P l ΔW l BP (i, j) + R; Among them, W l (i,j) represents each model parameter in the l-th layer updated based on the i-th sub-dataset of the j-th generated speech dataset, γ represents the learning rate, and P l represents the update direction of each model parameter in the l-th layer, and ΔW l BP (i,j) represents the gradient calculated by using the classical backpropagation algorithm based on the i-th sub-dataset of the j-th generated speech dataset.

4. The adaptive parameter direction and amplitude correction audio discrimination method based on continuous learning according to claim 1, wherein The speech discrimination source model is trained by the following steps: Obtaining a first generated speech dataset; Using the first generated speech dataset to train a deep learning model to obtain a speech discrimination source model.

5. The adaptive parameter direction and amplitude correction audio discrimination method based on continuous learning according to claim 4, characterized in that Using the first generated speech dataset to train a deep learning model includes: updating the parameters of the deep learning model based on the following formula: W l (i,j) = W l (i - 1,j) + γ(i,j)ΔW l BP (i,j); Among them, W l (i, j) represents each model parameter in the l-th layer updated based on the i-th sub-dataset of the j-th generated speech dataset, γ represents the learning rate, and ΔW l BP (i, j) represents the gradient calculated by using the classical backpropagation algorithm based on the i-th sub-dataset of the j-th generated speech dataset.

6. An audio discrimination device for adaptively correcting the direction and amplitude of parameters based on continuous learning, characterized in that, The device includes: An obtaining module, configured to obtain a second generated speech dataset and a speech discrimination source model, where the speech discrimination source model can discriminate the generated speech in the first generated speech dataset, and the first generated speech dataset and the second generated speech dataset are obtained based on different speech generation methods; An updating module, configured to use the second generated speech dataset to update the respective model parameters of the speech discrimination source model by correcting the parameter direction and amplitude, to obtain an updated speech discrimination model, where the updated speech discrimination model can discriminate the generated speech in the first generated speech dataset and the generated speech in the second generated speech dataset; Among them, the update by correcting the parameter direction and amplitude includes the following steps: Adjusting the update direction of each model parameter; Constraining the update amplitude of each model parameter based on the importance of each model parameter to the overall model; The constraining the update amplitude of each model parameter based on the importance of each model parameter to the overall model includes: Determining the importance of each model parameter to the overall model based on the following formula: where, Ω j The importance of each model parameter determined based on the j-th generated speech dataset to the overall model, g i,j x l (i, j) represents the influence on the entire model parameter set during the change of each model parameter when the input is x l (i, j); where, M(x l (i, j)) represents the output of the overall model M when the input is x l (i, j); ω represents the parameter values of each model parameter obtained after learning the j-th generated speech dataset; Determining the update amplitude of this parameter based on the following formula: Among them, λ represents the constraint strength constant, ω * represents the parameter values of each model parameter obtained after learning the (j - 1)-th generated speech dataset, ω - ω * represents the change amount of each model parameter of the speech discrimination source model during the process of learning the j-th generated speech dataset.

7. An electronic device, characterized in that, Includes: A memory, configured to store one or more programs; A processor; When the one or more programs are executed by the processor, implementing the adaptive parameter direction and amplitude correction audio discrimination method based on continuous learning as described in any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the audio discrimination method for adaptively correcting the direction and amplitude of parameters based on continuous learning as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Continuous learning voice identification model training method and device, equipment and medium

    CN117577116A

  • Method and device for training voice detection model of orthogonalization low-rank adaptation matrix

    CN117577117A