Model training method, voiceprint enhancement method, electronic equipment and computer medium
By extracting and analyzing the voiceprint characteristics of the target data and interfering data, the voiceprint enhancement model is trained based on the correlation degree, and the problem of target speech being filtered is solved and a better voice enhancement effect is achieved.
Patent Information
- Application Number
- CN202410097858.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-23
- Publication Date
- 2025-07-25
AI Technical Summary
The existing voiceprint enhancement method can easily filter out the target voice when the correlation between the target voice and the interfering voice is high, resulting in severe speech distortion after enhancement, reduced voice coherence, and poor user hearing.
By extracting the voiceprint characteristics of the target data and the interference data, performing correlation analysis, obtaining the label data of the training data set based on the correlation, training the voiceprint enhancement model, selectively fusing or removing the interference data, and calculating the loss function to optimize the model.
It improves the coherence of voice, reduces the hearing difference and distortion caused by excessive suppression of interference data, and improves the voiceprint enhancement effect.
Smart Images

Figure CN120375831A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice enhancement technologies, and particularly to a model training method, a voiceprint enhancement method, an electronic device, and a computer medium. Background Art
[0002] Voiceprint technology is a biometric technology based on voice features. Each person's voice has unique features, which are determined by the body structures such as the vocal cords, oral cavity, and nasal cavity of the human body, and this voice feature is the voiceprint feature. In the field of voice enhancement, when using an electronic device to process voice, voiceprint technology and voice enhancement technology are usually combined to achieve the enhancement of personal voice and improve the voice quality and clarity. For example, personal voice enhancement can be achieved by eliminating interference data such as background noise and echo, and this technology can be applied to many life scenarios, including telephone conferences, speech recognition, hearing aids, and multimedia content production.
[0003] Since the distinguishability of voiceprint features of different people is not large, especially the voiceprint features of people of the same gender and the same age group are highly correlated, it is difficult to distinguish the surrounding interfering voices from the target voice, and it is easy to have a situation where the voiceprint features of the interfering voice and the target voice are similar. In the process of voice enhancement, existing voiceprint enhancement methods usually use the target voice as the learning target of the model. When the correlation between the target voice and the interfering voice is high, the target voice in the audio data is easily filtered out, resulting in serious distortion of the enhanced voice, reduced coherence of the voice, poor listening experience of the user, and unsatisfactory enhancement effect. Summary of the Invention
[0004] To solve the above technical problems, this application provides a model training method, a voiceprint enhancement method, an electronic device, and a computer medium.
[0005] To solve the above problems, this application provides a first technical solution: providing a model training method applied to a voiceprint enhancement model, where the model training method includes: obtaining a training data set, where the training data set includes target data and interference data corresponding to the target data; extracting a first voiceprint feature from the target data, and extracting a second voiceprint feature from the interference data; performing a correlation analysis on the first voiceprint feature and the second voiceprint feature to obtain the correlation between the first voiceprint feature and the second voiceprint feature; and obtaining label data of the training data set based on the correlation, the interference data, and the target data, so as to train the voiceprint enhancement model based on the label data.
[0006] Optionally, the above training data set further includes first speech data to be enhanced; obtaining the label data of the above training data set based on the above relevance, the above interference data, and the above target data, and training the above voiceprint enhancement model based on the above label data, including: fusing the above interference data and the above target data based on the above relevance, and using the fused data as the above label data; inputting the above first speech data and the above first voiceprint feature into the above voiceprint enhancement model to obtain enhanced second speech data; calculating the loss function of the above voiceprint enhancement model based on the above fused data and the above second speech data, and training the above voiceprint enhancement model based on the above loss function.
[0007] Optionally, the above fusing the above interference data and the above target data based on the above relevance, and using the fused data as the above label data, includes: calculating the gain coefficient of the above interference data based on the above relevance; calculating the product of the above interference data and the above gain coefficient, and calculating the sum of the above target data and the above product, and using the above sum as the above label data.
[0008] Optionally, the above calculating the gain coefficient of the above interference data based on the above relevance includes: in response to the above relevance being greater than or equal to a first preset threshold, determining the above gain coefficient to be 1; in response to the above relevance being less than or equal to a second preset threshold, determining the above gain coefficient to be 0; in response to the above relevance being greater than the above second preset threshold and less than the above first preset threshold, obtaining the above gain coefficient corresponding to the above relevance based on the mapping relationship between the above relevance and the above gain coefficient; wherein, the above second preset threshold is less than the above first preset threshold.
[0009] Optionally, the above voiceprint enhancement model is used to extract features from the above first speech data corresponding to the above target data, fuse the voiceprint features of the extracted above first speech data and the above label data, and use the data after feature fusion as the above second speech data.
[0010] Optionally, the above calculating the loss function of the above voiceprint enhancement model based on the above fused data and the above second speech data, and training the above voiceprint enhancement model based on the above loss function, includes: calculating the above loss function based on the difference between the above label data and the above second speech data; updating the model parameters of the above voiceprint enhancement model based on the gradient of the above loss function to train the above voiceprint enhancement model.
[0011] To solve the above problems, the present application provides a second technical solution: providing a voiceprint enhancement method, which is applied to a voiceprint enhancement model, including: obtaining audio data to be enhanced, extracting a third voiceprint feature corresponding to target data from the above audio data, and extracting a fourth voiceprint feature corresponding to interference data; performing a correlation analysis on the above third voiceprint feature and the above fourth voiceprint feature to obtain the correlation degree between the above third voiceprint feature and the above fourth voiceprint feature; based on the above correlation degree, the above interference data, and the above target data, enhancing the above audio data to output the enhanced above audio data.
[0012] Optionally, the above enhancing the above audio data based on the above correlation degree, the above interference data, and the above target data to output the enhanced above audio data includes: calculating a gain coefficient of the above interference data based on the above correlation degree; calculating the product of the above interference data and the above gain coefficient; performing feature fusion on the sum of the above target data and the above product and the above audio data to obtain the enhanced above audio data.
[0013] To solve the above problems, the present application provides a third technical solution: providing an electronic device, the above electronic device includes a processor and a memory connected to the above processor, wherein the above memory stores program instructions; the above processor is used to execute the program instructions stored in the above memory to implement the above method.
[0014] To solve the above problems, the present application provides a fourth technical solution: providing a computer-readable storage medium, the above computer-readable storage medium stores program instructions, and the above program instructions can be executed by a processor to implement the above method.
[0015] The present application provides a model training method, a voiceprint enhancement method, an electronic device, and a computer medium. The model training method extracts a first voiceprint feature from target data and a second voiceprint feature from interference data; performs a correlation analysis on the first voiceprint feature and the second voiceprint feature to obtain the correlation degree between the first voiceprint feature and the second voiceprint feature; based on the correlation degree, interference data, and target data, obtains label data of a training data set to train a voiceprint enhancement model based on the label data. In the above manner, the model training method can obtain label data based on the correlation degree between the first voiceprint feature and the second voiceprint feature, so that the trained voiceprint enhancement model can partially retain or remove interference data based on the correlation degree when enhancing voice data, reducing situations such as poor listening perception and serious distortion caused by excessive suppression of interference data, improving the coherence of the voice, and thus improving the voiceprint enhancement effect. Description of the Drawings
[0016] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings. Among them:
[0017] Figure 1 is a schematic flowchart of the first embodiment of the model training method provided by the present application;
[0018] Figure 2 is a schematic flowchart of the second embodiment of the model training method provided by the present application;
[0019] Figure 3 is a schematic flowchart of the third embodiment of the model training method provided by the present application;
[0020] Figure 4 is a schematic structural diagram of the first embodiment of the electronic device provided by the present application;
[0021] Figure 5 is a schematic structural diagram of the second embodiment of the electronic device provided by the present application;
[0022] Figure 6 is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by the present application. Detailed implementation manners
[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0024] It should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present application, the directional indications are only used to explain the relative position relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.
[0025] In addition, if the embodiments of the present application involve descriptions such as "first" and "second", the descriptions of "first", "second", etc. are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments may be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.
[0026] Existing voiceprint enhancement methods are usually applied in proximal scenarios, such as near-field devices like headphones and hearing aids. As the distance increases, the degree of fuzziness of the voiceprint features of the target speech increases. Moreover, when the environmental noise increases or the voiceprint features of the target speech change, it is easy to have a situation where the correlation between the voiceprint features of the target speech and the interfering speech is high, which causes the voiceprint enhancement model to easily suppress the voiceprint features of the target speech as interfering speech, resulting in serious speech distortion.
[0027] In view of this, the embodiments of the present application propose a model training method. This model training method is applied to a voiceprint enhancement model, which is used to perform voiceprint recognition on the speaker in the voice data and enhance the voiceprint of the target speech in the voice data through the recognized voiceprint features to reduce interfering speech and improve the speech quality. The voiceprint enhancement method implemented by this voiceprint enhancement model can be applied to electronic devices, which can be, but are not limited to, electronic products such as mobile phones, computers, headphones, hearing aids, and speaker horns.
[0028] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the model training method provided by the present application. As Figure 1 shown, the model training method of this embodiment includes the following steps:
[0029] Step S11: Obtain a training data set, where the training data set includes target data and interfering data corresponding to the target data.
[0030] Specifically, the training dataset can be obtained from multiple standard test databases, or can be obtained by using a voice acquisition device to collect voices in relevant test areas. The training dataset includes at least first voice data, target data, and interference data corresponding to the target data. Among them, the first voice data is the voice data to be enhanced. For example, when the voiceprint enhancement model is applied to devices such as mobile phones, computers, headphones, hearing aids, and speaker speakers, the first voice data can be the voice data generated during a voice call. There is usually a preset speaker in the first voice data. The target data is voice data with relatively prominent voiceprint features of the same speaker, and the interference data is other interfering voices except this speaker. For example, the interference data can be, but is not limited to, human voice data other than the speaker, environmental noise data, etc.
[0031] Step S12: Extract the first voiceprint feature from the target data and the second voiceprint feature from the interference data.
[0032] After obtaining the training dataset, the first voiceprint feature is extracted from the target data through a voiceprint recognition algorithm, and this first voiceprint feature is used to represent the voice feature of the speaker; and, the second voiceprint feature is also extracted from the interference data through the voiceprint recognition algorithm, and the second voiceprint feature is used to represent the feature of the environmental sound other than the speaker.
[0033] Step S13: Perform a correlation analysis on the first voiceprint feature and the second voiceprint feature to obtain the correlation degree between the first voiceprint feature and the second voiceprint feature.
[0034] Perform a correlation analysis on the first voiceprint feature and the second voiceprint feature. Exemplarily, it can be done by comparing the first voiceprint feature and the second voiceprint feature, or by calculating the correlation coefficient between the first voiceprint feature and the second voiceprint feature, to obtain the correlation degree between the first voiceprint feature and the second voiceprint feature.
[0035] Among them, the value of the correlation degree can range from 0 to 1. When the correlation degree between the first voiceprint feature and the second voiceprint feature is 0, that is, the first voiceprint feature and the second voiceprint feature are completely uncorrelated, and the voice feature of the speaker and the environmental sound feature are completely different. At this time, the voiceprint enhancement model can easily separate the target voice from the interference voice in the first voice data; when the correlation degree between the first voiceprint feature and the second voiceprint feature is 1, that is, the first voiceprint feature and the second voiceprint feature are completely correlated, at this time, the voiceprint enhancement model is very difficult to separate the target voice from the interference voice in the first voice data.
[0036] Step S14: Based on the correlation degree, interference data, and target data, obtain the label data of the training dataset, and train the voiceprint enhancement model based on the label data.
[0037] Specifically, after obtaining the correlation between the first voiceprint feature and the second voiceprint feature, the interference data and the target data can be fused based on the correlation to obtain the labeled data of the training dataset, which is related to the correlation, the interference data, and the target data. Exemplarily, the interference data and the target data can be selectively fused based on the level of the correlation. For example, when the correlation is high, at least part of the interference data and the target data can be added together, and the sum value can be used as the labeled data to reduce the poor listening experience and severe distortion caused by excessive suppression of the interference data; or, when the correlation is low, the target data can be directly used as the labeled data, so that the voiceprint enhancement model can easily separate the target voice from the interference voice in the first voice data and reduce noise interference.
[0038] In the embodiment of the present application, this model training method can obtain labeled data based on the correlation between the first voiceprint feature and the second voiceprint feature, so that the trained voiceprint enhancement model can partially retain or remove the interference data based on the correlation when enhancing the voice data, reduce the poor listening experience and severe distortion caused by excessive suppression of the interference data, improve the coherence of the voice, and further improve the voiceprint enhancement effect.
[0039] In one embodiment, please refer to Figure 2 , Figure 2 which is a schematic flowchart of the second embodiment of the model training method provided by the present application. As Figure 2 shown, the training dataset further includes the first voice data to be enhanced. Step S14 further includes the following steps:
[0040] Step S21: Fuse the interference data and the target data based on the correlation, and use the fused data as the labeled data.
[0041] After obtaining the correlation between the first voiceprint feature and the second voiceprint feature, fuse the interference data and the target data based on the correlation, and use the fused data as the labeled data. Among them, the model training method of this embodiment determines the degree of fusing the interference data into the target data based on the correlation. For example, when the correlation is high, more interference data is fused into the target data; when the correlation is low, less interference data is fused into the target data, or the interference data is not retained and the target data is directly used as the labeled data.
[0042] Step S22: Input the first voice data and the first voiceprint feature into the voiceprint enhancement model to obtain the enhanced second voice data.
[0043] Input the first speech data to be enhanced in the training dataset into the voiceprint enhancement model. The voiceprint enhancement model synchronously obtains the first voiceprint feature of the same speaker in the first speech data. The voiceprint enhancement model enhances the voiceprint of the first speech data based on the first voiceprint feature to obtain the enhanced second speech data. In a possible implementation, the first speech data to be enhanced may be audio data obtained by mixing the target data and the interference data according to the signal-to-interference ratio. The first speech data may also be obtained by collecting the speech of the same speaker, which is not specifically limited herein.
[0044] Step S23: Calculate the loss function of the voiceprint enhancement model based on the fused data and the second speech data, and train the voiceprint enhancement model based on the loss function.
[0045] After obtaining the enhanced second speech data, calculate the loss function of the voiceprint enhancement model based on the fused data (i.e., the label data) and the second speech data, and iteratively train the voiceprint enhancement model based on the calculated loss function until the loss function is minimized.
[0046] In the embodiment of the present application, the model training method fuses the interference data and the target data based on the correlation, uses the fused data as the label data, and calculates the loss function of the voiceprint enhancement model through the label data and the second speech data, and trains the voiceprint enhancement model based on the loss function, so that the trained voiceprint enhancement model can partially retain or remove the interference data based on the correlation when enhancing the speech data, reducing the poor listening experience and serious distortion caused by excessive suppression of the interference data, improving the coherence of the speech, and thus improving the voiceprint enhancement effect.
[0047] Optionally, step S21 includes the following steps: calculate the gain coefficient of the interference data based on the correlation; calculate the product of the interference data and the gain coefficient, and calculate the sum of the target data and the product to use the sum as the label data.
[0048] Specifically, after obtaining the correlation between the first voiceprint feature and the second voiceprint feature, the gain coefficient of the interference data can be calculated based on the correlation. The gain coefficient represents the degree of fusing the interference data into the target data. Among them, there is a mapping relationship between the correlation and the gain coefficient. After obtaining the correlation, the corresponding gain coefficient can be directly obtained based on the mapping relationship. This mapping relationship can be preset in the memory through an algorithm or obtained through other algorithm models, which is not specifically limited herein.
[0049] After obtaining the gain coefficient, the model training method of this embodiment calculates the product of the interference data and the gain coefficient, and calculates the sum of the target data and the product to determine the label data, so that the model training method can use the label data as the learning target of the voiceprint enhancement model for training. The trained voiceprint enhancement model can select the gain coefficient of the interference data based on the relevance when enhancing the speech data, reduce the poor listening experience and severe distortion caused by excessive suppression of the interference data, improve the coherence of the speech, and thus improve the voiceprint enhancement effect.
[0050] For further information, see Figure 3 , Figure 3 is a flow chart of the third embodiment of the model training method provided by this application. Figure 3 As shown, the step of calculating the gain coefficient of the interference data based on the correlation also includes:
[0051] Step S31: In response to the correlation being greater than or equal to a first preset threshold, determining the gain coefficient to be 1.
[0052] Specifically, after obtaining the correlation between the first voiceprint feature and the second voiceprint feature, when the correlation is greater than or equal to the first preset threshold, that is, the correlation between the first voiceprint feature and the second voiceprint feature is high, the gain coefficient is determined to be 1. At this time, the model training method completely retains the interference data and uses it as label data, so that the trained voiceprint enhancement model does not suppress the interference data, so as to ensure the integrity of the target voice in the first voice data and reduce distortion.
[0053] Step S32: In response to the correlation being less than or equal to the second preset threshold, determining the gain coefficient to be 0.
[0054] Specifically, the second preset threshold is less than the first preset threshold. When the correlation is less than or equal to the second preset threshold, that is, the correlation between the first voiceprint feature and the second voiceprint feature is low, the gain coefficient is determined to be 0. At this time, the model training method does not retain the interference data at all and uses the target data as label data, so that the trained voiceprint enhancement model completely suppresses the interference data, so as to minimize the interference voice in the first voice data and improve the hearing of the target voice.
[0055] Step S33: in response to the correlation being greater than the second preset threshold and the correlation being less than the first preset threshold, acquiring a gain coefficient corresponding to the correlation based on a mapping relationship between the correlation and the gain coefficient.
[0056] Specifically, when the correlation is greater than the second preset threshold and less than the first preset threshold, that is, the first voiceprint feature and the second voiceprint feature have a certain correlation. At this time, the model training method can obtain the corresponding gain coefficient based on the mapping relationship between the correlation and the gain coefficient, so as to fuse part of the interference data into the target data based on the gain coefficient and obtain the labeled data. Through the above method, the trained voiceprint enhancement model can retain as much interference speech related to the target speech as possible, and eliminate as much interference speech unrelated to the target speech as possible, reducing the poor listening experience and serious distortion caused by excessive suppression of interference data, improving the coherence of the speech, and thus improving the voiceprint enhancement effect.
[0057] In a possible implementation, the first preset threshold can be a preset value close to 1 such as 0.7, 0.8, 0.9, etc., and the second preset threshold can be a value close to 0 such as 0.1, 0.2, 0.3, etc., which are not specifically limited herein.
[0058] Optionally, the voiceprint enhancement model is used to extract features from the first voice data corresponding to the target data, so as to perform feature fusion on the voiceprint features of the extracted first voice data and the labeled data, and use the data after feature fusion as the second voice data.
[0059] Specifically, the voiceprint enhancement model can be embedded with multiple layers of algorithm networks. The voiceprint enhancement model is used to receive the first voice data and also used to receive the voiceprint features of the labeled data. After obtaining the first voice data, the voiceprint enhancement model extracts and encodes the features of the first voice data, so that the features of the encoded first voice data can be on the same scale as the first voiceprint feature. Among them, the voiceprint enhancement model can, but is not limited to, extract the features of the first voice data at the levels of acoustics, morphology, prosody, language, accent, channel, etc.
[0060] After the voiceprint enhancement model extracts and encodes the features of the first voice data, it performs feature fusion on the voiceprint features of the first voice data and the labeled data, and performs feature decoding on the data after feature fusion to obtain the enhanced second voice data. Among them, the feature fusion can be in ways such as feature splicing, deep fusion, feature parameter mixing, etc., which are not specifically limited herein.
[0061] In the embodiment of the present application, the voiceprint enhancement model extracts and encodes the features of the first voice data, so that the voiceprint features of the first voice data and the labeled data can be fused on the time-frequency domain buto can scale, and after passing through the decoding network for feature transformation, the enhanced second voice data is output, realizing the voice enhancement of the target voiceprint, reducing the distortion degree of the voice data, and improving the readability of the voice data.
[0062] Optionally, step S23 includes the following steps: calculating a loss function based on the difference between the tag data and the second voice data; and updating the model parameters of the voiceprint enhancement model based on the gradient of the loss function to train the voiceprint enhancement model.
[0063] Specifically, after determining the fused data as the tag data, a loss function is calculated based on the difference between the tag data and the second voice data. The loss function is used to provide feedback to the voiceprint enhancement model, enabling the model training method to calculate the gradient based on the loss function and adjust or update the model parameters of the voiceprint enhancement model in a gradient descent manner to minimize errors and improve the training effect of the voiceprint enhancement model.
[0064] An embodiment of the present application also proposes a voiceprint enhancement method. The voiceprint enhancement method inputs the audio data to be enhanced into a voiceprint enhancement model to enhance the voiceprint of the audio data with the voiceprint enhancement data.
[0065] Specifically, the steps for the voiceprint enhancement model to perform voiceprint enhancement include: obtaining the audio data to be enhanced, extracting the third voiceprint feature corresponding to the target data from the audio data, and extracting the fourth voiceprint feature corresponding to the interference data; performing a correlation analysis on the third voiceprint feature and the fourth voiceprint feature to obtain the correlation degree between the third voiceprint feature and the fourth voiceprint feature; and enhancing the audio data based on the correlation degree, the interference data, and the target data to output the enhanced audio data.
[0066] After obtaining the audio data, feature extraction is performed on the audio data to extract the third voiceprint feature corresponding to the target data and the fourth voiceprint feature corresponding to the interference data from the audio data. Among them, the voiceprint enhancement model can be applied to an audio device. Exemplarily, the audio device can be used in application scenarios such as meetings, offices, live broadcasts, and social interactions, and is used to collect and transmit the audio data of the speaker; for example, the audio data can be the audio data generated in a conference device or a live broadcast device. It can be understood that the third voiceprint feature is a voice feature similar to or close to the speaker, and the fourth voiceprint feature is the voice feature other than the speaker in the speaking environment.
[0067] The third voiceprint feature and the fourth voiceprint feature are subjected to correlation analysis and the correlation degree is obtained. The value of the correlation degree can be between 0 and 1. When the correlation degree is 0, the third voiceprint feature and the fourth voiceprint feature are completely unrelated, and the voice feature of the speaker and the sound feature of the environment are completely different; when the correlation degree is 1, the third voiceprint feature and the fourth voiceprint feature are completely related; when the correlation degree is greater than 0 and less than 1, the third voiceprint feature and the fourth voiceprint feature have a certain correlation. Therefore, based on the correlation, the voiceprint features of the interference data and the target data can be selectively fused with the audio data to enhance the audio data, and the enhanced audio data can be output to a preset location. For example, when the audio data is input by the speaker through a conference device or a live broadcast device, the enhanced audio data can also be broadcast through the conference device or the live broadcast device.
[0068] In this embodiment, the voiceprint enhancement model enhances the audio data based on the correlation between the third voiceprint feature and the fourth voiceprint feature, so that the enhanced audio data will retain the interfering voice with high correlation with the target voice as much as possible, and filter out the interfering voice with low correlation with the target voice, thereby improving the continuity and listening experience of the audio data, and thus improving the user experience.
[0069] Optionally, the above-mentioned step of enhancing the audio data based on the correlation, interference data and target data to output the enhanced audio data includes: calculating the gain coefficient of the interference data based on the correlation; calculating the product of the interference data and the gain coefficient; and performing feature fusion of the sum of the target data and the product with the audio data to obtain the enhanced audio data.
[0070] Specifically, after obtaining the correlation between the third voiceprint feature and the fourth voiceprint feature, the gain coefficient of the interfering speech is calculated based on the correlation, and the gain coefficient indicates the degree of fusion of the interfering data into the target data, or the gain coefficient is used to indicate the gain effect of the interfering data on the target data. After obtaining the gain coefficient, the product of the interfering data and the gain coefficient is calculated, and the sum of the target data and the product is feature fused with the audio data to obtain the enhanced audio data. The voiceprint enhancement model further features the fused voice data with the audio data and realizes the enhancement of the voiceprint features.
[0071] Therefore, the voiceprint enhancement method of this embodiment takes into account the influence of the correlation between the target speech and the interfering speech, so that the voiceprint enhancement method of this embodiment can improve the hearing experience in the case where the voiceprint features are blurred due to the increase in distance. The voiceprint enhancement method of this embodiment can be used in remote scenarios, thereby improving the scope of application of the voiceprint enhancement method.
[0072] See also Figure 4 , Figure 4It is a schematic structural diagram of a first embodiment of the electronic device provided by this application. As Figure 4 shown, the electronic device of this embodiment includes a mutually connected memory 52 and a processor 51. The memory 52 is used to store program instructions for implementing the method described in any of the above embodiments. The processor 51 is used to execute the program instructions stored in the memory 52.
[0073] Among them, the processor 51 can also be called a CPU (Central Processing Unit, central processing unit). The processor 51 may be an integrated circuit chip with signaling processing capabilities. The processor 51 can also be a general-purpose processor, a digital signaling processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0074] The memory 52 can be a memory stick, a TF card, etc., and can store all the information in the electronic device. All the input original data, computer programs, intermediate operation results, and final operation results are stored in the memory. It stores and retrieves information according to the positions specified by the controller. With the memory, the string matching prediction device has a memory function and can ensure normal operation. According to the purpose, the memory of the string matching prediction device can be divided into a main memory (internal memory) and an auxiliary memory (external memory), and there is also a classification method of dividing it into an external memory and an internal memory. The external memory is usually a magnetic medium or an optical disc, etc., and can store information for a long time. The internal memory refers to the storage component on the motherboard, which is used to store the data and programs being currently executed, but only temporarily stores the programs and data. When the power is turned off or cut off, the data will be lost.
[0075] Please refer to Figure 5 , Figure 5 It is a schematic structural diagram of a second embodiment of the electronic device provided by this application. As Figure 5 shown, the electronic device of this embodiment includes an acquisition module 53, an extraction module 54, an analysis module 55, and a training module 56.
[0076] The acquisition module 53 is used to acquire a training data set, and the training data set includes target data and interference data corresponding to the target data. The extraction module 54 is used to extract first voiceprint features from the target data and second voiceprint features from the interference data. The analysis module 55 is used to perform a correlation analysis on the first voiceprint features and the second voiceprint features to obtain the correlation degree between the first voiceprint features and the second voiceprint features. The training module 56 is used to obtain label data of the training data set based on the correlation degree, the interference data, and the target data, so as to train the voiceprint enhancement model based on the label data.
[0077] In several embodiments provided by the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0078] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0079] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0080] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a system server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in each embodiment of the present application.
[0081] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by the present application. As Figure 6As shown in the figure, the computer-readable storage medium of the present application stores a program file 61 that can implement all of the above methods. Among them, the program file 61 can be stored in the above storage medium in the form of a software product, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage device includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, or electronic devices such as computers, servers, mobile phones, and tablets.
[0082] The above description is only for the embodiments of the present application and does not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A model training method, characterized in that, Applied to a voiceprint enhancement model, the model training method includes: Obtain a training data set, where the training data set includes target data and interference data corresponding to the target data; Extract first voiceprint features from the target data and second voiceprint features from the interference data; Conduct a correlation analysis on the first voiceprint features and the second voiceprint features to obtain the correlation degree between the first voiceprint features and the second voiceprint features; Based on the correlation degree, the interference data, and the target data, obtain the label data of the training data set, and train the voiceprint enhancement model based on the label data.
2. The model training method according to claim 1, wherein The training data set further includes first voice data to be enhanced; the obtaining of the label data of the training data set based on the correlation degree, the interference data, and the target data, and training the voiceprint enhancement model based on the label data includes: Fuse the interference data and the target data based on the correlation degree, and use the fused data as the label data; Input the first voice data and the first voiceprint features into the voiceprint enhancement model to obtain enhanced second voice data; Calculate the loss function of the voiceprint enhancement model based on the fused data and the second voice data, and train the voiceprint enhancement model based on the loss function.
3. The model training method according to claim 2, wherein The fusing of the interference data and the target data based on the correlation degree to use the fused data as the label data includes: Calculate the gain coefficient of the interference data based on the correlation degree; Calculate the product of the interference data and the gain coefficient, and calculate the sum value of the target data and the product, and use the sum value as the label data.
4. The model training method according to claim 3, characterized in that The calculating of the gain coefficient of the interference data based on the correlation degree includes: In response to the correlation degree being greater than or equal to a first preset threshold, determine the gain coefficient to be 1; In response to the correlation degree being less than or equal to a second preset threshold, determine the gain coefficient to be 0; In response to the correlation degree being greater than the second preset threshold and less than the first preset threshold, obtain the gain coefficient corresponding to the correlation degree based on the mapping relationship between the correlation degree and the gain coefficient; Wherein, the second preset threshold is less than the first preset threshold.
5. The model training method according to claim 2, wherein The voiceprint enhancement model is used to extract features from the first voice data corresponding to the target data, perform feature fusion on the extracted first voice data and the voiceprint features of the label data, and use the data after feature fusion as the second voice data.
6. The model training method according to claim 2, wherein The calculating of the loss function of the voiceprint enhancement model based on the fused data and the second voice data, and training the voiceprint enhancement model based on the loss function includes: Calculate the loss function based on the difference between the label data and the second voice data; Update the model parameters of the voiceprint enhancement model based on the gradient of the loss function to train the voiceprint enhancement model.
7. A voiceprint enhancement method, characterized in that, Applied to a voiceprint enhancement model, includes: Obtain the audio data to be enhanced, extract the third voiceprint feature corresponding to the target data from the audio data, and extract the fourth voiceprint feature corresponding to the interference data; Perform a correlation analysis on the third voiceprint feature and the fourth voiceprint feature to obtain the correlation degree between the third voiceprint feature and the fourth voiceprint feature; Enhance the audio data based on the correlation degree, the interference data, and the target data to output the enhanced audio data.
8. The voiceprint enhancement method according to claim 7, wherein The enhancing the audio data based on the correlation degree, the interference data, and the target data to output the enhanced audio data includes: Calculate the gain coefficient of the interference data based on the correlation degree; Calculate the product of the interference data and the gain coefficient; Perform feature fusion on the sum of the target data and the product and the audio data to obtain the enhanced audio data.
9. An electronic device, characterized in that, The electronic device includes a processor and a memory connected to the processor, wherein, The memory stores program instructions; The processor is configured to execute the program instructions stored in the memory to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions, and the program instructions can be executed by a processor to implement the method according to any one of claims 1 to 7.