Audio recognition model training method, device, storage medium and electronic device

By introducing uncertainty analysis models to screen prediction results with high confidence in audio recognition model training, combined with multiple rounds of training optimization models, the problem of low training efficiency in the existing technology is solved, and efficient audio recognition model training is achieved.

CN113763934BActive Publication Date: 2025-08-19TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110593500.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-28
Publication Date
2025-08-19
Estimated Expiration
2041-05-28

Smart Images

  • Figure CN113763934B_ABST
    Figure CN113763934B_ABST
Patent Text Reader

Abstract

The present invention discloses a training method, device, storage medium and electronic device for an audio recognition model. The method comprises: using a first training sample set to train an audio recognition model to obtain an initial audio recognition model; inputting the audio features of a second group of audio samples into the initial audio recognition model to obtain a first group of predicted audio recognition results; inputting the audio features of the second group of audio samples into an uncertainty analysis model to obtain a first group of uncertainty analysis results; and based on the first group of uncertainty analysis results, screening out a second group of predicted audio recognition results whose credibility meets a preset condition from the first group of predicted audio recognition results. The above method can also be applied in artificial intelligence scenarios, and specifically also involves technologies such as speech recognition and machine learning. The present invention solves the technical problem of low training efficiency of audio recognition models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computers, and in particular to a training method, device, storage medium, and electronic device for an audio recognition model. Background Art

[0002] In recent years, the application of audio recognition technology has become more and more extensive, such as in the fields of oral assessment and security testing, but how to improve the accuracy of audio recognition is still a topic under research.

[0003] In related technologies, audio recognition model training is often used to improve audio recognition accuracy. However, this training process often relies on a large amount of manually labeled sample data. In other words, when there is less manually labeled sample data, the training effect of the audio recognition model is often difficult to guarantee.

[0004] However, manual labeling itself consumes a lot of manpower and material resources. Furthermore, obtaining a large amount of manually labeled sample data not only incurs high labor and material costs but also requires long waiting times, thus reducing the training efficiency of audio recognition models. In other words, the existing technology suffers from the technical problem of low audio recognition model training efficiency.

[0005] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0006] Embodiments of the present invention provide a method, apparatus, storage medium, and electronic device for training an audio recognition model, so as to at least solve the technical problem of low training efficiency of the audio recognition model.

[0007] According to one aspect of an embodiment of the present invention, a method for training an audio recognition model is provided, comprising: using a first training sample set to train an audio recognition model to be trained to obtain an initial audio recognition model, wherein the first training sample set includes a first group of audio samples and a first group of actual audio recognition results obtained by annotating the first group of audio samples, and the initial audio recognition model is used to determine predicted audio recognition based on input audio features; inputting audio features of a second group of audio samples into the initial audio recognition model to obtain a first group of predicted audio recognition results, wherein the second group of audio samples are not annotated with corresponding actual audio recognition results; and inputting audio features of the second group of audio samples into the initial audio recognition model to obtain a first group of predicted audio recognition results. to the uncertainty analysis model to obtain a first set of uncertainty analysis results, wherein the above-mentioned first set of uncertainty analysis results is used to represent the credibility of the above-mentioned first set of predicted audio recognition results; based on the above-mentioned first set of uncertainty analysis results, a second set of predicted audio recognition results whose credibility meets the preset conditions are screened out from the above-mentioned first set of predicted audio recognition results, and a third set of audio samples corresponding to the above-mentioned second set of predicted audio recognition results are screened out from the above-mentioned second set of audio samples; based on the above-mentioned third set of audio samples and the above-mentioned second set of predicted audio recognition results, the above-mentioned initial audio recognition model is trained in the current round, wherein the above-mentioned initial audio recognition model is set to undergo multiple rounds of training until the preset convergence conditions are met.

[0008] According to another aspect of an embodiment of the present invention, an audio recognition method is provided, comprising: obtaining target audio input in a target application; obtaining a target audio recognition result determined by a target audio recognition model based on audio features of the target audio, wherein the target audio recognition model is an audio recognition model obtained by training an initial audio recognition model for multiple rounds until a preset convergence condition is met, the initial audio recognition model is a model obtained by training an audio recognition model to be trained using a first training sample set, the first training sample set includes a first group of audio samples and a first group of actual audio recognition results obtained by annotating the first group of audio samples, and the initial audio recognition model is used to determine the target audio recognition result based on the input audio features. The audio features input are used to determine the predicted audio recognition. In each round of training, the training sample set corresponding to each round is used to train the audio recognition model obtained by the previous round of training. The training sample set corresponding to each round includes the training sample set obtained by the previous round of training and the training sample set obtained by screening in this round. The training sample set obtained by screening in this round includes a group of audio samples and a group of predicted audio recognition results corresponding to the group of audio samples. The group of audio samples are not marked with corresponding actual audio recognition results. The group of predicted audio recognition results are the predicted audio recognition results determined by the audio recognition model after the previous round of training based on the audio features of the group of audio samples; the target audio recognition results are displayed in the target application.

[0009] As an alternative solution, according to the above first set of uncertainty analysis results, screening out a second set of predicted audio recognition results with credibility meeting the preset conditions from the above first set of predicted audio recognition results includes: when the difference between the predicted audio recognition results output by the audio recognition model after the above next round of training and the actual audio recognition results in the above third training sample set meets the above convergence condition, ending the training of the above initial audio recognition model to obtain a target audio recognition model, where a set of predicted audio recognition results in the above third training sample set is regarded as a set of actual audio recognition results.

[0010] As an alternative solution, according to the above first set of uncertainty analysis results, screening out a second set of predicted audio recognition results with credibility meeting the preset conditions from the above first set of predicted audio recognition results includes: when the above first set of uncertainty analysis results includes a set of uncertainty scores, sorting the above set of uncertainty scores from small to large to obtain an uncertainty score sequence, where the higher the above uncertainty score, the lower the credibility of the corresponding predicted audio recognition result; obtaining the first N uncertainty scores in the above uncertainty score sequence before sorting, where the above uncertainty score sequence includes M uncertainty scores and N < M; screening out the above second set of predicted audio recognition results corresponding to the first M uncertainty scores in the above uncertainty score sequence from the above first set of predicted audio recognition results.

[0011] As an alternative solution, obtaining the input target audio in the above target application includes: when the above reference text is displayed in the above target application, obtaining the above target audio generated by reading the above reference text in the above target application, or obtaining the above target audio generated by replying to the above reference text; displaying the above target audio recognition result in the above target application includes: displaying the evaluation score of the above target audio determined by the above target audio recognition model in the above target application.

[0012] According to another aspect of an embodiment of the present invention, a training device for an audio recognition model is also provided, comprising: a first training unit, for training the audio recognition model to be trained using a first training sample set to obtain an initial audio recognition model, wherein the first training sample set includes a first group of audio samples and a first group of actual audio recognition results obtained by annotating the first group of audio samples, and the initial audio recognition model is used to determine predicted audio recognition based on input audio features; a first input unit, for inputting audio features of a second group of audio samples into the initial audio recognition model to obtain a first group of predicted audio recognition results, wherein the second group of audio samples are not annotated with corresponding actual audio recognition results; and the audio features of the second group of audio samples are input into the initial audio recognition model. to the uncertainty analysis model to obtain a first set of uncertainty analysis results, wherein the above-mentioned first set of uncertainty analysis results is used to represent the credibility of the above-mentioned first set of predicted audio recognition results; a first screening unit is used to screen out a second set of predicted audio recognition results whose credibility meets the preset conditions from the above-mentioned first set of predicted audio recognition results based on the above-mentioned first set of uncertainty analysis results, and to screen out a third set of audio samples corresponding to the above-mentioned second set of predicted audio recognition results from the above-mentioned second group of audio samples; a second training unit is used to perform a current round of training on the above-mentioned initial audio recognition model based on the above-mentioned third group of audio samples and the above-mentioned second group of predicted audio recognition results, wherein the above-mentioned initial audio recognition model is set to undergo multiple rounds of training until the preset convergence conditions are met.

[0013] As an optional solution, the above-mentioned second training unit includes: a first merging module, used to merge the above-mentioned third group of audio samples and the above-mentioned second group of predicted audio recognition results into the above-mentioned first training sample set to obtain a second training sample set, wherein the above-mentioned second group of predicted audio recognition results in the above-mentioned second training sample set is regarded as a second group of actual audio recognition results; a first training module, used to use the above-mentioned second training sample set to perform a current round of training on the above-mentioned initial audio recognition model to obtain an audio recognition model after the current round of training.

[0014] As an optional solution, the above-mentioned device also includes: a first acquisition module, which is used to obtain a group of audio samples to be used in the next round of training and a group of predicted audio recognition results corresponding to the above-mentioned group of audio samples when the difference between the predicted audio recognition results output by the audio recognition model after the above-mentioned current round of training and the actual audio recognition results in the above-mentioned second training sample set does not meet the above-mentioned convergence condition, wherein the above-mentioned group of audio samples to be used are not marked with corresponding actual audio recognition results, and the above-mentioned group of predicted audio recognition results are predicted audio recognition results determined by the audio recognition model after the above-mentioned current round of training based on the audio features of the above-mentioned group of audio samples; a second merging module, which is used to merge the above-mentioned group of audio samples to be used and the corresponding group of predicted audio recognition results into the above-mentioned second training sample set to obtain a third training sample set; a second training module, which is used to use the above-mentioned third training sample set to perform the above-mentioned next round of training on the audio recognition model after the above-mentioned current round of training to obtain the audio recognition model after the next round of training.

[0015] As an optional solution, the above-mentioned acquisition module includes: an input submodule, which is used to input the audio features of the fourth group of audio samples into the audio recognition model after the current round of training to obtain a third group of predicted audio recognition results, wherein the above-mentioned fourth group of audio samples are not marked with corresponding actual audio recognition results; input the audio features of the above-mentioned fourth group of audio samples into the above-mentioned uncertainty analysis model to obtain a second group of uncertainty analysis results, wherein the above-mentioned second group of uncertainty analysis results is used to represent the credibility of the above-mentioned third group of predicted audio recognition results; a screening submodule, which is used to screen out the fourth group of predicted audio recognition results whose credibility meets the preset conditions from the above-mentioned third group of predicted audio recognition results based on the above-mentioned second group of uncertainty analysis results, and screen out the fifth group of audio samples corresponding to the above-mentioned fourth group of predicted audio recognition results from the above-mentioned fourth group of audio samples.

[0016] As an optional solution, the above-mentioned first screening unit includes: a second acquisition module, which is used to end the training of the above-mentioned initial audio recognition model and obtain the target audio recognition model when the difference between the predicted audio recognition results output by the audio recognition model after the above-mentioned next round of training and the actual audio recognition results in the above-mentioned third training sample set meets the above-mentioned convergence condition, wherein the above-mentioned set of predicted audio recognition results in the above-mentioned third training sample set is regarded as a set of actual audio recognition results.

[0017] As an optional solution, the above-mentioned first screening unit includes: a third acquisition module, configured to, when the above-mentioned first set of uncertainty analysis results includes a set of uncertainty scores, sort the above-mentioned set of uncertainty scores from smallest to largest to obtain an uncertainty score sequence, where the higher the above-mentioned uncertainty score, the lower the credibility of the corresponding predicted audio recognition result; a fourth acquisition module, configured to obtain the first N uncertainty scores in the above-mentioned uncertainty score sequence, where the above-mentioned uncertainty score sequence includes M uncertainty scores, and N < M; a screening module, configured to screen out the above-mentioned second set of predicted audio recognition results corresponding to the first M uncertainty scores in the above-mentioned uncertainty score sequence from the above-mentioned first set of predicted audio recognition results.

[0018] As an optional solution, the above-mentioned device further includes: a first acquisition unit, configured to acquire an input target audio in a target application; a second acquisition unit, configured to acquire a target audio recognition result determined by a target audio recognition model according to the audio features of the above-mentioned target audio, where the above-mentioned target audio recognition model is an audio recognition model obtained by performing multiple rounds of training on the above-mentioned initial audio recognition model until the preset above-mentioned convergence condition is satisfied; a first display unit, configured to display the above-mentioned target audio recognition result in the above-mentioned target application.

[0019] As an optional solution, it includes: the above-mentioned first acquisition unit, including: a target audio module, configured to, when a reference text is displayed in the above-mentioned target application, acquire the above-mentioned target audio generated by reading the above-mentioned reference text in the above-mentioned target application, or acquire the above-mentioned target audio generated by replying to the above-mentioned reference text; the above-mentioned first display unit, including: a first score module, configured to display the evaluation score of the above-mentioned target audio determined by the above-mentioned target audio recognition model in the above-mentioned target application.

[0020] According to another aspect of an embodiment of the present invention, a training device for an audio recognition model is also provided, including: a third acquisition unit, for acquiring an input target audio in a target application; a fourth acquisition unit, for acquiring a target audio recognition result determined by a target audio recognition model based on the audio features of the target audio, wherein the target audio recognition model is an audio recognition model obtained by training an initial audio recognition model for multiple rounds until a preset convergence condition is met, the initial audio recognition model is a model obtained by training an audio recognition model to be trained using a first training sample set, the first training sample set includes a first group of audio samples and a first group of actual audio recognition results obtained by annotating the first group of audio samples, and the initial audio recognition The model is used to determine predicted audio recognition based on the input audio features. In each round of training, the audio recognition model obtained by the previous round of training is trained using the training sample set corresponding to each round. The training sample set corresponding to each round includes the training sample set obtained by the previous round of training and the training sample set obtained by screening in this round. The training sample set obtained by screening in this round includes a group of audio samples and a group of predicted audio recognition results corresponding to the above group of audio samples. The above group of audio samples is not marked with corresponding actual audio recognition results. The above group of predicted audio recognition results are predicted audio recognition results determined by the audio recognition model after the previous round of training based on the audio features of the above group of audio samples; a third display unit is used to display the above target audio recognition results in the above target application.

[0021] As an optional solution, it includes: a third training unit, which is used to train the audio recognition model to be trained using the first training sample set before obtaining the input target audio in the above-mentioned target application to obtain an initial audio recognition model, wherein the above-mentioned first training sample set includes a first group of audio samples and a first group of actual audio recognition results obtained by annotating the above-mentioned first group of audio samples, and the above-mentioned initial audio recognition model is used to determine the predicted audio recognition according to the input audio features; a second input unit, which is used to input the audio features of the second group of audio samples into the above-mentioned initial audio recognition model before obtaining the input target audio in the above-mentioned target application to obtain a first group of predicted audio recognition results, wherein the above-mentioned second group of audio samples are not annotated with corresponding actual audio recognition results; the audio features of the above-mentioned second group of audio samples are input into the uncertainty analysis model to obtain to the first set of uncertainty analysis results, wherein the above-mentioned first set of uncertainty analysis results is used to represent the credibility of the above-mentioned first set of predicted audio recognition results; a second screening unit is used to, before obtaining the input target audio in the above-mentioned target application, screen out the second set of predicted audio recognition results whose credibility meets the preset conditions from the above-mentioned first set of predicted audio recognition results according to the above-mentioned first set of uncertainty analysis results, and screen out the third group of audio samples corresponding to the above-mentioned second group of predicted audio recognition results from the above-mentioned second group of audio samples; a fourth training unit is used to, before obtaining the input target audio in the above-mentioned target application, perform a current round of training on the above-mentioned initial audio recognition model according to the above-mentioned third group of audio samples and the above-mentioned second group of predicted audio recognition results, wherein the above-mentioned initial audio recognition model is set to undergo multiple rounds of training until the preset convergence conditions are met.

[0022] As an optional solution, it includes: a first merging unit, used to merge the above-mentioned third group of audio samples and the above-mentioned second group of predicted audio recognition results into the above-mentioned first training sample set before obtaining the input target audio in the above-mentioned target application, to obtain a second training sample set, wherein the above-mentioned second group of predicted audio recognition results in the above-mentioned second training sample set is regarded as a second group of actual audio recognition results; a fifth training unit, used to use the above-mentioned second training sample set to perform a current round of training on the above-mentioned initial audio recognition model before obtaining the input target audio in the above-mentioned target application, to obtain an audio recognition model after the current round of training.

[0023] As an optional solution, it includes: a fifth acquisition unit, which is used to obtain a group of audio samples to be used in the next round of training and a group of predicted audio recognition results corresponding to the above group of audio samples when the difference between the predicted audio recognition results output by the audio recognition model after the current round of training and the actual audio recognition results in the above second training sample set does not meet the above convergence condition, wherein the above group of audio samples to be used are not marked with corresponding actual audio recognition results, and the above group of predicted audio recognition results are predicted audio recognition results determined by the audio recognition model after the current round of training according to the audio features of the above group of audio samples; a second merging unit, which is used to merge the above group of audio samples to be used and the corresponding group of predicted audio recognition results into the above second training sample set to obtain a third training sample set; a fourth training unit, which is used to use the above third training sample set to perform the above next round of training on the audio recognition model after the current round of training to obtain the audio recognition model after the next round of training.

[0024] As an optional solution, it includes: a third input unit, which is used to input the audio features of the fourth group of audio samples into the audio recognition model after the current round of training before obtaining the input target audio in the above-mentioned target application, so as to obtain a third group of predicted audio recognition results, wherein the above-mentioned fourth group of audio samples are not marked with corresponding actual audio recognition results; input the audio features of the above-mentioned fourth group of audio samples into the above-mentioned uncertainty analysis model to obtain a second group of uncertainty analysis results, wherein the above-mentioned second group of uncertainty analysis results is used to represent the credibility of the above-mentioned third group of predicted audio recognition results; a third screening unit, which is used to screen out the fourth group of predicted audio recognition results whose credibility meets the preset conditions from the above-mentioned third group of predicted audio recognition results according to the above-mentioned second group of uncertainty analysis results before obtaining the input target audio in the above-mentioned target application, and screen out the fifth group of audio samples corresponding to the above-mentioned fourth group of predicted audio recognition results from the above-mentioned fourth group of audio samples.

[0025] As an optional solution, it includes: a sixth acquisition unit, which is used to end the training of the above-mentioned initial audio recognition model and obtain the target audio recognition model when the difference between the predicted audio recognition result output by the audio recognition model after the next round of training and the actual audio recognition result in the above-mentioned third training sample set meets the above-mentioned convergence condition before acquiring the input target audio in the above-mentioned target application, wherein the above-mentioned set of predicted audio recognition results in the above-mentioned third training sample set is regarded as a set of actual audio recognition results.

[0026] As an alternative solution, it includes: a sorting unit, configured to sort the above-mentioned set of uncertainty scores in ascending order to obtain an uncertainty score sequence before obtaining the input target audio in the above-mentioned target application when the above-mentioned first set of uncertainty analysis results includes a set of uncertainty scores, wherein the higher the above-mentioned uncertainty score, the lower the credibility of the corresponding predicted audio recognition result; a seventh acquisition unit, configured to obtain the top N uncertainty scores in the above-mentioned uncertainty score sequence before obtaining the input target audio in the above-mentioned target application, wherein the above-mentioned uncertainty score sequence includes M uncertainty scores and N < M; a fourth screening unit, configured to screen out the above-mentioned second set of predicted audio recognition results corresponding to the top N uncertainty scores in the above-mentioned first set of predicted audio recognition results before obtaining the input target audio in the above-mentioned target application.

[0027] As an alternative solution, the above-mentioned device further includes: the above-mentioned third acquisition unit, including: a second audio module, configured to obtain the above-mentioned target audio generated by reading the above-mentioned reference text in the above-mentioned target application when the above-mentioned reference text is displayed in the above-mentioned target application, or obtain the above-mentioned target audio generated by replying to the above-mentioned reference text; the above-mentioned third display unit, including: a second score module, configured to display the evaluation score of the above-mentioned target audio determined by the above-mentioned target audio recognition model in the above-mentioned target application.

[0028] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium storing a computer program, wherein the computer program is set to execute the above-mentioned training method of the audio recognition model when running.

[0029] According to another aspect of the embodiments of the present invention, there is also provided an electronic device including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the above-mentioned processor executes the above-mentioned training method of the audio recognition model through the computer program.

[0030] In an embodiment of the present invention, a first training sample set is used to train an audio recognition model to be trained to obtain an initial audio recognition model, wherein the first training sample set includes a first group of audio samples and a first group of actual audio recognition results obtained by annotating the first group of audio samples, and the initial audio recognition model is used to determine the predicted audio recognition based on the input audio features; the audio features of the second group of audio samples are input into the initial audio recognition model to obtain a first group of predicted audio recognition results, wherein the second group of audio samples are not annotated with corresponding actual audio recognition results; the audio features of the second group of audio samples are input into the uncertainty analysis model to obtain a first group of uncertainty analysis results, wherein the first group of uncertainty analysis results is used to represent the credibility of the first group of predicted audio recognition results; according to the above-mentioned A set of uncertainty analysis results, screening out a second group of predicted audio recognition results whose credibility meets the preset conditions from the above-mentioned first group of predicted audio recognition results, and screening out a third group of audio samples corresponding to the above-mentioned second group of predicted audio recognition results from the above-mentioned second group of audio samples; performing a current round of training on the above-mentioned initial audio recognition model based on the above-mentioned third group of audio samples and the above-mentioned second group of predicted audio recognition results, wherein the above-mentioned initial audio recognition model is set to undergo multiple rounds of training until the preset convergence conditions are met, and the training of the audio recognition model can be completed without ensuring that all audio samples are labeled, thereby achieving the purpose of reducing the influence of labeled samples on the audio recognition model, thereby achieving the technical effect of improving the training efficiency of the audio recognition model, and thus solving the technical problem of low training efficiency of the audio recognition model. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0032] Figure 1 is a schematic diagram of an application environment of an optional audio recognition model training method according to an embodiment of the present invention;

[0033] Figure 2 is a schematic diagram of a process of an optional audio recognition model training method according to an embodiment of the present invention;

[0034] Figure 3 is a schematic diagram of an optional training method for an audio recognition model according to an embodiment of the present invention;

[0035] Figure 4 is a schematic diagram of another optional audio recognition model training method according to an embodiment of the present invention;

[0036] Figure 5 is a schematic diagram of another optional audio recognition model training method according to an embodiment of the present invention;

[0037] Figure 6 is a schematic diagram of another optional audio recognition model training method according to an embodiment of the present invention;

[0038] Figure 7 is a schematic diagram of another optional audio recognition model training method according to an embodiment of the present invention;

[0039] Figure 8 is a schematic diagram of another optional audio recognition model training method according to an embodiment of the present invention;

[0040] Figure 9 is a schematic diagram of another optional audio recognition model training method according to an embodiment of the present invention;

[0041] Figure 10 is a schematic diagram of another optional audio recognition model training method according to an embodiment of the present invention;

[0042] Figure 11 is a schematic diagram of another optional audio recognition model training method according to an embodiment of the present invention;

[0043] Figure 12 is a schematic diagram of another optional audio recognition model training method according to an embodiment of the present invention;

[0044] Figure 13 is a schematic diagram of another optional audio recognition model training method according to an embodiment of the present invention;

[0045] Figure 14 is a schematic diagram of an optional audio recognition model training device according to an embodiment of the present invention;

[0046] Figure 15 is a schematic diagram of an optional audio recognition device according to an embodiment of the present invention;

[0047] Figure 16 FIG. 4 is a schematic structural diagram of an optional electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0048] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0049] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0050] First, to facilitate understanding of the embodiments of the present invention, some of the terms or nouns involved in the present invention are explained below:

[0051] Automatic speech recognition (ASR) is the process of converting audio into text.

[0052] Semi-Supervised Learning (SSL) is a learning method that combines supervised learning and unsupervised learning. It uses a large amount of unlabeled data and labeled data to perform pattern recognition.

[0053] Pearson correlation coefficient: It is used to measure the correlation (linear correlation) between two variables X and Y, and its value is between -1 and 1

[0054] SVR: support vector regression, a regression algorithm based on support vector machine

[0055] The K-Nearest Neighbors algorithm (KNN) finds K nearest neighbors for a new prediction instance, and then takes the average of the target values of these K samples as the prediction value of the new sample.

[0056] GBT: Gradient boost tree, a regression algorithm based on boosted trees. It uses the negative gradient of the loss function in the current model as an approximation of the residual to fit a regression tree.

[0057] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0058] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0059] Key technologies in speech technology include automatic speech recognition (ASR), text-to-speech (TTS), and voiceprint recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech becoming one of the most promising methods of human-computer interaction.

[0060] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0061] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, robots, smart medical care, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0062] The solutions provided in the embodiments of this application involve technologies such as artificial intelligence speech recognition and machine learning, which are specifically described through the following embodiments:

[0063] According to one aspect of an embodiment of the present invention, a method for training an audio recognition model is provided. Optionally, as an optional implementation, the method for training an audio recognition model can be applied to, but is not limited to, Figure 1 In the environment shown, it may include, but is not limited to, a user device 102, a network 110, and a server 112. The user device 102 may include, but is not limited to, a display 108, a processor 106, and a memory 104.

[0064] The specific process can be as follows:

[0065] Step S102: The user device 102 obtains a training instruction triggered by the virtual button "Start Training", wherein the training instruction is used to instruct to perform model training based on the first group of labeled audio samples 1022 and the second group of unlabeled audio samples 1024;

[0066] Steps S104-S106: User device 102 sends the training instruction to server 112 via network 110;

[0067] In step S108, the server 112 responds to the training instruction and processes the first set of audio samples 1022 and the second set of audio samples 1024 through the processing engine 116, thereby obtaining a trained audio recognition model and generating corresponding training results.

[0068] In steps S110 - S112 , the server 112 sends the training result to the user device 102 via the network 110 . The processor 106 in the user device 102 displays the training result on the display 108 and stores the training result in the memory 104 .

[0069] remove Figure 1In addition to the examples shown, the above steps can be independently completed by the user device 102, that is, the user device 102 performs steps such as processing the first group of audio samples 1022 and the second group of audio samples 1024, thereby reducing the processing pressure on the server. The user device 102 includes but is not limited to a handheld device (such as a mobile phone), a computer, an intelligent voice interaction device, a smart home appliance, an in-vehicle terminal, etc. The present invention does not limit the specific implementation of the user device 102.

[0070] Alternatively, as an optional implementation, Figure 2 As shown, the training method of the audio recognition model includes:

[0071] S202: training an audio recognition model to be trained using a first training sample set to obtain an initial audio recognition model, wherein the first training sample set includes a first set of audio samples and a first set of actual audio recognition results obtained by annotating the first set of audio samples, and the initial audio recognition model is used to determine a predicted audio recognition based on input audio features;

[0072] S204: Inputting the audio features of the second group of audio samples into the initial audio recognition model to obtain a first group of predicted audio recognition results, wherein the second group of audio samples are not labeled with corresponding actual audio recognition results; inputting the audio features of the second group of audio samples into the uncertainty analysis model to obtain a first group of uncertainty analysis results, wherein the first group of uncertainty analysis results is used to indicate the credibility of the first group of predicted audio recognition results;

[0073] S206, based on the first set of uncertainty analysis results, screening out a second set of predicted audio recognition results whose credibility meets a preset condition from the first set of predicted audio recognition results, and screening out a third set of audio samples corresponding to the second set of predicted audio recognition results from the second set of audio samples;

[0074] S208 , performing a current round of training on the initial audio recognition model based on the third group of audio samples and the second group of predicted audio recognition results, wherein the initial audio recognition model is set to undergo multiple rounds of training until a preset convergence condition is met.

[0075] Optionally, in this embodiment, the training method of the above-mentioned audio recognition model can be but is not limited to being applied in automatic oral evaluation scenarios. For example, an audio recognition model that can recognize spoken audio is trained through the training method of the above-mentioned audio recognition model, and the audio input by the user is recognized and evaluated, and the output of the audio recognition model is displayed as the evaluation result, so that the user can clearly know his or her own oral level.

[0076] Optionally, in this embodiment, the first group of audio samples and the second group of audio samples may be, but are not limited to, unlabeled audio samples, and the first group of actual audio recognition results is a group of audio samples obtained by labeling the first group of audio samples.

[0077] Optionally, in this embodiment, the initial audio recognition model may be, but is not limited to, an audio recognition model obtained by training using a few annotated audio samples. The initial audio recognition model may be, but is not limited to, a semi-finished audio recognition model with basic functions and whose training effect does not meet the convergence conditions.

[0078] Optionally, in this embodiment, the uncertainty analysis model may be, but is not limited to, a model that can automatically execute uncertainty methods, wherein the uncertainty methods may be, but are not limited to, at least one of the following: Gaussian process regression, Monte Carlo dropout method, deep mixture density network, etc. Among them, the Gaussian process uses Gaussian distribution modeling output to determine the mean and variance of each prediction result. This method uses variance as a measure of uncertainty. The larger the variance, the greater the uncertainty. The Monte Carlo dropout method uses multiple models to integrate and analyze the uncertainty of the model. It assumes that for uncertain data, the output of each model has diversity [8]. If the output is more diverse, the uncertainty is greater. The deep mixture density network is similar to the Gaussian process modeling, and models the mean and variance of the results [9]. This method also uses variance as a measure of uncertainty. The larger the variance, the greater the uncertainty.

[0079] It should be noted that an initial audio recognition model is first trained based on the first group of annotated audio samples, and initial audio recognition is performed on the second group of unlabeled audio samples to obtain a first group of predicted audio recognition results; then, the trained uncertainty analysis model is used to perform uncertainty analysis on the second group of unlabeled audio samples, and the analysis results are used to screen the first group of predicted audio recognition results to obtain a second group of predicted audio recognition results; furthermore, a third group of audio samples corresponding to the second group of predicted audio recognition results is obtained from the second group of audio samples, and the initial audio recognition model is iteratively trained using the third group of audio samples until the preset convergence conditions are met to obtain a trained audio recognition model.

[0080] To illustrate further, the optional Figure 3As shown, first, the audio recognition model to be trained 304 is trained using the first group of annotated audio samples 302 to obtain an initial audio recognition model 306; then, the audio features of the second group of audio samples 308 are input into the initial audio recognition model 306 to obtain a first group of predicted audio recognition results 312; and the audio features of the second group of audio samples 308 are input into the uncertainty analysis model 310 to obtain a first group of uncertainty analysis results 314, wherein the first group of uncertainty analysis results 314 is used to represent the credibility of the first group of predicted audio recognition results 312 Based on the first set of uncertainty analysis results 314, a second set of predicted audio recognition results 316 whose credibility meets a preset condition is selected from the first set of predicted audio recognition results 312, and a third set of audio samples 318 corresponding to the second set of predicted audio recognition results 316 is selected from the second set of audio samples 308; based on the third set of audio samples 318 and the second set of predicted audio recognition results 316, the initial audio recognition model 306 is trained in a current round, wherein the initial audio recognition model 306 is configured to undergo multiple rounds of training until a preset convergence condition is satisfied;

[0081] In addition, optionally, when a trained audio recognition model 320 is obtained, input audio 322 in a spoken language assessment scenario (such as audio of a reading reference text) is obtained, and the input audio 322 is input into the audio recognition model 320, and the audio recognition result 324 (such as a spoken language assessment score) is determined based on the output result of the audio recognition model 320.

[0082] To further illustrate, the training process of the optional audio recognition model is as follows: Figure 4 The specific steps are as follows:

[0083] Step S402: training the audio recognition model to be trained using the labeled samples to obtain an initial audio recognition model;

[0084] Step S404: Using the unlabeled samples and the uncertainty analysis model, the initial audio recognition model is initially trained and iteratively trained. The initial training includes screening the samples of the current round using the uncertainty analysis model. After the screening process is completed, the corresponding samples obtained are the samples used in the first round of training in the iterative training process. In the first round of training, in addition to using the samples obtained in the initial training to train the initial audio recognition model, the uncertainty analysis model is also used to perform screening again. After the screening process is completed, the corresponding samples obtained are the samples used in the second round of training in the iterative training process. In summary, in the iterative training, except for the samples used in the first round of training, which are the samples obtained in the initial training, the samples used in the remaining rounds of training are all the samples screened out in the previous round of training.

[0085] Step S406: After multiple rounds of training, until a preset convergence condition is met, a trained audio recognition model is obtained.

[0086] Through the embodiments provided by the present application, a first training sample set is used to train an audio recognition model to be trained to obtain an initial audio recognition model, wherein the first training sample set includes a first group of audio samples and a first group of actual audio recognition results obtained by annotating the first group of audio samples, and the initial audio recognition model is used to determine the predicted audio recognition based on the input audio features; the audio features of the second group of audio samples are input into the initial audio recognition model to obtain a first group of predicted audio recognition results, wherein the second group of audio samples are not annotated with corresponding actual audio recognition results; the audio features of the second group of audio samples are input into the uncertainty analysis model to obtain a first group of uncertainty analysis results, wherein the first group of uncertainty analysis results is used to represent the first group of predicted audio The credibility of the recognition result; based on the first set of uncertainty analysis results, a second set of predicted audio recognition results whose credibility meets the preset conditions are screened out from the first set of predicted audio recognition results, and a third set of audio samples corresponding to the second set of predicted audio recognition results are screened out from the second set of audio samples; based on the third set of audio samples and the second set of predicted audio recognition results, the initial audio recognition model is trained in the current round, wherein the initial audio recognition model is set to undergo multiple rounds of training until the preset convergence conditions are met, and the training of the audio recognition model can be completed without ensuring that all audio samples are labeled, thereby achieving the purpose of reducing the influence of labeled samples on the audio recognition model, thereby achieving the technical effect of improving the training efficiency of the audio recognition model.

[0087] As an optional solution, the initial audio recognition model is trained in a current round based on the third set of audio samples and the second set of predicted audio recognition results, including:

[0088] S1, merging the third set of audio samples and the second set of predicted audio recognition results into the first training sample set to obtain a second training sample set, wherein the second set of predicted audio recognition results in the second training sample set is regarded as the second set of actual audio recognition results;

[0089] S2: Perform a current round of training on the initial audio recognition model using the second training sample set to obtain an audio recognition model after the current round of training.

[0090] Optionally, in this embodiment, the third set of audio samples obtained after screening can be incorporated into the first set of training samples as labeled audio samples to jointly train the initial audio recognition model. In this way, even if limited human resources, material resources, or time prevent the acquisition of a large number of labeled audio samples, the audio recognition model can still be trained.

[0091] Through the embodiments provided in the present application, the third group of audio samples and the second group of predicted audio recognition results are merged into the first training sample set to obtain a second training sample set, wherein the second group of predicted audio recognition results in the second training sample set is regarded as the second group of actual audio recognition results; the initial audio recognition model is trained in the current round using the second training sample set to obtain an audio recognition model after the current round of training, thereby achieving the effect of improving the training efficiency of the audio recognition model.

[0092] As an optional solution, the method further includes:

[0093] S1, when the difference between the predicted audio recognition results output by the audio recognition model after the current round of training and the actual audio recognition results in the second training sample set does not meet the convergence condition, obtaining a group of audio samples to be used in the next round of training and a group of predicted audio recognition results corresponding to the group of audio samples, wherein the group of audio samples to be used are not labeled with corresponding actual audio recognition results, and the group of predicted audio recognition results are predicted audio recognition results determined by the audio recognition model after the current round of training based on audio features of the group of audio samples;

[0094] S2, merging the set of audio samples to be used and the corresponding set of predicted audio recognition results into the second training sample set to obtain a third training sample set;

[0095] S3, using the third training sample set to perform the next round of training on the audio recognition model after the current round of training to obtain the audio recognition model after the next round of training.

[0096] Optionally, in this embodiment, each round of training can use, but is not limited to, new audio samples, and these new audio samples can be, but are not limited to, unlabeled. Based on this, during the training of the audio recognition model, except for the initial construction of the initial audio recognition model, which requires the use of a small number of annotated audio samples, the remaining steps can directly use unlabeled audio samples for training, significantly saving the time required to label audio samples and the resources consumed by labeling for training the audio recognition model.

[0097] Through the embodiments provided by the present application, when the difference between the predicted audio recognition results output by the audio recognition model after the current round of training and the actual audio recognition results in the second training sample set does not meet the convergence condition, a group of audio samples to be used in the next round of training and a group of predicted audio recognition results corresponding to the group of audio samples are obtained, wherein the group of audio samples to be used are not marked with corresponding actual audio recognition results, and the group of predicted audio recognition results are predicted audio recognition results determined by the audio recognition model after the current round of training based on the audio features of the group of audio samples; the group of audio samples to be used and the corresponding group of predicted audio recognition results are merged into the second training sample set to obtain a third training sample set; the third training sample set is used to perform the next round of training on the audio recognition model after the current round of training to obtain the audio recognition model after the next round of training, thereby achieving the effect of improving the training efficiency of the audio recognition model.

[0098] As an optional solution, obtaining a set of audio samples to be used in the next round of training and a set of predicted audio recognition results corresponding to the set of audio samples includes:

[0099] S1, inputting the audio features of the fourth group of audio samples into the audio recognition model after the current round of training to obtain a third group of predicted audio recognition results, wherein the fourth group of audio samples are not labeled with corresponding actual audio recognition results; inputting the audio features of the fourth group of audio samples into the uncertainty analysis model to obtain a second group of uncertainty analysis results, wherein the second group of uncertainty analysis results are used to indicate the credibility of the third group of predicted audio recognition results;

[0100] S2. Based on the second group of uncertainty analysis results, select a fourth group of predicted audio recognition results whose credibility meets a preset condition from the third group of predicted audio recognition results, and select a fifth group of audio samples corresponding to the fourth group of predicted audio recognition results from the fourth group of audio samples.

[0101] Through the embodiments provided in the present application, the audio features of the fourth group of audio samples are input into the audio recognition model after the current round of training to obtain a third group of predicted audio recognition results, wherein the fourth group of audio samples are not labeled with corresponding actual audio recognition results; the audio features of the fourth group of audio samples are input into the uncertainty analysis model to obtain a second group of uncertainty analysis results, wherein the second group of uncertainty analysis results is used to represent the credibility of the third group of predicted audio recognition results; based on the second group of uncertainty analysis results, the fourth group of predicted audio recognition results whose credibility meets the preset conditions are screened out from the third group of predicted audio recognition results, and the fifth group of audio samples corresponding to the fourth group of predicted audio recognition results are screened out from the fourth group of audio samples, thereby achieving the effect of improving the training completeness of the audio recognition model.

[0102] As an optional solution, based on the first set of uncertainty analysis results, a second set of predicted audio recognition results whose credibility meets a preset condition is screened out from the first set of predicted audio recognition results, including:

[0103] When the difference between the predicted audio recognition results output by the audio recognition model after the next round of training and the actual audio recognition results in the third training sample set meets the convergence condition, the training of the initial audio recognition model is terminated to obtain the target audio recognition model, wherein a group of predicted audio recognition results in the third training sample set is regarded as a group of actual audio recognition results.

[0104] Optionally, in this embodiment, when the convergence condition is reached, the training of the initial audio recognition model is terminated to obtain a trained target audio recognition model.

[0105] Through the embodiments provided in the present application, when the difference between the predicted audio recognition results output by the audio recognition model after the next round of training and the actual audio recognition results in the third training sample set meets the convergence condition, the training of the initial audio recognition model is terminated and the target audio recognition model is obtained, wherein a group of predicted audio recognition results in the third training sample set is regarded as a group of actual audio recognition results, thereby achieving the effect of improving the training completeness of the audio recognition model.

[0106] As an optional solution, based on the first set of uncertainty analysis results, a second set of predicted audio recognition results whose credibility meets a preset condition is screened out from the first set of predicted audio recognition results, including:

[0107] S1, when the first set of uncertainty analysis results includes a set of uncertainty scores, sorting the set of uncertainty scores in ascending order to obtain an uncertainty score sequence, wherein a higher uncertainty score indicates a lower confidence level of the corresponding predicted audio recognition result;

[0108] S2, obtain the top N uncertainty scores in the uncertainty score sequence, where the uncertainty score sequence includes M uncertainty scores, N <M;

[0109] S3. Filter out a second group of predicted audio recognition results corresponding to the top N uncertainty scores from the first group of predicted audio recognition results.

[0110] Optionally, in this embodiment, the uncertainty score can be used as one of the screening methods but is not limited to, and each audio recognition result is sorted according to the uncertainty score, and the top N audio recognition results or the audio recognition results with uncertainty scores greater than or equal to the target threshold are taken to form the second group of predicted audio recognition results.

[0111] Through the embodiments provided in this application, when the first set of uncertainty analysis results includes a set of uncertainty scores, the set of uncertainty scores is sorted in ascending order to obtain an uncertainty score sequence. Among them, the higher the uncertainty score, the lower the credibility of the corresponding predicted audio recognition result. The first N uncertainty scores are obtained from the uncertainty score sequence, where the uncertainty score sequence includes M uncertainty scores and N < M. The second set of predicted audio recognition results corresponding to the first N uncertainty scores is screened out from the first set of predicted audio recognition results, achieving the effect of improving the screening efficiency of audio recognition results.

[0112] As an optional solution, the method further includes:

[0113] S1. Obtain the input target audio in the target application;

[0114] S2. Obtain the target audio recognition result determined by the target audio recognition model according to the audio features of the target audio, where the target audio recognition model is an audio recognition model obtained by training the initial audio recognition model for multiple rounds until the preset convergence condition is met;

[0115] S3. Display the target audio recognition result in the target application.

[0116] Among them, obtaining the input target audio in the target application may include, but is not limited to: when the reference text is displayed in the target application, obtaining the target audio generated by reading the reference text aloud in the target application, or obtaining the target audio generated by replying to the reference text;

[0117] Displaying the target audio recognition result in the target application may include, but is not limited to: displaying the evaluation score of the target audio determined by the target audio recognition model in the target application.

[0118] Optionally, in this embodiment, each reference text may include, but is not limited to, one or more reference audios. In the scenario of oral evaluation, the evaluation score of the target audio may be determined by comparing the similarity between the obtained target audio and the reference audio.

[0119] For further illustration, optionally, for example Figure 5 As shown, the reference text "Who are you?" and the prompt message "Please read the above text information aloud" are displayed on the target application interface 502. Then, as Figure 5As shown in (a), a touch operation is recognized on the virtual button "Start Reading", and then the audio signal within the target time period is collected, and the audio signal is input into the target audio recognition model as the target audio, so that the target audio recognition model outputs the corresponding recognition result according to the target audio. The performance of the recognition process in the foreground can be, but is not limited to, as follows Figure 5 As shown in (b); Furthermore, after obtaining the output result of the target audio recognition model, the output result is converted into an evaluation result, such as Figure 5 The evaluation result shown in (c) is expressed as an evaluation score of "85 points".

[0120] To illustrate further, the optional Figure 6 As shown, the target application interface 602 displays a reference text "How are you?" and a prompt message "Please answer the above text", and then Figure 6 As shown in (a), a touch operation is recognized on the virtual button "Start Answering", and then the audio signal within the target time period is collected, and the audio signal is input into the target audio recognition model as the target audio, so that the target audio recognition model outputs the corresponding recognition result according to the target audio. The performance of the recognition process in the foreground can be, but is not limited to, as follows Figure 6 As shown in (b); Furthermore, after obtaining the output result of the target audio recognition model, the output result is converted into an evaluation result, such as Figure 6 The evaluation result shown in (c) is the recognized answer text. In addition, an evaluation score (not shown in the figure) can be given based on whether the answer text is correct and whether the pronunciation is marked.

[0121] Through the embodiments provided in the present application, when the reference text is displayed in the target application, the target audio generated by reading the reference text aloud is obtained in the target application, or the target audio generated by replying to the reference text is obtained; the evaluation score of the target audio determined by the target audio recognition model is displayed in the target application, thereby achieving the effect of improving the accuracy of the audio evaluation.

[0122] Alternatively, as an optional implementation, Figure 7 As shown, the audio recognition method includes:

[0123] S702, obtaining input target audio in the target application;

[0124] S704, obtaining a target audio recognition result determined by a target audio recognition model according to the audio features of the target audio, wherein the target audio recognition model is an audio recognition model obtained by training the initial audio recognition model for multiple rounds until a preset convergence condition is met, the initial audio recognition model is a model obtained by training the audio recognition model to be trained using a first training sample set, the first training sample set includes a first group of audio samples and a first group of actual audio recognition results obtained by annotating the first group of audio samples, the initial audio recognition model is used to determine predicted audio recognition according to the input audio features, and in each round of training, the audio recognition model obtained by the previous round of training is trained using the training sample set corresponding to each round, the training sample set corresponding to each round includes the training sample set obtained by the previous round of training and the training sample set obtained by screening in this round, the training sample set obtained by screening in this round includes a group of audio samples and a group of predicted audio recognition results corresponding to the group of audio samples, a group of audio samples are not annotated with corresponding actual audio recognition results, and a group of predicted audio recognition results are predicted audio recognition results determined by the audio recognition model after the previous round of training according to the audio features of the group of audio samples;

[0125] S706: Display the target audio recognition result in the target application.

[0126] Optionally, in this embodiment, the above-mentioned audio recognition method can be applied, but is not limited to, in automatic oral evaluation scenarios. For example, through the above-mentioned audio recognition method, the audio input by the user is evaluated for oral proficiency, so that the user can clearly know his or her oral proficiency.

[0127] Optionally, in this embodiment, the first group of audio samples and the second group of audio samples may be, but are not limited to, unlabeled audio samples, and the first group of actual audio recognition results is a group of audio samples obtained by labeling the first group of audio samples.

[0128] Optionally, in this embodiment, the initial audio recognition model may be, but is not limited to, an audio recognition model obtained by training using a few annotated audio samples. The initial audio recognition model may be, but is not limited to, a semi-finished audio recognition model with basic functions and whose training effect does not meet the convergence conditions.

[0129] Optionally, in this embodiment, the uncertainty analysis model may be, but is not limited to, a model that can automatically execute uncertainty methods, wherein the uncertainty methods may be, but are not limited to, at least one of the following: Gaussian process regression, Monte Carlo dropout method, deep mixture density network, etc. Among them, the Gaussian process uses Gaussian distribution modeling output to determine the mean and variance of each prediction result. This method uses variance as a measure of uncertainty. The larger the variance, the greater the uncertainty. The Monte Carlo dropout method uses multiple models to integrate and analyze the uncertainty of the model. It assumes that for uncertain data, the output of each model has diversity [8]. If the output is more diverse, the uncertainty is greater. The deep mixture density network is similar to the Gaussian process modeling, and models the mean and variance of the results [9]. This method also uses variance as a measure of uncertainty. The larger the variance, the greater the uncertainty.

[0130] Through the embodiments provided by the present application, an input target audio is obtained in a target application; a target audio recognition result determined by a target audio recognition model according to the audio features of the target audio is obtained, wherein the target audio recognition model is an audio recognition model obtained by training the initial audio recognition model for multiple rounds until a preset convergence condition is met, and the initial audio recognition model is a model obtained by training the audio recognition model to be trained using a first training sample set, the first training sample set including a first group of audio samples and a first group of actual audio recognition results obtained by labeling the first group of audio samples, the initial audio recognition model is used to determine the predicted audio recognition based on the input audio features, and the training sample set corresponding to each round is used in each round of training to train the previous round of training. The obtained audio recognition model is trained, and the training sample set corresponding to each round includes the training sample set obtained in the previous round of training and the training sample set obtained in this round of screening. The training sample set obtained in this round of screening includes a group of audio samples and a group of predicted audio recognition results corresponding to the group of audio samples. A group of audio samples are not labeled with corresponding actual audio recognition results. A group of predicted audio recognition results are predicted audio recognition results determined by the audio recognition model after the previous round of training based on the audio features of a group of audio samples; the target audio recognition results are displayed in the target application, and through a model training method that does not require a large number of labeled audio samples, an audio recognition model that meets the convergence conditions is quickly obtained for audio recognition, thereby achieving a technical effect of improving the efficiency of audio recognition.

[0131] As an optional solution, before obtaining the input target audio in the target application, include:

[0132] S1, training an audio recognition model to be trained using a first training sample set to obtain an initial audio recognition model, wherein the first training sample set includes a first group of audio samples and a first group of actual audio recognition results obtained by annotating the first group of audio samples, and the initial audio recognition model is used to determine a predicted audio recognition based on input audio features;

[0133] S2: Inputting the audio features of the second group of audio samples into the initial audio recognition model to obtain a first group of predicted audio recognition results, wherein the second group of audio samples are not labeled with corresponding actual audio recognition results; inputting the audio features of the second group of audio samples into the uncertainty analysis model to obtain a first group of uncertainty analysis results, wherein the first group of uncertainty analysis results is used to indicate the credibility of the first group of predicted audio recognition results;

[0134] S3, based on the first set of uncertainty analysis results, screening out a second set of predicted audio recognition results whose credibility meets a preset condition from the first set of predicted audio recognition results, and screening out a third set of audio samples corresponding to the second set of predicted audio recognition results from the second set of audio samples;

[0135] The initial audio recognition model is trained in a current round according to the third group of audio samples and the second group of predicted audio recognition results, wherein the initial audio recognition model is set to undergo multiple rounds of training until a preset convergence condition is met.

[0136] It should be noted that an initial audio recognition model is first trained based on the first group of annotated audio samples, and initial audio recognition is performed on the second group of unlabeled audio samples to obtain a first group of predicted audio recognition results; then, the trained uncertainty analysis model is used to perform uncertainty analysis on the second group of unlabeled audio samples, and the analysis results are used to screen the first group of predicted audio recognition results to obtain a second group of predicted audio recognition results; furthermore, a third group of audio samples corresponding to the second group of predicted audio recognition results is obtained from the second group of audio samples, and the initial audio recognition model is iteratively trained using the third group of audio samples until the preset convergence conditions are met to obtain a trained audio recognition model.

[0137] To illustrate further, the optional Figure 3As shown, first, the audio recognition model to be trained 304 is trained using the first group of annotated audio samples 302 to obtain an initial audio recognition model 306; then, the audio features of the second group of audio samples 308 are input into the initial audio recognition model 306 to obtain a first group of predicted audio recognition results 312; and the audio features of the second group of audio samples 308 are input into the uncertainty analysis model 310 to obtain a first group of uncertainty analysis results 314, wherein the first group of uncertainty analysis results 314 is used to represent the credibility of the first group of predicted audio recognition results 312 Based on the first set of uncertainty analysis results 314, a second set of predicted audio recognition results 316 whose credibility meets a preset condition is selected from the first set of predicted audio recognition results 312, and a third set of audio samples 318 corresponding to the second set of predicted audio recognition results 316 is selected from the second set of audio samples 308; based on the third set of audio samples 318 and the second set of predicted audio recognition results 316, the initial audio recognition model 306 is trained in a current round, wherein the initial audio recognition model 306 is configured to undergo multiple rounds of training until a preset convergence condition is satisfied;

[0138] In addition, optionally, when a trained audio recognition model 320 is obtained, input audio 322 in a spoken language assessment scenario (such as audio of a reading reference text) is obtained, and the input audio 322 is input into the audio recognition model 320, and the audio recognition result 324 (such as a spoken language assessment score) is determined based on the output result of the audio recognition model 320.

[0139] To further illustrate, the training process of the optional audio recognition model is as follows: Figure 4 The specific steps are as follows:

[0140] Step S402: training the audio recognition model to be trained using the labeled samples to obtain an initial audio recognition model;

[0141] Step S404, using unlabeled samples and an uncertainty analysis model, performs initial training (round 0 training) and iterative training (rounds 1 to n training) on the initial audio recognition model, wherein the initial training includes screening the samples of the current round using the uncertainty analysis model, and after the screening process is completed, the corresponding samples obtained are the samples used in the first round of training in the iterative training process, and in the first round of training, in addition to using the samples obtained by the initial training to train the initial audio recognition model, the uncertainty analysis model is also used to perform screening again, and after the screening process is completed, the corresponding samples obtained are the samples used in the second round of training in the iterative training process; in summary, in the iterative training, except for the samples used in the first round of training, which are the samples obtained in the initial training, the samples used in the remaining rounds of training are all the samples screened out by the previous round of training, for details, please refer to Figure 4 As shown in the multi-round training diagram on the right side;

[0142] Step S406: After multiple rounds of training, until a preset convergence condition is met, a trained audio recognition model is obtained.

[0143] Through the embodiments provided by the present application, a first training sample set is used to train an audio recognition model to be trained to obtain an initial audio recognition model, wherein the first training sample set includes a first group of audio samples and a first group of actual audio recognition results obtained by annotating the first group of audio samples, and the initial audio recognition model is used to determine the predicted audio recognition based on the input audio features; the audio features of the second group of audio samples are input into the initial audio recognition model to obtain a first group of predicted audio recognition results, wherein the second group of audio samples are not annotated with corresponding actual audio recognition results; the audio features of the second group of audio samples are input into the uncertainty analysis model to obtain a first group of uncertainty analysis results, wherein the first group of uncertainty analysis results is used to represent the first group of predicted audio The credibility of the recognition result; based on the first set of uncertainty analysis results, a second set of predicted audio recognition results whose credibility meets the preset conditions are screened out from the first set of predicted audio recognition results, and a third set of audio samples corresponding to the second set of predicted audio recognition results are screened out from the second set of audio samples; based on the third set of audio samples and the second set of predicted audio recognition results, the initial audio recognition model is trained in the current round, wherein the initial audio recognition model is set to undergo multiple rounds of training until the preset convergence conditions are met, and the training of the audio recognition model can be completed without ensuring that all audio samples are labeled, thereby achieving the purpose of reducing the influence of labeled samples on the audio recognition model, thereby achieving the technical effect of improving the training efficiency of the audio recognition model.

[0144] As an optional solution, the initial audio recognition model is trained in a current round based on the third set of audio samples and the second set of predicted audio recognition results, including:

[0145] S1, merging the third set of audio samples and the second set of predicted audio recognition results into the first training sample set to obtain a second training sample set, wherein the second set of predicted audio recognition results in the second training sample set is regarded as the second set of actual audio recognition results;

[0146] S2: Perform a current round of training on the initial audio recognition model using the second training sample set to obtain an audio recognition model after the current round of training.

[0147] Optionally, in this embodiment, the third set of audio samples obtained after screening can be incorporated into the first set of training samples as labeled audio samples to jointly train the initial audio recognition model. In this way, even if limited human resources, material resources, or time prevent the acquisition of a large number of labeled audio samples, the audio recognition model can still be trained.

[0148] Through the embodiments provided in the present application, the third group of audio samples and the second group of predicted audio recognition results are merged into the first training sample set to obtain a second training sample set, wherein the second group of predicted audio recognition results in the second training sample set is regarded as the second group of actual audio recognition results; the initial audio recognition model is trained in the current round using the second training sample set to obtain an audio recognition model after the current round of training, thereby achieving the effect of improving the training efficiency of the audio recognition model.

[0149] As an optional solution, the method further includes:

[0150] When the difference between the predicted audio recognition results output by the audio recognition model after the current round of training and the actual audio recognition results in the second training sample set does not meet the convergence condition, obtaining a group of audio samples to be used in the next round of training and a group of predicted audio recognition results corresponding to the group of audio samples, wherein the group of audio samples to be used are not labeled with corresponding actual audio recognition results, and the group of predicted audio recognition results are predicted audio recognition results determined by the audio recognition model after the current round of training based on audio features of the group of audio samples;

[0151] Merging a set of audio samples to be used and a corresponding set of predicted audio recognition results into the second training sample set to obtain a third training sample set;

[0152] The audio recognition model after the current round of training is trained using the third training sample set to obtain the audio recognition model after the next round of training.

[0153] Optionally, in this embodiment, each round of training can use, but is not limited to, new audio samples, and these new audio samples can be, but are not limited to, unlabeled. Based on this, during the training of the audio recognition model, except for the initial construction of the initial audio recognition model, which requires the use of a small number of annotated audio samples, the remaining steps can directly use unlabeled audio samples for training, significantly saving the time required to label audio samples and the resources consumed by labeling for training the audio recognition model.

[0154] Through the embodiments provided by the present application, when the difference between the predicted audio recognition results output by the audio recognition model after the current round of training and the actual audio recognition results in the second training sample set does not meet the convergence condition, a group of audio samples to be used in the next round of training and a group of predicted audio recognition results corresponding to the group of audio samples are obtained, wherein the group of audio samples to be used are not marked with corresponding actual audio recognition results, and the group of predicted audio recognition results are predicted audio recognition results determined by the audio recognition model after the current round of training based on the audio features of the group of audio samples; the group of audio samples to be used and the corresponding group of predicted audio recognition results are merged into the second training sample set to obtain a third training sample set; the third training sample set is used to perform the next round of training on the audio recognition model after the current round of training to obtain the audio recognition model after the next round of training, thereby achieving the effect of improving the training efficiency of the audio recognition model.

[0155] As an optional solution, obtaining a set of audio samples to be used in the next round of training and a set of predicted audio recognition results corresponding to the set of audio samples includes:

[0156] Inputting the audio features of the fourth group of audio samples into the audio recognition model after the current round of training to obtain a third group of predicted audio recognition results, wherein the fourth group of audio samples are not labeled with corresponding actual audio recognition results; inputting the audio features of the fourth group of audio samples into the uncertainty analysis model to obtain a second group of uncertainty analysis results, wherein the second group of uncertainty analysis results are used to indicate the credibility of the third group of predicted audio recognition results;

[0157] Based on the second group of uncertainty analysis results, a fourth group of predicted audio recognition results whose credibility meets preset conditions is screened out from the third group of predicted audio recognition results, and a fifth group of audio samples corresponding to the fourth group of predicted audio recognition results is screened out from the fourth group of audio samples.

[0158] Through the embodiments provided in the present application, the audio features of the fourth group of audio samples are input into the audio recognition model after the current round of training to obtain a third group of predicted audio recognition results, wherein the fourth group of audio samples are not labeled with corresponding actual audio recognition results; the audio features of the fourth group of audio samples are input into the uncertainty analysis model to obtain a second group of uncertainty analysis results, wherein the second group of uncertainty analysis results is used to represent the credibility of the third group of predicted audio recognition results; based on the second group of uncertainty analysis results, the fourth group of predicted audio recognition results whose credibility meets the preset conditions are screened out from the third group of predicted audio recognition results, and the fifth group of audio samples corresponding to the fourth group of predicted audio recognition results are screened out from the fourth group of audio samples, thereby achieving the effect of improving the training completeness of the audio recognition model.

[0159] As an optional solution, based on the first set of uncertainty analysis results, a second set of predicted audio recognition results whose credibility meets a preset condition is screened out from the first set of predicted audio recognition results, including:

[0160] When the difference between the predicted audio recognition results output by the audio recognition model after the next round of training and the actual audio recognition results in the third training sample set meets the convergence condition, the training of the initial audio recognition model is terminated to obtain the target audio recognition model, wherein a group of predicted audio recognition results in the third training sample set is regarded as a group of actual audio recognition results.

[0161] Optionally, in this embodiment, when the convergence condition is reached, the training of the initial audio recognition model is terminated to obtain a trained target audio recognition model.

[0162] Through the embodiments provided in the present application, when the difference between the predicted audio recognition results output by the audio recognition model after the next round of training and the actual audio recognition results in the third training sample set meets the convergence condition, the training of the initial audio recognition model is terminated and the target audio recognition model is obtained, wherein a group of predicted audio recognition results in the third training sample set is regarded as a group of actual audio recognition results, thereby achieving the effect of improving the training completeness of the audio recognition model.

[0163] As an optional solution, based on the first set of uncertainty analysis results, a second set of predicted audio recognition results whose credibility meets a preset condition is screened out from the first set of predicted audio recognition results, including:

[0164] S1, when the first set of uncertainty analysis results includes a set of uncertainty scores, sorting the set of uncertainty scores in ascending order to obtain an uncertainty score sequence, wherein a higher uncertainty score indicates a lower confidence level of the corresponding predicted audio recognition result;

[0165] S2, obtain the top N uncertainty scores in the uncertainty score sequence, where the uncertainty score sequence includes M uncertainty scores, N <M;

[0166] S3. Filter out a second group of predicted audio recognition results corresponding to the top N uncertainty scores from the first group of predicted audio recognition results.

[0167] Optionally, in this embodiment, the uncertainty score can be used as one of the screening methods but is not limited to it. Each audio recognition result is sorted according to the uncertainty score, and the top M audio recognition results or the audio recognition results with uncertainty scores greater than or equal to the target threshold are taken to form the second group of predicted audio recognition results.

[0168] Through the embodiments provided in this application, when the first set of uncertainty analysis results includes a set of uncertainty scores, the set of uncertainty scores is sorted from smallest to largest to obtain an uncertainty score sequence. Among them, the higher the uncertainty score, the lower the credibility of the corresponding predicted audio recognition result; the top N uncertainty scores are obtained in the uncertainty score sequence, where the uncertainty score sequence includes M uncertainty scores and N < M; the second set of predicted audio recognition results corresponding to the top N uncertainty scores is screened out from the first set of predicted audio recognition results, achieving the effect of improving the screening efficiency of audio recognition results.

[0169] As an optional solution, obtaining the input target audio in the target application includes: when the reference text is displayed in the target application, obtaining the target audio generated by reading the reference text aloud in the target application, or obtaining the target audio generated by replying to the reference text.

[0170] Displaying the target audio recognition result in the target application includes: displaying the evaluation score of the target audio determined by the target audio recognition model in the target application.

[0171] Optionally, in this embodiment, each reference text may or may not correspond to one or more reference audios respectively. In the scenario of oral evaluation, the evaluation score of the target audio may or may not be determined by comparing the similarity between the obtained target audio and the reference audio.

[0172] For further illustration by example, optionally, for example Figure 5 As shown, the reference text "Who are you?" and the prompt message "Please read the above text aloud" are displayed on the target application interface 502. Then, as shown in (a) of Figure 5 , a touch operation is recognized on the virtual button "Start Reading", and then the audio signal within the target time period is collected and used as the target audio to be input into the target audio recognition model, so that the target audio recognition model outputs the corresponding recognition result according to the target audio. The performance of its recognition process in the foreground may or may not be as shown in (b) of Figure 5 ; furthermore, after obtaining the output result of the target audio recognition model, the output result is converted into an evaluation result, and the evaluation result shown in (c) of Figure 5 is shown as the evaluation score "85 points".

[0173] For further illustration by example, optionally, for example Figure 6 As shown, the reference text "How are you?" and the prompt message "Please answer the above text" are displayed on the target application interface 602. Then, as shown in Figure 6As shown in (a), a touch operation is recognized on the virtual button "Start Answering", and then the audio signal within the target time period is collected, and the audio signal is input into the target audio recognition model as the target audio, so that the target audio recognition model outputs the corresponding recognition result according to the target audio. The performance of the recognition process in the foreground can be, but is not limited to, as follows Figure 6 As shown in (b); Furthermore, after obtaining the output result of the target audio recognition model, the output result is converted into an evaluation result, such as Figure 6 The evaluation result shown in (c) is the recognized answer text. In addition, an evaluation score (not shown in the figure) can be given based on whether the answer text is correct and whether the pronunciation is marked.

[0174] Through the embodiments provided in the present application, when the reference text is displayed in the target application, the target audio generated by reading the reference text aloud is obtained in the target application, or the target audio generated by replying to the reference text is obtained; the evaluation score of the target audio determined by the target audio recognition model is displayed in the target application, thereby achieving the effect of improving the accuracy of the audio evaluation.

[0175] As an optional solution, for ease of understanding, the training method of the audio recognition model and the audio recognition method described above are described using an automatic oral assessment scenario. The automatic oral assessment scenario can include, but is not limited to, scenarios related to oral exams, such as objective questions, such as reading aloud questions, and subjective questions, such as picture description and oral composition. The details are as follows:

[0176] Automatic oral evaluation often relies on a large amount of manually annotated data. When there is less manually annotated data, the effect is difficult to guarantee. Based on this, an algorithm using a semi-supervised learning pseudo-labeling algorithm for oral evaluation training is proposed, which effectively alleviates the demand for data. First, a oral evaluation model is trained based on a small amount of labeled oral evaluation data. The model is used to predict the scores of the remaining unlabeled oral evaluation audio. Since the predicted scores contain a lot of incorrect labels or noise data, an uncertainty analysis algorithm is used to obtain uncertainty parameters. Based on the uncertainty analysis results, the predicted test data is filtered, and the filtered data is used to expand the training set. Based on the new expanded training set, the oral evaluation model is retrained to improve the oral evaluation effect.

[0177] Specifically, first of all, Figure 8 As shown in (a), click the start reading button 802 to start reading the sentence; Figure 8 As shown in (b), click the end follow-up button 804 to end the follow-up sentence; Figure 9 As shown, the screen returns the evaluation result 902 and displays it to the user, such as the sentence evaluation result is 4 stars.

[0178] For example Figure 10 As shown, click the start recording button 1002 to start answering questions; click the end recording button 1004 to end answering questions. Figure 11 As shown, the screen returns the evaluation result 1102 and displays it to the user, such as the sentence evaluation result is 4 stars.

[0179] In addition, the overall process can refer to Figure 12 The specific steps are as follows:

[0180] S1202, the user opens the application 1202, and the screen displays the question; clicks "Start Recording" in the application 1202 to answer the question;

[0181] S1204, the application 1202 sends the audio and the read text to the server 1204;

[0182] S1206, the server 1204 sends the audio and question information to the pseudo-label-based spoken language evaluation model 1206;

[0183] S1208, the oral evaluation module 1206 returns the scoring result to the server 1204;

[0184] S1210 , the server 1204 returns the final score to the application 1202 , and the user views the final score on the application 1202 .

[0185] Furthermore, the pseudo-label spoken language evaluation model combined with uncertainty analysis can be referenced Figure 13 The specific steps are as follows:

[0186] First, input the audio and input the audio into ASR (automatic speech recognition) to obtain the text of the speech recognition and the start and end time of each phoneme and each word in the audio. Input the audio, alignment results and recognized text into the feature extraction module to extract acoustic features and text features. Input these features into a trained base model to predict the spoken score. At the same time, input these features into a score uncertainty analysis module to obtain the uncertainty analysis results. Finally, input the predicted score and uncertainty analysis results into the pseudo sample screening module to filter out the available pseudo samples. Pseudo sample D u (10 questions, 10 audios, 10 model prediction scores) and the training data D of the base model L (10 questions, 10 audios, 10 manual scores) are combined as shown in the following formula (1). Retrain the oral evaluation model. This process can be repeated multiple times, continuously incorporating new pseudo-labeled samples and continuously retraining the oral evaluation model until convergence.

[0187] D′ L =D u ∪D L (1)

[0188] Among them, the text features extracted based on ASR recognition mainly include semantic features, pragmatic features, keyword features, and text disfluency features. Keyword features mainly include extracting keywords from standard answers and keywords from the answer content, and calculating precision and recall rates. Pragmatic features include the diversity of words and sentence patterns in the answer content, as well as the grammatical accuracy of the answer content based on language model analysis. Semantic features include the topic features of the answer content, tf-idf features, etc. The text disfluency feature is the statistical proportion of the identified disfluent components in the text;

[0189] Acoustic features are mainly divided into pronunciation accuracy, pronunciation fluency, and pronunciation rhythm. Pronunciation accuracy is based on the confidence level of speech recognition and is evaluated at the phoneme, word, and sentence levels in the corresponding pronunciation content. Pronunciation fluency includes speech rate characteristics during the pronunciation process, characteristics based on duration statistics such as the average duration of pronunciation segments and the average pause duration between pronunciation segments. Pronunciation rhythm includes the evaluation of pronunciation rhythm, the correctness of word stress in sentences, and the evaluation of sentence boundary tones.

[0190] Based on the extracted acoustic and text features, a regression model is constructed and fitted with manual scoring. The regression model can be a traditional regression model such as KNN, SVR, GBT tree model, etc., or a deep neural network model, which obtains the final score through multi-layer network forward propagation;

[0191] Based on the extracted text features and acoustic features, an uncertainty analysis model is constructed. At present, there are many types of uncertainty methods, and typical methods include Gaussian process regression, Monte Carlo dropout method, deep mixture density network, etc. Among them, the Gaussian process uses Gaussian distribution modeling output to determine the mean and variance of each prediction result. This method uses variance as a measure of uncertainty. The larger the variance, the greater the uncertainty. The Monte Carlo dropout method uses multiple models to integrate and analyze the uncertainty of the model. It assumes that for uncertain data, the output of each model has diversity [8]. If the output is more diverse, the uncertainty is greater. Deep mixture density network is similar to Gaussian process modeling. It models the mean and variance of the results. This method also uses variance as a measure of uncertainty. The larger the variance, the greater the uncertainty.

[0192] Based on the scores output by the spoken language evaluation model and the uncertainty scores output by the uncertainty analysis module, it is determined whether to use the pseudo sample to retrain the model. Assume that the predicted score of the i-th speech is p i , the uncertainty score is C iThen whether the final sample is selected is R i It is expressed as the following formula (2). Where T1 and T2 are preset thresholds, representing the minimum uncertainty and maximum uncertainty of the selected samples. Such thresholds can be determined by manual setting or search algorithms. T1 < T2, and the smaller part is taken;

[0193] R i = I[C i > T1 & C i < T2] (2)

[0194] In addition, in this embodiment, two data sets can be but are not limited to being adopted. One data set is the answer data of the reading question type in the oral test, and the other is the answer data of the picture description question type in the oral test. There are 1500 pieces of each type of data, which are labeled by three experts. The final measurement effect is mainly through the Pearson correlation coefficient and the accuracy rate (that is, the proportion of the label and the predicted score less than or equal to 1 level). It can be seen from the results that the effect of the oral evaluation model can be greatly improved by combining the uncertainty analysis results based on the pseudo-label algorithm.

[0195] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0196] According to another aspect of the embodiments of the present invention, there is also provided a training device for an audio recognition model for implementing the training method of the above audio recognition model. As Figure 14 shown, the device includes:

[0197] A first training unit 1402, configured to train the to-be-trained audio recognition model using a first training sample set to obtain an initial audio recognition model. The first training sample set includes a first group of audio samples and a first group of actual audio recognition results obtained by annotating the first group of audio samples. The initial audio recognition model is used to determine a predicted audio recognition according to the input audio features;

[0198] A first input unit 1404, configured to input the audio features of a second group of audio samples into the initial audio recognition model to obtain a first group of predicted audio recognition results. The second group of audio samples is not annotated with corresponding actual audio recognition results; input the audio features of the second group of audio samples into the uncertainty analysis model to obtain a first group of uncertainty analysis results, where the first group of uncertainty analysis results is used to represent the credibility of the first group of predicted audio recognition results;

[0199] A first screening unit 1406 is configured to screen, based on the first set of uncertainty analysis results, a second set of predicted audio recognition results whose credibility meets a preset condition from the first set of predicted audio recognition results, and to screen, from the second set of audio samples, a third set of audio samples corresponding to the second set of predicted audio recognition results;

[0200] The second training unit 1408 is configured to perform a current round of training on the initial audio recognition model based on the third set of audio samples and the second set of predicted audio recognition results, wherein the initial audio recognition model is configured to undergo multiple rounds of training until a preset convergence condition is met.

[0201] Optionally, in this embodiment, the training device of the above-mentioned audio recognition model can be used, but is not limited to, in automatic oral evaluation scenarios. For example, an audio recognition model that can recognize spoken audio is trained by the training device of the above-mentioned audio recognition model, and the audio input by the user is recognized and evaluated, and the output of the audio recognition model is displayed as the evaluation result, so that the user can clearly know his or her oral level.

[0202] Optionally, in this embodiment, the first group of audio samples and the second group of audio samples may be, but are not limited to, unlabeled audio samples, and the first group of actual audio recognition results is a group of audio samples obtained by labeling the first group of audio samples.

[0203] Optionally, in this embodiment, the initial audio recognition model may be, but is not limited to, an audio recognition model obtained by training using a few annotated audio samples. The initial audio recognition model may be, but is not limited to, a semi-finished audio recognition model with basic functions and whose training effect does not meet the convergence conditions.

[0204] Optionally, in this embodiment, the uncertainty analysis model may be, but is not limited to, a model that can automatically execute an uncertainty device, wherein the uncertainty device may be, but is not limited to, at least one of the following: Gaussian process regression, Monte Carlo dropout device, deep mixed density network, etc. Among them, the Gaussian process uses Gaussian distribution modeling output to determine the mean and variance of each prediction result. The device uses variance as a measure of uncertainty. The larger the variance, the greater the uncertainty. The Monte Carlo dropout device uses multiple models to integrate and analyze the uncertainty of the model. It assumes that for uncertain data, the output of each model has diversity [8]. If the output is more diverse, the uncertainty is greater. The deep mixed density network is similar to the Gaussian process modeling, and models the mean and variance of the results [9]. The device also uses variance as a measure of uncertainty. The larger the variance, the greater the uncertainty.

[0205] It should be noted that an initial audio recognition model is first trained based on the first group of annotated audio samples, and initial audio recognition is performed on the second group of unlabeled audio samples to obtain a first group of predicted audio recognition results; then, the trained uncertainty analysis model is used to perform uncertainty analysis on the second group of unlabeled audio samples, and the analysis results are used to screen the first group of predicted audio recognition results to obtain a second group of predicted audio recognition results; furthermore, a third group of audio samples corresponding to the second group of predicted audio recognition results is obtained from the second group of audio samples, and the initial audio recognition model is iteratively trained using the third group of audio samples until the preset convergence conditions are met to obtain a trained audio recognition model.

[0206] For specific embodiments, reference may be made to the examples shown in the above-mentioned audio recognition model training method, which will not be described in detail in this example.

[0207] Through the embodiments provided by the present application, a first training sample set is used to train an audio recognition model to be trained to obtain an initial audio recognition model, wherein the first training sample set includes a first group of audio samples and a first group of actual audio recognition results obtained by annotating the first group of audio samples, and the initial audio recognition model is used to determine the predicted audio recognition based on the input audio features; the audio features of the second group of audio samples are input into the initial audio recognition model to obtain a first group of predicted audio recognition results, wherein the second group of audio samples are not annotated with corresponding actual audio recognition results; the audio features of the second group of audio samples are input into the uncertainty analysis model to obtain a first group of uncertainty analysis results, wherein the first group of uncertainty analysis results is used to represent the first group of predicted audio The credibility of the recognition result; based on the first set of uncertainty analysis results, a second set of predicted audio recognition results whose credibility meets the preset conditions are screened out from the first set of predicted audio recognition results, and a third set of audio samples corresponding to the second set of predicted audio recognition results are screened out from the second set of audio samples; based on the third set of audio samples and the second set of predicted audio recognition results, the initial audio recognition model is trained in the current round, wherein the initial audio recognition model is set to undergo multiple rounds of training until the preset convergence conditions are met, and the training of the audio recognition model can be completed without ensuring that all audio samples are labeled, thereby achieving the purpose of reducing the influence of labeled samples on the audio recognition model, thereby achieving the technical effect of improving the training efficiency of the audio recognition model.

[0208] As an optional solution, the second training unit 1408 includes:

[0209] a first merging module, configured to merge the third set of audio samples and the second set of predicted audio recognition results into the first training sample set to obtain a second training sample set, wherein the second set of predicted audio recognition results in the second training sample set is regarded as the second set of actual audio recognition results;

[0210] The first training module is configured to perform a current round of training on the initial audio recognition model using the second training sample set to obtain an audio recognition model after the current round of training.

[0211] For specific embodiments, reference may be made to the examples shown in the above-mentioned audio recognition model training method, which will not be described in detail in this example.

[0212] As an optional solution, the device further includes:

[0213] a first acquisition module, configured to acquire, when a difference between a predicted audio recognition result output by the audio recognition model after a current round of training and an actual audio recognition result in a second training sample set does not satisfy a convergence condition, a set of audio samples to be used in a next round of training and a set of predicted audio recognition results corresponding to the set of audio samples, wherein the set of audio samples to be used are not labeled with corresponding actual audio recognition results, and the set of predicted audio recognition results are predicted audio recognition results determined by the audio recognition model after the current round of training based on audio features of the set of audio samples;

[0214] A second merging module is configured to merge a set of audio samples to be used and a corresponding set of predicted audio recognition results into the second training sample set to obtain a third training sample set;

[0215] The second training module is used to use the third training sample set to perform the next round of training on the audio recognition model after the current round of training to obtain the audio recognition model after the next round of training.

[0216] For specific embodiments, reference may be made to the examples shown in the above-mentioned audio recognition model training method, which will not be described in detail in this example.

[0217] As an optional solution, obtain modules, including:

[0218] an input submodule, configured to input the audio features of the fourth group of audio samples into the audio recognition model after the current round of training to obtain a third group of predicted audio recognition results, wherein the fourth group of audio samples are not labeled with corresponding actual audio recognition results; and input the audio features of the fourth group of audio samples into the uncertainty analysis model to obtain a second group of uncertainty analysis results, wherein the second group of uncertainty analysis results are used to indicate the credibility of the third group of predicted audio recognition results;

[0219] The screening submodule is used to screen out a fourth group of predicted audio recognition results whose credibility meets preset conditions from the third group of predicted audio recognition results based on the second group of uncertainty analysis results, and to screen out a fifth group of audio samples corresponding to the fourth group of predicted audio recognition results from the fourth group of audio samples.

[0220] For specific embodiments, reference may be made to the examples shown in the above-mentioned audio recognition model training method, which will not be described in detail in this example.

[0221] As an optional solution, the first screening unit 1406 includes:

[0222] The second acquisition module is used to end the training of the initial audio recognition model and obtain the target audio recognition model when the difference between the predicted audio recognition results output by the audio recognition model after the next round of training and the actual audio recognition results in the third training sample set meets the convergence condition, wherein a group of predicted audio recognition results in the third training sample set is regarded as a group of actual audio recognition results.

[0223] For specific embodiments, reference may be made to the examples shown in the above-mentioned audio recognition model training method, which will not be described in detail in this example.

[0224] As an optional solution, the first screening unit 1406 includes:

[0225] a third acquisition module, configured to, when the first set of uncertainty analysis results includes a set of uncertainty scores, sort the set of uncertainty scores in ascending order to obtain a sequence of uncertainty scores, wherein a higher uncertainty score indicates a lower credibility of the corresponding predicted audio recognition result;

[0226] The fourth acquisition module is used to obtain the top N uncertainty scores in the uncertainty score sequence, wherein the uncertainty score sequence includes M uncertainty scores, N <M;

[0227] The screening module is used to screen out a second group of predicted audio recognition results corresponding to the top N uncertainty scores from the first group of predicted audio recognition results.

[0228] For specific embodiments, reference may be made to the examples shown in the above-mentioned audio recognition model training method, which will not be described in detail in this example.

[0229] As an optional solution, the device further includes:

[0230] A first acquiring unit, configured to acquire input target audio in a target application;

[0231] a second acquisition unit, configured to acquire a target audio recognition result determined by a target audio recognition model based on audio features of the target audio, wherein the target audio recognition model is an audio recognition model obtained by training the initial audio recognition model for multiple rounds until a preset convergence condition is satisfied;

[0232] The first display unit is configured to display a target audio recognition result in a target application.

[0233] For specific embodiments, reference may be made to the examples shown in the above-mentioned audio recognition model training method, which will not be described in detail in this example.

[0234] As an optional solution, it includes:

[0235] The first acquisition unit includes: a target audio module, configured to acquire, in the target application, target audio generated by reading aloud the reference text, or acquire target audio generated by replying to the reference text, when the reference text is displayed in the target application;

[0236] The first display unit includes: a first score module, which is used to display the evaluation score of the target audio determined by the target audio recognition model in the target application.

[0237] For specific embodiments, reference may be made to the examples shown in the above-mentioned audio recognition model training method, which will not be described in detail in this example.

[0238] According to another aspect of the embodiments of the present invention, an audio recognition device for implementing the above audio recognition method is also provided. Figure 15 As shown, the device includes:

[0239] The third acquiring unit 1502 is configured to acquire the input target audio in the target application;

[0240] The fourth acquisition unit 1504 is used to obtain a target audio recognition result determined by the target audio recognition model according to the audio features of the target audio, wherein the target audio recognition model is an audio recognition model obtained by training the initial audio recognition model for multiple rounds until a preset convergence condition is met, the initial audio recognition model is a model obtained by training the audio recognition model to be trained using a first training sample set, the first training sample set includes a first group of audio samples and a first group of actual audio recognition results obtained by annotating the first group of audio samples, the initial audio recognition model is used to determine predicted audio recognition according to the input audio features, and in each round of training, the audio recognition model obtained by the previous round of training is trained using the training sample set corresponding to each round, the training sample set corresponding to each round includes the training sample set obtained by the previous round of training and the training sample set obtained by screening in this round, the training sample set obtained by screening in this round includes a group of audio samples and a group of predicted audio recognition results corresponding to the group of audio samples, a group of audio samples are not annotated with corresponding actual audio recognition results, and a group of predicted audio recognition results are predicted audio recognition results determined by the audio recognition model after the previous round of training according to the audio features of the group of audio samples;

[0241] The third display unit 1506 is configured to display the target audio recognition result in the target application.

[0242] Optionally, in this embodiment, the above-mentioned audio recognition device can be used, but is not limited to, in automatic oral evaluation scenarios. For example, through the above-mentioned audio recognition method, the audio input by the user is evaluated for oral proficiency, so that the user can clearly know his or her oral proficiency.

[0243] Optionally, in this embodiment, the first group of audio samples and the second group of audio samples may be, but are not limited to, unlabeled audio samples, and the first group of actual audio recognition results is a group of audio samples obtained by labeling the first group of audio samples.

[0244] Optionally, in this embodiment, the initial audio recognition model may be, but is not limited to, an audio recognition model obtained by training using a few annotated audio samples. The initial audio recognition model may be, but is not limited to, a semi-finished audio recognition model with basic functions and whose training effect does not meet the convergence conditions.

[0245] Optionally, in this embodiment, the uncertainty analysis model may be, but is not limited to, a model that can automatically execute uncertainty methods, wherein the uncertainty methods may be, but are not limited to, at least one of the following: Gaussian process regression, Monte Carlo dropout method, deep mixture density network, etc. Among them, the Gaussian process uses Gaussian distribution modeling output to determine the mean and variance of each prediction result. This method uses variance as a measure of uncertainty. The larger the variance, the greater the uncertainty. The Monte Carlo dropout method uses multiple models to integrate and analyze the uncertainty of the model. It assumes that for uncertain data, the output of each model has diversity [8]. If the output is more diverse, the uncertainty is greater. The deep mixture density network is similar to the Gaussian process modeling, and models the mean and variance of the results [9]. This method also uses variance as a measure of uncertainty. The larger the variance, the greater the uncertainty.

[0246] Through the embodiments provided by the present application, an input target audio is obtained in a target application; a target audio recognition result determined by a target audio recognition model according to the audio features of the target audio is obtained, wherein the target audio recognition model is an audio recognition model obtained by training the initial audio recognition model for multiple rounds until a preset convergence condition is met, and the initial audio recognition model is a model obtained by training the audio recognition model to be trained using a first training sample set, the first training sample set including a first group of audio samples and a first group of actual audio recognition results obtained by labeling the first group of audio samples, the initial audio recognition model is used to determine the predicted audio recognition based on the input audio features, and the training sample set corresponding to each round is used in each round of training to train the previous round of training. The obtained audio recognition model is trained, and the training sample set corresponding to each round includes the training sample set obtained in the previous round of training and the training sample set obtained in this round of screening. The training sample set obtained in this round of screening includes a group of audio samples and a group of predicted audio recognition results corresponding to the group of audio samples. A group of audio samples are not labeled with corresponding actual audio recognition results. A group of predicted audio recognition results are predicted audio recognition results determined by the audio recognition model after the previous round of training based on the audio features of a group of audio samples; the target audio recognition results are displayed in the target application, and through a model training method that does not require a large number of labeled audio samples, an audio recognition model that meets the convergence conditions is quickly obtained for audio recognition, thereby achieving a technical effect of improving the efficiency of audio recognition.

[0247] As an optional solution, it includes:

[0248] a third training unit, configured to, before obtaining the input target audio in the target application, train the audio recognition model to be trained using the first training sample set to obtain an initial audio recognition model, wherein the first training sample set includes a first group of audio samples and a first group of actual audio recognition results obtained by annotating the first group of audio samples, and the initial audio recognition model is configured to determine a predicted audio recognition based on the input audio features;

[0249] a second input unit, configured to input the audio features of the second group of audio samples into an initial audio recognition model before obtaining the input target audio in the target application to obtain a first group of predicted audio recognition results, wherein the second group of audio samples are not labeled with corresponding actual audio recognition results; and input the audio features of the second group of audio samples into an uncertainty analysis model to obtain a first group of uncertainty analysis results, wherein the first group of uncertainty analysis results is used to indicate the credibility of the first group of predicted audio recognition results;

[0250] a second screening unit configured to, before acquiring the input target audio in the target application, screen, based on the first set of uncertainty analysis results, a second set of predicted audio recognition results whose credibility meets a preset condition from the first set of predicted audio recognition results, and screen a third set of audio samples corresponding to the second set of predicted audio recognition results from the second set of audio samples;

[0251] The fourth training unit is used to perform a current round of training on the initial audio recognition model based on the third group of audio samples and the second group of predicted audio recognition results before obtaining the input target audio in the target application, wherein the initial audio recognition model is set to undergo multiple rounds of training until a preset convergence condition is met.

[0252] For specific embodiments, reference may be made to the examples shown in the above-mentioned audio recognition model training method, which will not be described in detail in this example.

[0253] As an optional solution, it includes:

[0254] a first merging unit, configured to merge the third group of audio samples and the second group of predicted audio recognition results into the first training sample set before obtaining the input target audio in the target application, to obtain a second training sample set, wherein the second group of predicted audio recognition results in the second training sample set is regarded as the second group of actual audio recognition results;

[0255] The fifth training unit is configured to perform a current round of training on the initial audio recognition model using the second training sample set before obtaining the input target audio in the target application to obtain the audio recognition model after the current round of training.

[0256] For specific embodiments, reference may be made to the examples shown in the above-mentioned audio recognition model training method, which will not be described in detail in this example.

[0257] As an optional solution, it includes:

[0258] a fifth acquisition unit, configured to acquire, when a difference between a predicted audio recognition result output by the audio recognition model after a current round of training and an actual audio recognition result in the second training sample set does not satisfy a convergence condition, a group of audio samples to be used in a next round of training and a group of predicted audio recognition results corresponding to the group of audio samples, wherein the group of audio samples to be used are not labeled with corresponding actual audio recognition results, and the group of predicted audio recognition results are predicted audio recognition results determined by the audio recognition model after the current round of training based on audio features of the group of audio samples;

[0259] A second merging unit is configured to merge the set of audio samples to be used and the corresponding set of predicted audio recognition results into the second training sample set to obtain a third training sample set;

[0260] The fourth training unit is configured to perform a next round of training on the audio recognition model after the current round of training using the third training sample set to obtain the audio recognition model after the next round of training.

[0261] For specific embodiments, reference may be made to the examples shown in the above-mentioned audio recognition model training method, which will not be described in detail in this example.

[0262] As an optional solution, it includes:

[0263] a third input unit, configured to input the audio features of the fourth group of audio samples into the audio recognition model after the current training before obtaining the input target audio in the target application, to obtain a third group of predicted audio recognition results, wherein the fourth group of audio samples are not labeled with corresponding actual audio recognition results; and input the audio features of the fourth group of audio samples into the uncertainty analysis model to obtain a second group of uncertainty analysis results, wherein the second group of uncertainty analysis results are used to indicate the credibility of the third group of predicted audio recognition results;

[0264] The third screening unit is used to screen out a fourth group of predicted audio recognition results whose credibility meets a preset condition from the third group of predicted audio recognition results based on the second group of uncertainty analysis results before obtaining the input target audio in the target application, and to screen out a fifth group of audio samples corresponding to the fourth group of predicted audio recognition results from the fourth group of audio samples.

[0265] For specific embodiments, reference may be made to the examples shown in the above-mentioned audio recognition model training method, which will not be described in detail in this example.

[0266] As an optional solution, it includes:

[0267] The sixth acquisition unit is used to end the training of the initial audio recognition model before acquiring the input target audio in the target application, when the difference between the predicted audio recognition result output by the audio recognition model after the next round of training and the actual audio recognition result in the third training sample set meets the convergence condition, and obtain the target audio recognition model, wherein a group of predicted audio recognition results in the third training sample set is regarded as a group of actual audio recognition results.

[0268] For specific embodiments, reference may be made to the examples shown in the above-mentioned audio recognition model training method, which will not be described in detail in this example.

[0269] As an optional solution, it includes:

[0270] a sorting unit, configured to, before obtaining the input target audio in the target application, sort the uncertainty scores in ascending order when the first set of uncertainty analysis results includes a set of uncertainty scores to obtain a sequence of uncertainty scores, wherein a higher uncertainty score indicates a lower confidence level of the corresponding predicted audio recognition result;

[0271] The seventh acquisition unit is configured to acquire the first N uncertainty scores in the uncertainty score sequence before acquiring the input target audio in the target application, wherein the uncertainty score sequence includes M uncertainty scores, N <M;

[0272] The fourth screening unit is used to screen out the second group of predicted audio recognition results corresponding to the top N uncertainty scores from the first group of predicted audio recognition results before obtaining the input target audio in the target application.

[0273] For specific embodiments, reference may be made to the examples shown in the above-mentioned audio recognition model training method, which will not be described in detail in this example.

[0274] As an optional solution, the device further includes:

[0275] The third acquisition unit includes: a second audio module, configured to acquire, in the target application, target audio generated by reading the reference text aloud, or acquiring target audio generated by replying to the reference text, when the reference text is displayed in the target application;

[0276] The third display unit includes: a second score module, which is used to display the evaluation score of the target audio determined by the target audio recognition model in the target application.

[0277] For specific embodiments, reference may be made to the examples shown in the above-mentioned audio recognition model training method, which will not be described in detail in this example.

[0278] According to another aspect of the present invention, an electronic device for implementing the above-mentioned training method of the audio recognition model is also provided. Figure 16 As shown, the electronic device includes a memory 1602 and a processor 1604. The memory 1602 stores a computer program, and the processor 1604 is configured to execute the steps in any of the above method embodiments through the computer program.

[0279] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.

[0280] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:

[0281] S1, training an audio recognition model to be trained using a first training sample set to obtain an initial audio recognition model, wherein the first training sample set includes a first group of audio samples and a first group of actual audio recognition results obtained by annotating the first group of audio samples, and the initial audio recognition model is used to determine a predicted audio recognition based on input audio features;

[0282] S2: Inputting the audio features of the second group of audio samples into the initial audio recognition model to obtain a first group of predicted audio recognition results, wherein the second group of audio samples are not labeled with corresponding actual audio recognition results; inputting the audio features of the second group of audio samples into the uncertainty analysis model to obtain a first group of uncertainty analysis results, wherein the first group of uncertainty analysis results is used to indicate the credibility of the first group of predicted audio recognition results;

[0283] S3, based on the first set of uncertainty analysis results, screening out a second set of predicted audio recognition results whose credibility meets a preset condition from the first set of predicted audio recognition results, and screening out a third set of audio samples corresponding to the second set of predicted audio recognition results from the second set of audio samples;

[0284] S4, performing a current round of training on the initial audio recognition model based on the third set of audio samples and the second set of predicted audio recognition results, wherein the initial audio recognition model is set to undergo multiple rounds of training until a preset convergence condition is met. Or,

[0285] S1, obtain the input target audio in the target application;

[0286] S2. Obtain a target audio recognition result determined by a target audio recognition model according to the audio features of the target audio, wherein the target audio recognition model is an audio recognition model obtained by training the initial audio recognition model for multiple rounds until a preset convergence condition is met, the initial audio recognition model is a model obtained by training the audio recognition model to be trained using a first training sample set, the first training sample set includes a first group of audio samples and a first group of actual audio recognition results obtained by annotating the first group of audio samples, the initial audio recognition model is used to determine predicted audio recognition according to the input audio features, and in each round of training, the audio recognition model obtained by the previous round of training is trained using the training sample set corresponding to each round, the training sample set corresponding to each round includes the training sample set obtained by the previous round of training and the training sample set obtained by screening in this round, the training sample set obtained by screening in this round includes a group of audio samples and a group of predicted audio recognition results corresponding to the group of audio samples, a group of audio samples are not annotated with corresponding actual audio recognition results, and a group of predicted audio recognition results are predicted audio recognition results determined by the audio recognition model after the previous round of training according to the audio features of the group of audio samples;

[0287] S3, display the target audio recognition result in the target application. Optionally, those skilled in the art will understand that Figure 16 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 16 It does not limit the structure of the above electronic device. For example, the electronic device may also include Figure 16 More or fewer components (such as network interfaces, etc.) as shown in, or with Figure 16 Different configurations shown.

[0288] Among them, the memory 1602 can be used to store software programs and modules, such as the program instructions / modules corresponding to the training method and device of the audio recognition model in the embodiment of the present invention. The processor 1604 executes various functional applications and data processing by running the software programs and modules stored in the memory 1602, that is, realizes the above-mentioned training method of the audio recognition model. The memory 1602 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1602 may further include a memory remotely located relative to the processor 1604, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include but are not limited to the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 1602 can be used to store information such as the first group of audio samples, the second group of audio samples, and the audio recognition model, but is not limited to it. As an example, if Figure 16 As shown, the memory 1602 may include, but is not limited to, the first training unit 1402, the first screening unit 1406, the input unit 1404, and the second training unit 1408 in the training device for the audio recognition model. Furthermore, the memory 1602 may also include, but is not limited to, other module units in the training device for the audio recognition model, which will not be described in detail in this example.

[0289] Optionally, the transmission device 1606 is configured to receive or send data via a network. Specific examples of the network may include a wired network and a wireless network. In one embodiment, the transmission device 1606 includes a network interface controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In one embodiment, the transmission device 1606 is a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0290] In addition, the electronic device further includes: a display 1608 for displaying information such as the first group of audio samples, the second group of audio samples, and the audio recognition model; and a connection bus 1610 for connecting various module components in the electronic device.

[0291] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting multiple nodes via network communication. The nodes may form a peer-to-peer (P2P) network, and any computing device, such as a server, terminal, or other electronic device, may become a node in the blockchain system by joining the peer-to-peer network.

[0292] According to one aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned audio recognition model training and audio recognition method. The computer program is configured to execute the steps of any of the aforementioned method embodiments when executed.

[0293] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:

[0294] S1, training an audio recognition model to be trained using a first training sample set to obtain an initial audio recognition model, wherein the first training sample set includes a first group of audio samples and a first group of actual audio recognition results obtained by annotating the first group of audio samples, and the initial audio recognition model is used to determine a predicted audio recognition based on input audio features;

[0295] S2: Inputting the audio features of the second group of audio samples into the initial audio recognition model to obtain a first group of predicted audio recognition results, wherein the second group of audio samples are not labeled with corresponding actual audio recognition results; inputting the audio features of the second group of audio samples into the uncertainty analysis model to obtain a first group of uncertainty analysis results, wherein the first group of uncertainty analysis results is used to indicate the credibility of the first group of predicted audio recognition results;

[0296] S3, based on the first set of uncertainty analysis results, screening out a second set of predicted audio recognition results whose credibility meets a preset condition from the first set of predicted audio recognition results, and screening out a third set of audio samples corresponding to the second set of predicted audio recognition results from the second set of audio samples;

[0297] S4, performing a current round of training on the initial audio recognition model based on the third set of audio samples and the second set of predicted audio recognition results, wherein the initial audio recognition model is set to undergo multiple rounds of training until a preset convergence condition is met. Or,

[0298] S1, obtain the input target audio in the target application;

[0299] S2. Obtain a target audio recognition result determined by a target audio recognition model according to the audio features of the target audio, wherein the target audio recognition model is an audio recognition model obtained by training the initial audio recognition model for multiple rounds until a preset convergence condition is met, the initial audio recognition model is a model obtained by training the audio recognition model to be trained using a first training sample set, the first training sample set includes a first group of audio samples and a first group of actual audio recognition results obtained by annotating the first group of audio samples, the initial audio recognition model is used to determine predicted audio recognition according to the input audio features, and in each round of training, the audio recognition model obtained by the previous round of training is trained using the training sample set corresponding to each round, the training sample set corresponding to each round includes the training sample set obtained by the previous round of training and the training sample set obtained by screening in this round, the training sample set obtained by screening in this round includes a group of audio samples and a group of predicted audio recognition results corresponding to the group of audio samples, a group of audio samples are not annotated with corresponding actual audio recognition results, and a group of predicted audio recognition results are predicted audio recognition results determined by the audio recognition model after the previous round of training according to the audio features of the group of audio samples;

[0300] S3, displaying the target audio recognition result in the target application.

[0301] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing the hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0302] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0303] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing one or more computer devices (such as personal computers, servers, or network devices) to execute all or part of the steps of the methods described in various embodiments of the present invention.

[0304] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0305] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, and can be electrical or other forms.

[0306] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0307] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0308] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A training method for an audio recognition model, characterized in that: include: Training the audio recognition model to be trained using a first training sample set to obtain an initial audio recognition model, wherein the first training sample set includes a first group of audio samples and a first group of actual audio recognition results obtained by annotating the first group of audio samples, and the initial audio recognition model is used to determine a predicted audio recognition based on input audio features; Inputting the audio features of the second group of audio samples into the initial audio recognition model to obtain a first group of predicted audio recognition results, wherein the second group of audio samples are not labeled with corresponding actual audio recognition results; inputting the audio features of the second group of audio samples into the uncertainty analysis model to obtain a first group of uncertainty analysis results, wherein the first group of uncertainty analysis results is used to indicate the credibility of the first group of predicted audio recognition results; Based on the first group of uncertainty analysis results, screening out a second group of predicted audio recognition results whose credibility meets a preset condition from the first group of predicted audio recognition results, and screening out a third group of audio samples corresponding to the second group of predicted audio recognition results from the second group of audio samples; The initial audio recognition model is trained in a current round according to the third group of audio samples and the second group of predicted audio recognition results, wherein the initial audio recognition model is configured to undergo multiple rounds of training until a preset convergence condition is met.

2. The method according to claim 1, characterized in that The performing a current round of training on the initial audio recognition model according to the third group of audio samples and the second group of predicted audio recognition results includes: Merging the third set of audio samples and the second set of predicted audio recognition results into the first training sample set to obtain a second training sample set, wherein the second set of predicted audio recognition results in the second training sample set is regarded as a second set of actual audio recognition results; The initial audio recognition model is trained in a current round using the second training sample set to obtain an audio recognition model after the current round of training.

3. The method according to claim 2, characterized in that The method further comprises: When the difference between the predicted audio recognition results output by the audio recognition model after the current round of training and the actual audio recognition results in the second training sample set does not meet the convergence condition, obtaining a group of audio samples to be used in the next round of training and a group of predicted audio recognition results corresponding to the group of audio samples, wherein the group of audio samples to be used are not labeled with corresponding actual audio recognition results, and the group of predicted audio recognition results are predicted audio recognition results determined by the audio recognition model after the current round of training based on audio features of the group of audio samples; Merging the set of audio samples to be used and the corresponding set of predicted audio recognition results into the second training sample set to obtain a third training sample set; The audio recognition model after the current round of training is trained using the third training sample set to obtain the audio recognition model after the next round of training.

4. The method according to claim 3, characterized in that The obtaining of a set of audio samples to be used in the next round of training and a set of predicted audio recognition results corresponding to the set of audio samples includes: Inputting the audio features of the fourth group of audio samples into the audio recognition model after the current round of training to obtain a third group of predicted audio recognition results, wherein the fourth group of audio samples are not labeled with corresponding actual audio recognition results; inputting the audio features of the fourth group of audio samples into the uncertainty analysis model to obtain a second group of uncertainty analysis results, wherein the second group of uncertainty analysis results is used to indicate the credibility of the third group of predicted audio recognition results; According to the second group of uncertainty analysis results, a fourth group of predicted audio recognition results whose credibility meets preset conditions is screened out from the third group of predicted audio recognition results, and a fifth group of audio samples corresponding to the fourth group of predicted audio recognition results is screened out from the fourth group of audio samples.

5. The method according to claim 3, characterized in that The step of screening out a second group of predicted audio recognition results whose credibility meets a preset condition from the first group of predicted audio recognition results based on the first group of uncertainty analysis results includes: When the difference between the predicted audio recognition results output by the audio recognition model after the next round of training and the actual audio recognition results in the third training sample set meets the convergence condition, the training of the initial audio recognition model is terminated to obtain the target audio recognition model, wherein the set of predicted audio recognition results in the third training sample set is regarded as a set of actual audio recognition results.

6. The method according to any one of claims 1 to 5, characterized in that The step of screening out a second group of predicted audio recognition results whose credibility meets a preset condition from the first group of predicted audio recognition results based on the first group of uncertainty analysis results includes: When the first set of uncertainty analysis results includes a set of uncertainty scores, sorting the set of uncertainty scores in ascending order to obtain an uncertainty score sequence, wherein a higher uncertainty score indicates a lower credibility of the corresponding predicted audio recognition result; Obtain the top N uncertainty scores in the uncertainty score sequence, wherein the uncertainty score sequence includes M uncertainty scores, N <M; The second group of predicted audio recognition results corresponding to the top N uncertainty scores in the ranking are screened out from the first group of predicted audio recognition results.

7. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Get the input target audio in the target application; Obtaining a target audio recognition result determined by a target audio recognition model based on the audio features of the target audio, wherein the target audio recognition model is an audio recognition model obtained by training the initial audio recognition model for multiple rounds until a preset convergence condition is satisfied; The target audio recognition result is displayed in the target application.

8. The method according to claim 7, characterized in that The acquiring of the input target audio in the target application includes: when a reference text is displayed in the target application, acquiring the target audio generated by reading the reference text in the target application, or acquiring the target audio generated by replying to the reference text; Displaying the target audio recognition result in the target application includes: displaying an evaluation score of the target audio determined by the target audio recognition model in the target application.

9. An audio recognition method, characterized in that: include: Get the input target audio in the target application; Obtain a target audio recognition result determined by a target audio recognition model based on the audio features of the target audio, wherein the target audio recognition model is an audio recognition model obtained by training an initial audio recognition model for multiple rounds until a preset convergence condition is met, the initial audio recognition model is a model obtained by training the audio recognition model to be trained using a first training sample set, the first training sample set including a first group of audio samples and a first group of actual audio recognition results obtained by annotating the first group of audio samples, the initial audio recognition model is used to determine predicted audio recognition based on the input audio features, the initial audio recognition model is further used to perform initial audio recognition on a second group of unannotated audio samples to obtain a first group of predicted audio recognition results, the second group of audio samples is further used to perform uncertainty analysis using a trained uncertainty analysis model to obtain an analysis result, the analysis result is used to screen the first group of predicted audio recognition results to obtain a second group of predicted audio recognition results, and a third group of audio samples in the second group of audio samples corresponding to the second group of predicted audio recognition results is used to iteratively train the initial audio recognition model until the preset convergence condition is met to obtain a trained target audio recognition model; The target audio recognition result is displayed in the target application.

10. The method according to claim 9, characterized in that Before obtaining the input target audio in the target application, include: Training the audio recognition model to be trained using the first training sample set to obtain the initial audio recognition model, wherein the first training sample set includes a first group of audio samples and a first group of actual audio recognition results obtained by annotating the first group of audio samples, and the initial audio recognition model is used to determine predicted audio recognition based on input audio features; Inputting the audio features of the second group of audio samples into the initial audio recognition model to obtain the first group of predicted audio recognition results, wherein the second group of audio samples are not labeled with corresponding actual audio recognition results; Inputting the audio features of the second group of audio samples into the uncertainty analysis model to obtain a first group of uncertainty analysis results, wherein the first group of uncertainty analysis results is used to indicate the credibility of the first group of predicted audio recognition results; Based on the first group of uncertainty analysis results, screening out a second group of predicted audio recognition results whose credibility meets a preset condition from the first group of predicted audio recognition results, and screening out a third group of audio samples corresponding to the second group of predicted audio recognition results from the second group of audio samples; The initial audio recognition model is trained in a current round according to the third group of audio samples and the second group of predicted audio recognition results, wherein the initial audio recognition model is configured to undergo multiple rounds of training until a preset convergence condition is met.

11. The method according to claim 10, characterized in that The performing a current round of training on the initial audio recognition model according to the third group of audio samples and the second group of predicted audio recognition results includes: Merging the third set of audio samples and the second set of predicted audio recognition results into the first training sample set to obtain a second training sample set, wherein the second set of predicted audio recognition results in the second training sample set is regarded as a second set of actual audio recognition results; The initial audio recognition model is trained in a current round using the second training sample set to obtain an audio recognition model after the current round of training.

12. The method according to claim 11, characterized in that The method further comprises: When the difference between the predicted audio recognition results output by the audio recognition model after the current round of training and the actual audio recognition results in the second training sample set does not meet the convergence condition, obtaining a group of audio samples to be used in the next round of training and a group of predicted audio recognition results corresponding to the group of audio samples, wherein the group of audio samples to be used are not labeled with corresponding actual audio recognition results, and the group of predicted audio recognition results are predicted audio recognition results determined by the audio recognition model after the current round of training based on audio features of the group of audio samples; Merging the set of audio samples to be used and the corresponding set of predicted audio recognition results into the second training sample set to obtain a third training sample set; The audio recognition model after the current round of training is trained using the third training sample set to obtain the audio recognition model after the next round of training.

13. The method according to claim 12, characterized in that The obtaining of a set of audio samples to be used in the next round of training and a set of predicted audio recognition results corresponding to the set of audio samples includes: Inputting the audio features of the fourth group of audio samples into the audio recognition model after the current round of training to obtain a third group of predicted audio recognition results, wherein the fourth group of audio samples are not labeled with corresponding actual audio recognition results; inputting the audio features of the fourth group of audio samples into the uncertainty analysis model to obtain a second group of uncertainty analysis results, wherein the second group of uncertainty analysis results is used to indicate the credibility of the third group of predicted audio recognition results; According to the second group of uncertainty analysis results, a fourth group of predicted audio recognition results whose credibility meets preset conditions is screened out from the third group of predicted audio recognition results, and a fifth group of audio samples corresponding to the fourth group of predicted audio recognition results is screened out from the fourth group of audio samples.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, wherein the program executes the method described in any one of claims 1 to 11 when executed.

15. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 11 through the computer program.

Citation Information

Patent Citations

  • An anti-disturbance generation method and device for an object detection model

    CN109902705A

  • Voice data processing method and device, electronic equipment and computer readable medium

    CN110992938A