A data processing method, device, computer device and readable storage medium

By performing multiple noise data addition processing on the original audio data, and generating multiple sample audio data for feature extraction and model training, the problem of insufficient training data for oral examination scoring model in the existing technology is solved, and the scoring accuracy is improved.

CN115206342BActive Publication Date: 2025-05-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110396808.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-13
Publication Date
2025-05-27
Estimated Expiration
2041-04-13

AI Technical Summary

Technical Problem

In the prior art, the training data of the oral examination scoring model is insufficient, resulting in low scoring accuracy and affecting the test scores.

Method used

By performing multiple noise data noise processing on the original audio data, multiple sample audio data are generated, feature extraction and model training are performed, and audio evaluation models are generated to improve scoring accuracy.

Benefits of technology

The amount of data for model training is amplified, the accuracy of model prediction is improved, and the accuracy of oral test scores is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115206342B_ABST
    Figure CN115206342B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses a data processing method, apparatus, computer device, and readable storage medium, which relate to voice technology in artificial intelligence. The method includes: obtaining the original audio data of a target user, performing noise addition processing on the original audio data respectively using N types of noise data to obtain sample audio data respectively corresponding to the N types of noise data; N is a positive integer; respectively performing feature extraction on the N sample audio data to obtain sample data features respectively corresponding to the N sample audio data; obtaining a score label corresponding to the original audio data, and training an initial evaluation model based on the sample data features respectively corresponding to the N sample audio data and the score label to generate an audio evaluation model; the audio evaluation model is used to predict a target score corresponding to target audio data. By using the embodiment of the present application, the data volume can be amplified, thereby improving the accuracy of model prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech technology in artificial intelligence, and particularly to a data processing method, apparatus, computer device, and readable storage medium. Background Art

[0002] Artificial intelligence has been widely applied to various devices, enabling the devices to have functions such as question-and-answer and scoring according to users' answers. Question-and-answer scoring can be applied to various scoring systems, for example, a Putonghua examination system, an oral English examination system, and various systems that require language tests. Currently, more and more provinces and cities have included oral English examinations in the scope of high school entrance examinations and college entrance examinations, and the importance of oral examinations is self-evident. How to improve the accuracy of oral examination scoring is an urgent problem to be solved.

[0003] In the prior art, generally, the data corresponding to the oral examination is used to train an evaluation model, and the trained evaluation model is used to score the data to be evaluated in the oral examination. However, due to the small amount of data for each set of oral test questions, the model training process is insufficient, resulting in low accuracy of using the model for score prediction, which in turn has a greater impact on students' examination results. Summary of the Invention

[0004] Embodiments of this application provide a data processing method, apparatus, computer device, and readable storage medium, which can increase the amount of data and improve the accuracy of model prediction.

[0005] On the one hand, an embodiment of this application provides a data processing method, including:

[0006] Obtain the original audio data of the target user, and perform noise addition processing on the original audio data using N kinds of noise data respectively to obtain the sample audio data corresponding to the N kinds of noise data respectively; N is a positive integer;

[0007] Extract features from the N sample audio data respectively to obtain the sample data features corresponding to the N sample audio data respectively;

[0008] Obtain the score label corresponding to the original audio data, and train an initial evaluation model based on the sample data features corresponding to the N sample audio data respectively and the score label to generate an audio evaluation model; the audio evaluation model is used to predict the target score corresponding to the target audio data.

[0009] On the one hand, an embodiment of this application provides a data processing apparatus, including:

[0010] An original data acquisition module, configured to obtain the original audio data of the target user, and perform noise addition processing on the original audio data using N kinds of noise data respectively to obtain the sample audio data corresponding to the N kinds of noise data respectively; N is a positive integer;

[0011] A feature extraction module for separately extracting features from N sample audio data to obtain sample data features corresponding to the N sample audio data respectively;

[0012] A model generation module for obtaining the score label corresponding to the original audio data, training an initial evaluation model based on the sample data features corresponding to the N sample audio data and the score label, and generating an audio evaluation model; the audio evaluation model is used to predict the target score corresponding to the target audio data.

[0013] Optionally, the N sample audio data includes sample audio data i; i is a positive integer; the feature extraction module includes:

[0014] A speech extraction unit for extracting speech features from the sample audio data i to obtain sample speech features corresponding to the sample audio data i;

[0015] A text extraction unit for performing speech conversion processing on the sample audio data i to obtain sample text data corresponding to the sample audio data i, and extracting text features from the sample text data to obtain sample text features corresponding to the sample audio data i;

[0016] A feature splicing unit for splicing the sample speech features and the sample text features to generate sample data features corresponding to the sample audio data i.

[0017] Optionally, the speech extraction unit is specifically used for:

[0018] Obtaining the speech fluency corresponding to the sample audio data i, and determining the first speech feature based on the speech fluency;

[0019] Obtaining the phoneme sequence corresponding to the sample audio data i, and determining the second speech feature based on the phoneme sequence corresponding to the sample audio data i;

[0020] Obtaining the pronunciation accuracy corresponding to the sample audio data i, and determining the third speech feature based on the pronunciation accuracy;

[0021] Determining the sample speech features corresponding to the sample audio data i based on the first speech feature, the second speech feature and the third speech feature.

[0022] Optionally, the N sample audio data all include the audio data to be evaluated and the reference audio data, and the sample data features corresponding to the N sample audio data respectively include the data features to be evaluated corresponding to the audio data to be evaluated and the reference audio features corresponding to the reference audio data; the model generation module includes:

[0023] A similarity determination unit, configured to input the sample data features respectively corresponding to the N sample audio data into the initial evaluation model, determine the audio similarity between the data features to be evaluated corresponding to each sample audio data and the reference audio feature based on the initial evaluation model, and obtain a sample prediction score according to the audio similarity;

[0024] A model adjustment unit, configured to adjust the initial evaluation model based on the difference value between the score label and the sample prediction score, and generate the audio evaluation model.

[0025] Optionally, the apparatus further includes:

[0026] A sample division module, configured to divide the N sample audio data into training sample data and validation sample data;

[0027] The model generation module includes:

[0028] A model training unit, configured to train the initial evaluation model based on the training sample data to generate a to-be-detected evaluation model;

[0029] A quality determination unit, configured to detect the to-be-detected evaluation model based on the validation sample data to obtain the model quality corresponding to the to-be-detected evaluation model;

[0030] A model determination unit, configured to, if the model quality is greater than or equal to the model effective threshold, determine the to-be-detected evaluation model as the audio evaluation model.

[0031] Optionally, the apparatus further includes a model adjustment module, including:

[0032] A data acquisition unit, configured to acquire target audio data generated by the target user for the target service;

[0033] A feature extraction unit, configured to extract features from the target audio data to obtain target audio features corresponding to the target audio data;

[0034] A score determination unit, configured to input the target audio features into the audio evaluation model, and predict the target audio features based on the audio evaluation model to obtain a target score corresponding to the target audio features;

[0035] Specifically, if the target score is greater than or equal to the service qualification threshold, the score determination unit is configured to send a service processing success message to the target user;

[0036] Specifically, if the target score is less than the service qualification threshold, the score determination unit is configured to send a service processing failure message to the target user, and the service processing failure message is used to instruct the target user to regenerate audio data for the target service within the target time range.

[0037] Optionally, the device further includes a model optimization module, including:

[0038] A request acquisition unit, configured to receive an appeal request from the target user for the target score;

[0039] A request review unit, configured to send the target score to a test and evaluation terminal for review based on the appeal request;

[0040] A model optimization unit, configured to receive the review result sent by the test and evaluation terminal and adjust the audio evaluation model based on the review result.

[0041] On the one hand, the present application provides a computer device, including: a processor, a memory, and a network interface;

[0042] The above-mentioned processor is connected to the memory and the network interface. Among them, the network interface is used to provide a data communication function, the above-mentioned memory is used to store a computer program, and the above-mentioned processor is used to call the above-mentioned computer program so that the computer device including the processor executes the above-mentioned method.

[0043] On the one hand, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. The computer program is suitable for being loaded and executed by a processor so that a computer device having the processor executes the above-mentioned method.

[0044] On the one hand, an embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various alternative manners in an embodiment of the present application.

[0045] In the embodiments of the present application, by obtaining the original audio data of the target user, adding noise to the original audio data using N types of noise data respectively to obtain the sample audio data corresponding to the N types of noise data respectively; extracting features from the N sample audio data respectively to obtain the sample data features corresponding to the N sample audio data respectively; obtaining the score label corresponding to the original audio data, and training the initial evaluation model based on the sample data features and the score label corresponding to the N sample audio data respectively to generate an audio evaluation model; the audio evaluation model is used to predict the target score corresponding to the target audio data. Since adding noise to the original audio data using multiple types of noise data respectively to obtain the sample audio data corresponding to the multiple types of noise data respectively can increase the amount of data for training the audio evaluation model; therefore, using the amplified large amount of sample audio data to train and predict the model can improve the accuracy of model prediction, and further improve the accuracy of audio data scoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0047] Figure 1 It is a schematic diagram of the architecture of a data processing system provided by an embodiment of the present application;

[0048] Figure 2 It is a schematic diagram of the application scenario of a data processing method provided by an embodiment of the present application;

[0049] Figure 3 It is a schematic diagram of the flow of a data processing method provided by an embodiment of the present application;

[0050] Figure 4 It is a schematic diagram of the process of model training and predicting scores provided by an embodiment of the present application;

[0051] Figure 5 It is a schematic diagram of the flow of another data processing method provided by an embodiment of the present application;

[0052] Figure 6 It is a schematic diagram of the process of adjusting the audio evaluation model provided by an embodiment of the present application;

[0053] Figure 7 It is a schematic diagram of the process of determining the data volume provided by an embodiment of the present application;

[0054] Figure 8 It is a schematic diagram of the composition structure of a data processing device provided by an embodiment of the present application;

[0055] Figure 9 It is a schematic diagram of the composition structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0056] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without making creative efforts shall fall within the protection scope of the present application.

[0057] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, mechatronics, etc. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech technology, natural language processing technology, and machine learning / deep learning.

[0058] Among them, the key technologies of speech technology include automatic speech recognition technology (ASR), text-to-speech technology (TTS), and voiceprint recognition technology. Enabling the computer to listen, see, speak, and feel is the future development direction of human-computer interaction, and among them, speech has become one of the most promising human-computer interaction methods in the future.

[0059] This application relates to voice technology in artificial intelligence. The sample audio data is used to extract features through voice technology to obtain the sample data features corresponding to the sample audio data; the score label corresponding to the original audio data is obtained, and the initial evaluation model is trained based on the sample data features and the score label corresponding to the sample audio data to generate an audio evaluation model, and the target score corresponding to the target audio data is predicted based on the audio evaluation model. The technical solution of this application can be used in scenarios where the sample audio data is predicted to obtain the predicted score corresponding to the sample audio data. For example, it can be used in Putonghua proficiency tests, spoken English tests, other language tests, or other scenarios that require language tests. By obtaining the original audio data of the target user, N kinds of noise data are used to add noise to the original audio data respectively to obtain the sample audio data corresponding to the N kinds of noise data respectively; N is a positive integer. Further, feature extraction is performed on the N sample audio data respectively to obtain the sample data features corresponding to the N sample audio data respectively. Further, the score label corresponding to the original audio data is obtained, and the initial evaluation model is trained based on the sample data features and the score label corresponding to the N sample audio data respectively to generate an audio evaluation model; the audio evaluation model is used to predict the target score corresponding to the target audio data. Since N kinds of noise data are used to add noise to the original audio data respectively to obtain the sample audio data corresponding to the N kinds of noise data respectively, the data volume for training the audio evaluation model can be amplified; therefore, using the amplified large amount of sample audio data to train and predict the model can improve the accuracy of model prediction, and further improve the accuracy of audio data scoring.

[0060] Please refer to Figure 1 , Figure 1 which is the network architecture diagram of a data processing system provided by an embodiment of this application. As Figure 1 shown, the computer device 101 can interact with the user terminal corresponding to the target user. The number of user terminals can be one or more. For example, when the number of user terminals is multiple, the user terminals can include Figure 1User terminals 102a, 102b, 102c, etc. among them. Among them, taking the user terminal corresponding to the target user as the user terminal 102a as an example, the computer device 101 can obtain the original audio data sent by the user terminal 102a. The computer device 101 can perform noise addition processing on the original audio data respectively using N kinds of noise data to obtain sample audio data corresponding to the N kinds of noise data respectively; where N is a positive integer. Further, the computer device 101 can perform feature extraction on the N sample audio data respectively to obtain sample data features corresponding to the N sample audio data respectively. Further, the computer device 101 can obtain the score label corresponding to the original audio data, and train the initial evaluation model based on the sample data features and the score label corresponding to the N sample audio data respectively to generate an audio evaluation model; where the audio evaluation model is used to predict the target score corresponding to the target audio data. Optionally, the computer device 101 can also send the target score to the user terminal 102a, and the target user can view the corresponding target score through the user terminal 102a.

[0061] Since noise addition processing is performed on the original audio data respectively using multiple kinds of noise data to obtain sample audio data corresponding to the multiple kinds of noise data respectively, the data volume for training the audio evaluation model can be amplified; therefore, using the amplified large amount of sample audio data to train and predict the model can improve the accuracy of model prediction, and further improve the accuracy of audio data scoring.

[0062] It can be understood that the computer device mentioned in the embodiments of the present application includes but is not limited to a terminal device or a server. In other words, the computer device or the user terminal can be a server or a terminal device, or a system composed of a server and a terminal device. Among them, the above-mentioned terminal device can be an electronic device, including but not limited to a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a vehicle-mounted device, an Augmented Reality / Virtual Reality (AR / VR) device, a head-mounted display, a wearable device, a smart speaker, a digital camera, a camera, and other mobile internet devices (MID) with network access capabilities, etc. Among them, the above-mentioned server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, vehicle-road collaboration, Content Delivery Network (CDN), as well as big data and artificial intelligence platforms.

[0063] Further, please refer to Figure 2 , Figure 2 , which is a schematic diagram of an application scenario of a data processing method provided by an embodiment of the present application. As Figure 2 shown, the user terminal 20 corresponding to the target user sends the original audio data 21 to the computer device 22, and the computer device 22 performs noise addition processing on the original audio data 21 using N kinds of noise data respectively to obtain sample audio data 23 corresponding to the N kinds of noise data respectively; the computer device 22 performs feature extraction on the sample audio data 23 to obtain sample data features 24 corresponding to the sample audio data 23, where the sample data features may include sample voice features and sample text features. Further, the computer device 22 obtains a score label corresponding to the original audio data 21, and trains an initial evaluation model based on the sample data features 24 corresponding to the sample audio data 23 and the score label to generate an audio evaluation model. Specifically, the computer device 22 may input the sample data features 24 into the initial evaluation model to obtain sample prediction scores, and train the initial evaluation model based on the difference value between the score label and the sample prediction scores to generate an audio evaluation model. Optionally, the computer device 22 may also predict a target score corresponding to the target audio data based on the audio evaluation model and send the target score to the user terminal 20 corresponding to the target user.

[0064] Further, please refer to Figure 3 , Figure 3 , which is a schematic flowchart of a data processing method provided by an embodiment of the present application. This method can be applied to a computer device; as Figure 3 shown, this method includes:

[0065] S101, obtain the original audio data of the target user, and perform noise addition processing on the original audio data using N kinds of noise data respectively to obtain sample audio data corresponding to the N kinds of noise data respectively.

[0066] In the embodiments of the present application, the computer device may obtain the original audio data of the target user from the user terminal corresponding to the target user, or may obtain the original video data of one or more target users from the audio database storing the original audio data of multiple target users. Among them, the original audio data may refer to the audio data generated by the target user for the target service. For example, the target service may include, but is not limited to, the Putonghua proficiency test, the spoken English test, or other language test corresponding services. Then, the target audio data may refer to the audio data corresponding to the Putonghua proficiency test, the spoken English test, or other language test corresponding audio data, etc. Further, the computer device may perform noise addition processing on the original audio data respectively using N kinds of noise data to obtain the sample audio data corresponding to the N kinds of noise data respectively. Among them, N is a positive integer. That is to say, if the number of the original audio data is 1, the number of the sample audio data obtained after performing noise addition processing on the original audio data respectively using N kinds of noise data is N; if the number of the original audio data is M, the number of the sample audio data obtained after performing noise addition processing on the original audio data respectively using N kinds of noise data is N*M, and M is a positive integer. Here, the value of N can be determined according to specific requirements. For example, the value of N can be a multiple of the number of the original audio data, such as 5 times, 10 times, or 20 times the number of the original audio data, etc. For example, the number of the original audio data is 237, and N is 5, then the number of the sample audio data after the noise addition processing is 237*(1 + 5) = 1422.

[0067] In a specific implementation, after obtaining the original audio data, the computer device can obtain multiple different types of noise data to perform noise addition processing on the original audio data respectively, and obtain the sample audio data. Among them, the N types of noise data can include noise data of different noise types, including but not limited to Gaussian noise data, constant noise data, or other types of noise data, etc., or can also include noise data of the same type but different decibels. That is to say, the noise types of the N types of noise data are not completely the same. Optionally, the computer device can convert the noise data into a signal with the same signal strength as the original audio data, so as to add the noise data to the original audio data to obtain the sample audio data after noise addition processing. Or, the computer device can also use other noise addition methods to implement noise addition processing on the original audio data, which is not limited in the embodiments of this application. Optionally, the computer device can also receive the original audio data after audio encoding, and perform noise addition processing on the original audio data after audio encoding to obtain the sample audio data. Since the audio features corresponding to the audio data are generally composed of multi-dimensional vectors, each dimension vector corresponds to an element, and there is an inherent correlation between different elements corresponding to the audio features. In the embodiments of this application, noise addition processing is directly performed on the original audio data, that is, noise addition processing is performed before feature extraction of the audio data. Therefore, when subsequent feature extraction is performed on the sample audio data after noise addition processing, it will not affect the correlation between the elements corresponding to the audio features, so that the extracted audio features can retain as much data information of the sample audio data as possible, thereby improving the accuracy of model training.

[0068] Optionally, for example, in the scenario of an oral English test, a variety of question types may be included, such as a follow-up question type, a question-and-answer question type, a semi-open question type, a situational question type, and a picture description question type, etc. Taking the follow-up question type as an example, the sentence played in the computer device is "What's the weather yesterday?" ("What was the weather like yesterday?"), and the examinee needs to read the sentence aloud, and the recording device in the examination room can record the sound of the examinee reading the sentence aloud to obtain the original audio data of the target user. Or, taking the question-and-answer question type as an example, the question sentence played in the computer device is "What's the weather yesterday?" The examinee needs to answer based on the question sentence, for example, the answer sentence answered by the examinee is "I trained heavily yesterday" ("It rained heavily yesterday"), then the recording device in the examination room can record the sound of the examinee's answer "It rained heavily yesterday" to obtain the original audio data of the target user, etc. The examinee makes corresponding responses to different question types, and the recording device records the examinee's response data to obtain the original audio data of the target user. The recording device can be a part of the computer device, and when the recording device records the candidate's voice data, the computer device obtains the original audio data of the target user. Alternatively, the recording device can also be a device independent of the computer device, and the recording device records the user's voice data and sends it to the computer device, so that the computer device obtains the original audio data of the target user.

[0069] S102, performing feature extraction on N sample audio data respectively to obtain sample data features corresponding to the N sample audio data respectively.

[0070] In an embodiment of the present application, the computer device may perform feature extraction on each of the N sample audio data to obtain sample data features corresponding to each of the N sample audio data. The sample data features may include sample audio features and sample text features, the sample audio features may be determined based on the sample audio data, and the sample text features may be determined based on the sample text data corresponding to the sample audio data.

[0071] Optionally, the N sample audio data includes sample audio data i; i is a positive integer, that is, sample audio data i is any one of the N sample audio data. The computer device can extract features from the sample audio data i to obtain the sample data features corresponding to the sample audio data i. Specifically, the computer device can perform speech feature extraction on the sample audio data i to obtain the sample speech features corresponding to the sample audio data i; perform speech conversion processing on the sample audio data i to obtain the sample text data corresponding to the sample audio data i, perform text feature extraction on the sample text data to obtain the sample text features corresponding to the sample audio data i; splice the sample speech features and the sample text features to generate the sample data features corresponding to the sample audio data i.

[0072] In a specific implementation, the computer device can obtain the speech fluency corresponding to the sample audio data i, and determine the first speech feature based on the speech fluency; obtain the phoneme sequence corresponding to the sample audio data i, and determine the second speech feature based on the phoneme sequence corresponding to the sample audio data i; obtain the pronunciation accuracy corresponding to the sample audio data i, and determine the third speech feature based on the pronunciation accuracy; and determine the sample speech feature corresponding to the sample audio data i based on the first speech feature, the second speech feature, and the third speech feature. Among them, the computer device can determine the speech fluency of the sample audio data i according to the pause duration between every two words in the sample audio data i. The pause duration can be determined according to the duration without sound between the end of the pronunciation of the previous word and the start of the next word between two adjacent words. When the pause duration is greater than the pause duration threshold, it indicates low speech fluency; when the pause duration is less than or equal to the pause duration threshold, it indicates high speech fluency. By determining the pause duration between words in the sample audio data i, the speech fluency corresponding to the sample audio data i can be determined, thereby obtaining the first speech feature. The phoneme sequence can be composed of each smallest speech unit included in the sample audio data i. The computer device can obtain each smallest speech unit in the sample audio data i and combine the smallest speech units included in the sample audio data i into a sequence to obtain the phoneme sequence. For example, there are 48 phonemes in English, including vowel phonemes and consonant phonemes. The vowel phonemes include monophthong phonemes and diphthong phonemes. The computer device can determine the pronunciation accuracy corresponding to the sample audio data i according to the pronunciation of the smallest speech unit in the sample audio data i. It can be seen that the first speech feature, the second speech feature, and the third speech feature are the sample speech features corresponding to the sample audio data i obtained by extracting features from the sample audio data i from three different dimensions. Optionally, the computer device can also obtain the integrity, prosody, etc. corresponding to the sample audio data i, obtain the corresponding features based on the integrity and the prosody respectively, and combine the features corresponding to the integrity of the audio data i, the features corresponding to the prosody, the first speech feature, the second speech feature, and the third speech feature to determine the sample speech feature corresponding to the sample audio data i, so as to realize feature extraction from more dimensions for the sample audio data i, make the extracted features more complete, and thus obtain a more accurate score. Among them, the integrity can refer to the completeness of each word in the sample audio data i, and the prosody can refer to the weak stress, strong stress, and liaison of words in the sample audio data i, etc.

[0073] In a specific implementation, the computer device may perform speech conversion processing on the sample audio data i based on automatic speech recognition technology, convert the audio data into text data, and obtain the sample text data corresponding to the sample audio data i. Alternatively, the computer device may also use other methods to implement the conversion of the sample audio data into the sample text data, which is not limited in the embodiments of the present application. The computer device may splice the sample speech features and the sample text features to generate the sample data features corresponding to the sample audio data i. Alternatively, the computer device may also fuse the sample speech features and the sample text features based on a fusion algorithm to generate the sample data features corresponding to the sample audio data i. The computer device may also use other methods to implement the splicing of the sample speech features and the sample text features, which is not limited in the embodiments of the present application.

[0074] S103. Obtain the score label corresponding to the original audio data, and train the initial evaluation model based on the sample data features and the score labels respectively corresponding to the N sample audio data to generate an audio evaluation model.

[0075] In the embodiments of the present application, the computer device may obtain the score label corresponding to the original audio data, and train the initial evaluation model based on the sample data features and the score labels respectively corresponding to the N sample audio data to generate an audio evaluation model. The computer device may also predict the target score corresponding to the target audio data based on the audio evaluation model, so as to implement the scoring of the audio data. Among them, the score label corresponding to the original audio data may refer to the score given by an expert to the original audio data.

[0076] Optionally, the N sample audio data all include the audio data to be evaluated and the reference audio data. The sample data features respectively corresponding to the N sample audio data include the data features to be evaluated corresponding to the audio data to be evaluated and the reference audio features corresponding to the reference audio data. Among them, the audio data to be evaluated may refer to the audio data corresponding to the target user, such as the audio data generated by a candidate during an exam in response to exam questions. The reference audio data may refer to the audio data corresponding to the reference answer, such as the audio data recorded when an expert reads the reference answer aloud in standard spoken English. The method for obtaining the data features to be evaluated corresponding to the audio data to be evaluated and the reference audio features corresponding to the reference audio data may refer to the method for obtaining the sample data features corresponding to the sample audio data in the foregoing steps, and will not be described in detail here. Specifically, the computer device inputs the sample data features respectively corresponding to the N sample audio data into the initial evaluation model, determines the audio similarity between the data features to be evaluated and the reference audio features corresponding to each sample audio data based on the initial evaluation model, and obtains the sample prediction score according to the audio similarity; based on the difference value between the score label and the sample prediction score, adjust the initial evaluation model to generate an audio evaluation model.

[0077] In a specific implementation, the computer device can obtain the sample data features corresponding to the sample audio data i in the N sample audio data, input the sample data features corresponding to the sample audio data i into the initial evaluation model, determine the audio similarity between the data features to be evaluated corresponding to the sample audio data i and the reference audio features based on the initial evaluation model, and obtain the sample prediction score according to the audio similarity. Optionally, the computer device can separately obtain the feature vector corresponding to the data features to be evaluated and the feature vector corresponding to the reference audio features, calculate the audio similarity between the feature vector corresponding to the data features to be evaluated and the feature vector corresponding to the reference audio features. The higher the audio similarity, the higher the similarity between the data features to be evaluated and the reference audio features; the lower the audio similarity, the lower the similarity between the data features to be evaluated and the reference audio features. Thus, the sample prediction score can be determined according to the audio similarity between the data features to be evaluated and the reference audio features. For example, the higher the audio similarity, the higher the corresponding sample prediction score; the lower the audio similarity, the lower the corresponding sample prediction score. The computer device adjusts the initial evaluation model according to the difference value between the score label and the sample prediction score to generate an audio evaluation model. The difference value between the score label and the sample prediction score can refer to the absolute difference between the score label and the sample prediction score. Optionally, the closer the score value corresponding to the score label is to the sample prediction score, the higher the accuracy of scoring based on the initial evaluation model. Then, the initial evaluation model at this time can be saved to obtain the audio evaluation model. The greater the difference value between the score value corresponding to the score label and the sample prediction score, the lower the accuracy of scoring based on the initial evaluation model. Therefore, continue to adjust the initial evaluation model at this time, adjust the parameters in the initial evaluation model to improve the accuracy of the model.

[0078] The above process is for the processing of the sample audio data i in the N sample audio data. The processing method for other sample audio data in the N sample audio data can refer to the processing method of the sample audio data i, so as to obtain the sample prediction scores corresponding to each sample audio data in the N sample audio data. Then, according to the difference value between the sample prediction score corresponding to each sample audio data and the score label corresponding to each sample audio data, the initial evaluation model is adjusted to generate an audio evaluation model.

[0079] Optionally, the computer device may divide the N sample audio data into training sample data and validation sample data, then train the initial evaluation model based on the training sample data, and validate the initial evaluation model based on the validation sample data to obtain an audio evaluation model. Specifically, the computer device divides the N sample audio data into training sample data and validation sample data; trains the initial evaluation model based on the training sample data to generate a to-be-detected evaluation model; detects the to-be-detected evaluation model based on the validation sample data to obtain the model quality corresponding to the to-be-detected evaluation model; if the model quality is greater than or equal to the model effective threshold, determines the to-be-detected evaluation model as the audio evaluation model. Among them, the ratio between the training sample data and the validation sample data can be determined according to requirements. For example, in the case of high precision requirements for the model, the number of training sample data can be increased, etc., which is not limited here. The model quality corresponding to the to-be-detected evaluation model is used to indicate the accuracy of the model. When the model quality is greater than or equal to the model effective threshold, it means that the accuracy of the model is high and can be used normally. Therefore, the to-be-detected evaluation model can be determined as the audio evaluation model. When the model quality is less than the model effective threshold, it means that the accuracy of the model is low, and using the model at this time will lead to inaccurate model prediction. Therefore, the initial evaluation model can be continuously trained to adjust the parameters in the initial evaluation model, so that the model quality corresponding to the generated to-be-detected evaluation model is greater than the model effective threshold, thereby improving the accuracy of model prediction.

[0080] During the model training process, in order to evaluate the model training effect and prevent overfitting, a training set (training sample data) and a validation set (validation sample data) are usually divided. Since the available training set data is too small before data augmentation, in order to ensure the model training effect, the proportion of the validation set will be very small to increase the amount of data in the training set; however, if the validation set is too small, it may lead to insufficient model validation and the training process ends in an unstable state. Therefore, after using the technical solution in this application to perform data enhancement on the original audio data, the obtained data quantity increases, so the data quantity used for model training increases significantly. By increasing the proportion of the validation set, the model can converge to a more accurate state, so that when using the model for score prediction, the prediction accuracy can be improved.

[0081] Optionally, refer to Figure 4 , Figure 4 which is a schematic flowchart of a model training and prediction score provided by an embodiment of this application. As Figure 4 shown, the method includes:

[0082] S201, obtain sample audio data, reference text data, and score labels.

[0083] Among them, the score label can be obtained by an expert scoring the sample audio data.

[0084] S202. Extract features from the sample audio data and the reference text data to obtain the sample data features corresponding to the sample audio data.

[0085] S203. Input the sample data features and the score label into the initial evaluation model for training to generate an audio evaluation model.

[0086] S204. Obtain the audio data to be evaluated and the reference text data, input the audio data to be evaluated and the reference text data into the audio evaluation model for prediction, and output the target score.

[0087] Optionally, the processes of model training and model prediction score can be implemented by different computer devices respectively, or can also be implemented using the same computer device, which is not limited in the embodiments of the present application. By using the sample data features and the score label corresponding to the sample audio data to input into the initial evaluation model for training to generate an audio evaluation model, the accuracy of model prediction can be improved through training the model, and the scoring efficiency can be improved in the subsequent process of using the audio evaluation model to predict the score of audio data.

[0088] In the embodiments of the present application, by obtaining the original audio data of the target user, adding noise to the original audio data respectively with N kinds of noise data to obtain the sample audio data corresponding to the N kinds of noise data respectively; extracting features from the N sample audio data respectively to obtain the sample data features corresponding to the N sample audio data respectively; obtaining the score label corresponding to the original audio data, and training the initial evaluation model based on the sample data features and the score label corresponding to the N sample audio data respectively to generate an audio evaluation model; the audio evaluation model is used to predict the target score corresponding to the target audio data. Since adding noise to the original audio data respectively with multiple kinds of noise data to obtain the sample audio data corresponding to the multiple kinds of noise data respectively can increase the amount of data for training the audio evaluation model; therefore, using the amplified large amount of sample audio data to train and predict the model can improve the accuracy of model prediction, and further improve the accuracy of audio data scoring.

[0089] Optionally, after training the audio evaluation model, the audio evaluation model can be applied to a specific application scenario, please refer to Figure 5 , Figure 5 which is a schematic flowchart of another data processing method provided by the embodiments of the present application. This method can be applied to a computer device; as Figure 5 shown, this method includes:

[0090] S301. Obtain the target audio data generated by the target user for the target service.

[0091] In the embodiments of the present application, the audio evaluation model can be applied to various scenarios that require voice testing, including but not limited to the Putonghua proficiency test, the spoken English test, or the question-and-answer system, etc. Correspondingly, the target service can refer to the Putonghua proficiency test service, the spoken English test service, or the service corresponding to the question-and-answer system, etc. The computer device obtains the target audio data generated by the target user for the target service. The target user can refer to the user who needs to perform the target service, and the target audio data refers to the audio data when the target user handles the target service. For example, in the spoken English test, the target audio data is the audio data obtained by recording the voice of the target user's response to the questions of the computer device. Optionally, the computer device can obtain the reference text data corresponding to the target audio data.

[0092] S302. Extract features from the target audio data to obtain the target audio features corresponding to the target audio data.

[0093] In the embodiments of the present application, the specific implementation manner of step S302 can refer to Figure 3 the description of extracting features from the sample audio data in step S101 in the corresponding embodiment to obtain the sample data features corresponding to the sample audio data, which will not be elaborated here. Optionally, the computer device can also extract features from the reference text data corresponding to the target audio data to obtain the reference audio features corresponding to the target audio data. The specific feature extraction method can refer to the description of extracting features from the sample audio data to obtain the sample data features corresponding to the sample audio data, which will not be elaborated here. Optionally, the computer device can also input the target audio data into the audio evaluation model, and based on the audio evaluation model, extract features from the target audio data to obtain the target audio features corresponding to the target audio data.

[0094] S303. Input the target audio features into the audio evaluation model, and based on the audio evaluation model, predict the target audio features to obtain the target score corresponding to the target audio features.

[0095] In the embodiments of the present application, the computer device can input the target audio features and the reference audio features corresponding to the target audio data into the audio evaluation model, and based on the audio evaluation model, predict the target audio features and the reference audio features corresponding to the target audio data to obtain the target score corresponding to the target audio features. Specifically, the computer device can determine the audio similarity between the target audio features and the reference audio features corresponding to the target audio data based on the audio evaluation model, and obtain the target score corresponding to the target audio features according to the audio similarity.

[0096] S304. If the target score is greater than or equal to the service passing threshold, send a service processing success message to the target user.

[0097] In the embodiments of the present application, the service qualification threshold can be determined according to specific circumstances. For example, in some oral English tests, if the oral score is less than 81, it is considered a failure, then the service qualification threshold can be 81. Optionally, the service qualification threshold can also be represented by a grade. Then, the computer device can also determine the grade corresponding to the target score according to the target score. When the grade corresponding to the target score is greater than the grade corresponding to the service qualification threshold, it indicates that the service processing of the target user is successful. For example, in the scenario of an exam, it means that the target user's score is passing, and a service processing success message is sent to the target user. Optionally, the computer device can also annotate the text data corresponding to the low-score audio data in the target audio data to obtain score analysis data, or annotate the text data corresponding to the high-score audio data to obtain score analysis data, and send the annotated score analysis data to the target user, so that the target user can improve the target audio data based on the score analysis data.

[0098] S305, if the target score is less than the service qualification threshold, send a service processing failure message to the target user.

[0099] In the embodiments of the present application, if the target score is less than the service qualification threshold, it indicates that the service processing of the target user fails. For example, in the scenario of an exam, it means that the target user's score is failing, and a service processing failure message is sent to the target user. The service processing failure message is used to instruct the target user to regenerate the audio data for the target service within the target time range. For example, the service processing failure message can include "Your score is failing. Please take a make-up exam within the XX time period."

[0100] In the embodiments of the present application, by predicting the target audio data of the target user and sending the predicted target score to the target user for viewing, it is convenient for the target user to understand their own deficiencies and facilitate subsequent improvement of the target audio data.

[0101] Optionally, during the process of using the audio evaluation model for score prediction, the audio evaluation model can also be adjusted according to user feedback to improve the accuracy of the audio evaluation model. Please refer to Figure 6 , Figure 6 which is a schematic flowchart of a process for adjusting an audio evaluation model provided by an embodiment of the present application. This method can be applied to a computer device; as Figure 6 shown, this method includes:

[0102] S401, obtain the original audio data of the target user, and perform noise addition processing on the original audio data respectively with N kinds of noise data to obtain N kinds of sample audio data corresponding to the N kinds of noise data.

[0103] S402. Extract features from each of the N sample audio data to obtain the sample data features corresponding to each of the N sample audio data.

[0104] S403. Obtain the score label corresponding to the original audio data. Train the initial evaluation model based on the sample data features and the score label corresponding to each of the N sample audio data to generate an audio evaluation model. Predict the target audio data of the target user based on the audio evaluation model to obtain a target score and send it to the target user.

[0105] In the embodiments of the present application, the specific implementation manners of steps S401 to S403 may refer to Figure 3 the descriptions in steps S101 to S103 in the corresponding embodiments, and refer to Figure 5 the description in step S303 in the corresponding embodiment of predicting the target audio data of the target user based on the audio evaluation model to obtain the target score, which will not be elaborated here.

[0106] S404. Receive the appeal request from the target user for the target score.

[0107] In the embodiments of the present application, the target user may, based on the target score received by the user terminal from the computer device, send an appeal request for the target score if the target user deems the target score abnormal. For example, if the target score is much lower than the target user's usual mock score.

[0108] S405. Send the target score to the evaluation terminal for review based on the appeal request.

[0109] In the embodiments of the present application, when the computer device receives the appeal request sent by the target user, it may send the target score to the evaluation terminal for review based on the appeal request. The evaluation terminal may refer to the terminal corresponding to an authoritative institution. The features of the target audio data of the target user may be extracted through the evaluation terminal to conduct a review and obtain a review result; alternatively, an expert may obtain the target audio data of the target user through the evaluation terminal to conduct a review and obtain a review result. The review result may be used to indicate that the target score is incorrect or the target score is correct. For example, when reviewing the target audio data, if the difference between the obtained score and the target score is less than or equal to the score validity threshold, the review result indicates that the target score is correct; if the difference between the obtained score and the target score is greater than the score validity threshold when reviewing the target audio data, the review result indicates that the target score is incorrect.

[0110] S406. Receive the review result sent by the evaluation terminal and adjust the audio evaluation model based on the review result.

[0111] In the embodiments of the present application, the computer device receives the review result sent by the evaluation terminal and adjusts the audio evaluation model based on the review result. If the review result indicates that the target score is correct, it means that the accuracy of the audio evaluation model is relatively high; if the review result indicates that the target score is incorrect, it means that the accuracy of the audio evaluation model is relatively low, and then the audio evaluation model is adjusted to improve the accuracy of model prediction.

[0112] In the embodiments of the present application, during the process of using the audio evaluation model to predict scores, if a complaint request for a user is received, the target score is reviewed based on the complaint request, and the audio evaluation model is adjusted based on the review result, so as to optimize the audio evaluation model and improve the accuracy of model prediction.

[0113] Optionally, referring to Figure 7 , Figure 7 FIG. is a schematic flowchart of a process for determining the data volume provided by the embodiments of the present application. This method can be applied to a computer device; as Figure 7 shown, the method includes:

[0114] S501, obtaining the original audio data of the target user.

[0115] S502, adding noise to the original audio data to obtain sample audio data.

[0116] S503, extracting features from the sample audio data to obtain the sample speech features and sample text features corresponding to the sample audio data.

[0117] S504, performing feature splicing on the sample speech features and sample text features to obtain the sample data features corresponding to the sample audio data, and increasing the number of sample audio data by 1.

[0118] S505, determining whether the number of sample audio data is equal to the target number threshold.

[0119] If yes, execute step S506; if no, execute step S502. The target number threshold can be determined according to specific circumstances. For example, the target number threshold can be a multiple of the number of original audio data, such as 5 times, 10 times, or 20 times the number of original audio data, etc.

[0120] S506, ending the step of extracting features from the sample audio data, and training the initial evaluation model based on the sample audio data to generate an audio evaluation model.

[0121] In the embodiments of the present application, by adding noise to the original audio data, the data volume for model training can be increased.

[0122] Optionally, three types of features corresponding to the audio data are provided in the embodiments of the present application, including directly extracting the features of the original audio data; the features enhanced by the prior art, that is, after extracting the features of the original audio data, adding noise to the extracted features to obtain enhanced features; and the features enhanced by the present application, that is, the enhanced features obtained by using the solutions in the embodiments of the present application. The three types of features are as follows:

[0123] Features of the original audio:

[0124] (01)1.00,0.00,2.33,1.67,3.00,0.00,3.00,0.00,0.00,1.00,

[0125] (02)1.00,0.00,2.00,2.00,0.00,0.00,0.00,3.00,0.00,1.00,

[0126] (03)2.00,0.00,0.00,0.00,2.00,2.00,2.00,1.00,2.00,2.67,

[0127] (04)0.00,0.00,2.00,0.00,2.00,2.00,0.00,3.00,2.00,0.06,

[0128] (05)0.00,0.09,0.09,0.03,0.00,0.03,0.00,0.00,0.03,0.03,

[0129] (06)0.00,0.03,0.03,0.00,0.00,0.00,0.09,0.00,0.03,0.03,

[0130] (07)0.00,0.00,0.00,0.03,0.03,0.03,0.03,0.03,0.09,0.00,

[0131] (08)0.00,0.06,0.00,0.03,0.03,0.00,0.03,0.03,0.00,0.40,

[0132] (09)0.07,0.00,1.00,0.00,1.00,0.68,1.00,0.82,0.81,0.00,

[0133] (10)0.00,0.00,0.00,0.00,1.00,0.00,1.00,1.00,0.00,0.00,

[0134] (11)0.00,0.00,1.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,

[0135] (12)0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,

[0136] (13)0.00,0.00,1.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,

[0137] (14)0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,

[0138] (15)0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,

[0139] (16)0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,

[0140] (17)0.00,0.00,4.00,0.00,1.00,0.00,0.00,0.00,0.00,0.00,

[0141] (18)0.00,0.00,0.00,0.00,0.00,0.00,0.00,4.00,0.00,0.00,

[0142] (19)0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,

[0143] (20)0.00,1.00,1.00,2.00,4.00,1.00,1.00,1.00,4.00,0.00,

[0144] (21)1.00,0.00,4.00,0.00,9.00,0.00,4.00,0.00,9.00,0.00,

[0145] (22)0.00,9.00,0.00,0.00,0.00,1.00,0.00,0.00,0.00,0.00,

[0146] (23)4.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,

[0147] (24)0.00,

[0148] Features enhanced from the prior art:

[0149] (01)0.92,0.13,2.28,1.53,3.12,-0.05,3.06,-0.10,-0.11,1.02,

[0150] ……

[0151] (09)0.16,0.08,0.95,0.03,1.04,0.67,1.13,0.68,0.95,0.04,

[0152] (10)-0.09,0.03,0.10,-0.12,0.99,-0.10,0.83,0.94,0.00,-0.06,

[0153] ……

[0154] Features enhanced in this application:

[0155] (01)1.00,0.00,2.33,2.00,3.00,0.00,3.00,0.00,0.00,1.00,

[0156] ……

[0157] (09)0.07,0.00,1.00,0.00,1.00,0.71,1.00,0.81,0.81,0.00,

[0158] (10)0.00,0.00,0.00,0.00,1.00,0.00,2.00,1.00,0.00,0.00,

[0159] ……

[0160] It can be seen that the solution in the prior art will add noise to each element in the original audio features and does not consider the internal correlation between each element in the original audio features. In contrast, the technical solution in this application only changes some of the speech features in (1) and the speech and text features related to the original audio features in (9) and (10), taking into account the internal correlation between each element in the original audio features. Therefore, using the data amplified in this way for model training can improve the accuracy of the model.

[0161] Optionally, in the application scenario of oral English tests, the evaluation indicators of oral English test questions usually include the consistency rate and acceptability, etc. The consistency rate refers to the proportion of the machine scoring and manual scoring of the intelligent marking system being exactly the same, and the acceptability refers to the proportion that the difference between the machine scoring and manual scoring is within one scoring range. Correspondingly, in the embodiments of the present application, a comparison table of the effects of model training using the original audio data, the audio data enhanced by the method of the prior art, and the audio data enhanced by the present application is provided, as shown in Table 1:

[0162] Table 1 Comparison Table of Effects

[0163]

[0164] As can be seen from Table 1, the consistency rate corresponding to the audio data enhanced by the present application has increased by 1 point, exceeding 90%; the acceptability has increased by 0.88 points, exceeding 96%. Since the effect of the original audio data is already good, achieving an improvement of 1 point in the present application is very prominent, and the corresponding score prediction accuracy has also increased accordingly.

[0165] The method of the embodiments of the present application has been introduced above. Next, the device of the embodiments of the present application will be introduced.

[0166] See Figure 8 , Figure 8 is a schematic structural diagram of the composition of a data processing device provided by the embodiments of the present application. The above-mentioned video data processing device may be a computer program (including program code) running in a computer device. For example, this data processing device is an application software; this device can be used to execute the corresponding steps in the method provided by the embodiments of the present application. The device 80 includes:

[0167] An original data acquisition module 81, configured to acquire the original audio data of the target user, and perform noise addition processing on the original audio data using N kinds of noise data respectively to obtain the sample audio data corresponding to the N kinds of noise data respectively; N is a positive integer;

[0168] A feature extraction module 82, configured to perform feature extraction on the N sample audio data respectively to obtain the sample data features corresponding to the N sample audio data respectively;

[0169] A model generation module 83, configured to obtain the score label corresponding to the original audio data, and train an initial evaluation model based on the sample data features corresponding to the N sample audio data respectively and the score label to generate an audio evaluation model; this audio evaluation model is used to predict the target score corresponding to the target audio data.

[0170] Optionally, the N sample audio data includes sample audio data i; i is a positive integer; the feature extraction module 82 includes:

[0171] A voice extraction unit 821, configured to extract voice features from the sample audio data i to obtain sample voice features corresponding to the sample audio data i;

[0172] A text extraction unit 822, configured to perform voice conversion processing on the sample audio data i to obtain sample text data corresponding to the sample audio data i, extract text features from the sample text data, and obtain sample text features corresponding to the sample audio data i;

[0173] A feature splicing unit 823, configured to splice the sample voice features and the sample text features to generate sample data features corresponding to the sample audio data i.

[0174] Optionally, the voice extraction unit 821 is specifically configured to:

[0175] Obtain the voice fluency corresponding to the sample audio data i, and determine the first voice feature based on the voice fluency;

[0176] Obtain the phoneme sequence corresponding to the sample audio data i, and determine the second voice feature based on the phoneme sequence corresponding to the sample audio data i;

[0177] Obtain the pronunciation accuracy corresponding to the sample audio data i, and determine the third voice feature based on the pronunciation accuracy;

[0178] Based on the first voice feature, the second voice feature, and the third voice feature, determine the sample voice features corresponding to the sample audio data i.

[0179] Optionally, the N sample audio data all include the audio data to be evaluated and reference audio data, and the sample data features respectively corresponding to the N sample audio data include the data features to be evaluated corresponding to the audio data to be evaluated and the reference audio features corresponding to the reference audio data; the model generation module 83 includes:

[0180] A similarity determination unit 831, configured to input the sample data features respectively corresponding to the N sample audio data into the initial evaluation model, determine the audio similarity between the data features to be evaluated and the reference audio features corresponding to each sample audio data based on the initial evaluation model, and obtain sample prediction scores according to the audio similarity;

[0181] A model adjustment unit 832, configured to adjust the initial evaluation model based on the difference value between the score label and the sample prediction scores to generate the audio evaluation model.

[0182] Optionally, the device 80 further includes:

[0183] A sample division module 84, configured to divide the N sample audio data into training sample data and verification sample data;

[0184] The model generation module 83 includes:

[0185] A model training unit 833, configured to train the initial evaluation model based on the training sample data to generate a to-be-detected evaluation model;

[0186] A quality determination unit 834, configured to detect the to-be-detected evaluation model based on the verification sample data to obtain the model quality corresponding to the to-be-detected evaluation model;

[0187] A model determination unit 835, configured to, if the model quality is greater than or equal to a model effective threshold, determine the to-be-detected evaluation model as the audio evaluation model.

[0188] Optionally, the device 80 further includes a model adjustment module 85, including:

[0189] A data acquisition unit 851, configured to acquire target audio data generated by the target user for a target service;

[0190] A feature extraction unit 852, configured to extract features from the target audio data to obtain target audio features corresponding to the target audio data;

[0191] A score determination unit 853, configured to input the target audio features into the audio evaluation model, and predict the target audio features based on the audio evaluation model to obtain a target score corresponding to the target audio features;

[0192] Specifically, if the target score is greater than or equal to a service qualification threshold, the score determination unit 853 is configured to send a service processing success message to the target user;

[0193] Specifically, if the target score is less than the service qualification threshold, the score determination unit 853 is configured to send a service processing failure message to the target user, where the service processing failure message is used to instruct the target user to regenerate audio data for the target service within a target time range.

[0194] Optionally, the device 80 further includes a model optimization module 86, including:

[0195] A request acquisition unit 861, configured to receive an appeal request of the target user for the target score;

[0196] A request review unit 862, configured to send the target score to an evaluation terminal for review based on the appeal request;

[0197] The model optimization unit 863 is configured to receive the review result sent by the evaluation terminal and adjust the audio evaluation model based on the review result.

[0198] It should be noted that Figure 8 For the content not mentioned in the corresponding embodiments, reference can be made to the description of the method embodiments, which will not be elaborated here.

[0199] In the embodiments of the present application, by obtaining the original audio data of the target user, adding noise to the original audio data respectively using N kinds of noise data to obtain the sample audio data corresponding to the N kinds of noise data respectively; extracting features from the N sample audio data respectively to obtain the sample data features corresponding to the N sample audio data respectively; obtaining the score label corresponding to the original audio data, and training the initial evaluation model based on the sample data features and the score label corresponding to the N sample audio data respectively to generate an audio evaluation model; the audio evaluation model is used to predict the target score corresponding to the target audio data. Since adding noise to the original audio data respectively using multiple kinds of noise data to obtain the sample audio data corresponding to multiple kinds of noise data respectively can increase the amount of data for training the audio evaluation model; therefore, using the amplified large number of sample audio data to train and predict the model can improve the accuracy of model prediction, and further improve the accuracy of audio data scoring.

[0200] See Figure 9 , Figure 9 is a schematic structural diagram of a computer device provided by the embodiments of the present application. As Figure 9 shown, the above computer device 90 may include: a processor 901, a network interface 904, and a memory 905. In addition, the above computer device 90 may further include: a user interface 903, and at least one communication bus 902. Among them, the communication bus 902 is used to realize the connection and communication between these components. Among them, the user interface 903 may include a display screen (Display), a keyboard (Keyboard), and optionally the user interface 903 may further include a standard wired interface and a wireless interface. The network interface 904 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 905 may be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk memory. The memory 905 may optionally be at least one storage device located far from the aforementioned processor 901. As Figure 9 shown, the memory 905, as a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.

[0201] In Figure 9In the computer device 90 shown, the network interface 904 can provide network communication functions; the user interface 903 is mainly used to provide an interface for users to input; and the processor 901 can be used to call the device control application program stored in the memory 905 to achieve:

[0202] Obtain the original audio data of the target user, and perform noise addition processing on the original audio data using N kinds of noise data respectively to obtain the sample audio data corresponding to the N kinds of noise data respectively; N is a positive integer;

[0203] Extract features from the N sample audio data respectively to obtain the sample data features corresponding to the N sample audio data respectively;

[0204] Obtain the score label corresponding to the original audio data, and train the initial evaluation model based on the sample data features corresponding to the N sample audio data respectively and the score label to generate an audio evaluation model; the audio evaluation model is used to predict the target score corresponding to the target audio data.

[0205] In some possible implementation manners, the processor 901 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0206] The memory 905 may include a read-only memory and a random access memory, and provide instructions and data to the processor 901 and the network interface 904. A part of the memory 905 may also include a non-volatile random access memory. For example, the memory 905 may also store information about the device type.

[0207] In specific implementation, the computer device can execute the implementation manners provided in each step of the foregoing method embodiment through its built-in various functional modules. For details, refer to the implementation manners provided in each step of the foregoing method embodiment, which will not be elaborated here.

[0208] Embodiments of the present application provide a computer device, including: a processor, a memory, and a network interface. The processor obtains computer instructions in the memory and executes each step of the information processing method to perform information processing operations. In the embodiments of the present application, since multiple pieces of noise data are respectively used to perform noise addition processing on the original audio data to obtain sample audio data corresponding to the multiple pieces of noise data respectively, the data volume for training the audio evaluation model can be amplified. Therefore, using the amplified large amount of sample audio data to train and predict the model can improve the accuracy of model prediction, and further improve the accuracy of audio data scoring.

[0209] Embodiments of the present application further provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by the processor, the information processing method provided by each step in the foregoing method embodiments can be implemented. Specifically, reference can be made to the implementation manners provided by each step in the foregoing method embodiments, which will not be elaborated herein. In addition, the description of the beneficial effects of using the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiments of the present application. As an example, the program instructions can be deployed to be executed on a computer device, or on multiple computer devices located at one place. Or, on multiple computer devices distributed at multiple places and interconnected through a communication network.

[0210] The computer-readable storage medium can be the information processing device provided in any of the foregoing embodiments or an internal storage unit of the computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.

[0211] The embodiments of the present application also provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the various alternative ways in the foregoing method embodiments. As a result, by performing noise addition processing on the original audio data using multiple pieces of noise data respectively, sample audio data corresponding to the multiple pieces of noise data can be obtained, and the amount of data for training the audio evaluation model can be amplified. Therefore, using the amplified large amount of sample audio data to train and predict the model can improve the accuracy of model prediction, and further improve the accuracy of audio data scoring.

[0212] In the description of the embodiments of the present application, the terms "first", "second", etc. in the specification, claims and drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the listed steps or modules, but optionally further includes steps or modules not listed, or optionally further includes other step units inherent to these processes, methods, devices, products or equipment.

[0213] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in this description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0214] The methods and related devices provided in the embodiments of the present application are described with reference to the method flowcharts and / or structure diagrams provided in the embodiments of the present application. Specifically, each process and / or block of the method flowchart and / or structure diagram, and the combination of the processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 a process or multiple processes and / or structure schematic Figure 1means for the functions specified in one or more boxes. These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means, and the instruction means implements in the process Figure 1 one process or more processes and / or structural schematic Figure 1 means for the functions specified in one or more boxes. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide for implementing in the process Figure 1 steps for the functions specified in one process or more processes and / or structural schematic one or more boxes.

[0215] The above disclosure is only for the preferred embodiments of this application. Of course, the scope of rights of this application cannot be limited by this. Therefore, equivalent changes made according to the claims of this application still fall within the scope covered by this application.

Claims

1. A data processing method, characterized in that, it includes: Obtain the original audio data of the target user, convert N kinds of noise data into N kinds of noise signals with the same signal intensity as the original audio data, and use the N kinds of noise signals to perform noise addition processing on the original audio data respectively to obtain N sample audio data; N is a positive integer; Perform feature extraction on the N sample audio data respectively to obtain the sample data features corresponding to the N sample audio data; wherein, each sample audio data includes the audio data to be evaluated and the reference audio data; the sample data feature corresponding to each sample audio data includes the data feature to be evaluated corresponding to the audio data to be evaluated and the reference audio feature corresponding to the reference audio data; Obtain the score label corresponding to the original audio data, and train the initial evaluation model based on the sample data features corresponding to the N sample audio data and the score label to generate an audio evaluation model; the audio evaluation model is used to predict the target score corresponding to the target audio data; Annotate the text data corresponding to the low-score audio data or the text data corresponding to the high-score audio data in the target audio data to obtain score analysis data; Send the score analysis data to the target user, and the score analysis data is used for the target user to improve the target audio data; Among them, the training of the initial evaluation model based on the sample data features corresponding to the N sample audio data and the score label to generate an audio evaluation model includes: Input the sample data features corresponding to the N sample audio data into the initial evaluation model respectively, determine the audio similarity between the data feature to be evaluated and the reference audio feature corresponding to each sample audio data based on the initial evaluation model, and obtain the sample prediction score according to the audio similarity; Adjust the initial evaluation model based on the difference value between the score label and the sample prediction score to generate an audio evaluation model.

2. The method according to claim 1, characterized in that, the N sample audio data includes sample audio data i; i is a positive integer; The performing feature extraction on the N sample audio data respectively to obtain the sample data features corresponding to the N sample audio data includes: Perform speech feature extraction on the sample audio data i to obtain the sample speech feature corresponding to the sample audio data i; Perform speech conversion processing on the sample audio data i to obtain the sample text data corresponding to the sample audio data i, and perform text feature extraction on the sample text data to obtain the sample text feature corresponding to the sample audio data i; Perform feature splicing on the sample speech feature and the sample text feature to generate the sample data feature corresponding to the sample audio data i.

3. The method according to claim 2, characterized in that, The performing speech feature extraction on the sample audio data i to obtain the sample speech feature corresponding to the sample audio data i includes: Obtain the speech fluency corresponding to the sample audio data i, and determine the first speech feature based on the speech fluency; Obtain the phoneme sequence corresponding to the sample audio data i, and determine the second speech feature based on the phoneme sequence corresponding to the sample audio data i; Obtain the pronunciation accuracy corresponding to the sample audio data i, and determine the third speech feature based on the pronunciation accuracy; Based on the first speech feature, the second speech feature, and the third speech feature, determine the sample speech feature corresponding to the sample audio data i.

4. The method according to claim 1, wherein, the method further includes: Dividing the N sample audio data into training sample data and verification sample data; The training of the initial evaluation model based on the sample data features and the score labels respectively corresponding to the N sample audio data to generate an audio evaluation model includes: Training the initial evaluation model based on the training sample data to generate a to-be-detected evaluation model; Detecting the to-be-detected evaluation model based on the verification sample data to obtain the model quality corresponding to the to-be-detected evaluation model; If the model quality is greater than or equal to the model effective threshold, determine the to-be-detected evaluation model as the audio evaluation model.

5. The method according to claim 1, wherein, the method further includes: Obtain the target audio data generated by the target user for the target service; Extract features from the target audio data to obtain the target audio feature corresponding to the target audio data; Input the target audio feature into the audio evaluation model, and predict the target audio feature based on the audio evaluation model to obtain the target score corresponding to the target audio feature; If the target score is greater than or equal to the service qualification threshold, send a service processing success message to the target user; If the target score is less than the service qualification threshold, send a service processing failure message to the target user, and the service processing failure message is used to instruct the target user to regenerate the audio data for the target service within the target time range.

6. The method according to claim 1, wherein, the method further includes: Receive the appeal request of the target user for the target score; Send the target score to the evaluation terminal for review based on the appeal request; Receive the review result sent by the evaluation terminal, and adjust the audio evaluation model based on the review result.

7. A data processing device, wherein, it includes: An original data acquisition module, configured to acquire the original audio data of the target user, convert N kinds of noise data into N kinds of noise signals with the same signal intensity as the original audio data, and use the N kinds of noise signals to perform noise addition processing on the original audio data respectively to obtain N sample audio data; N is a positive integer; A feature extraction module for separately extracting features from N sample audio data to obtain sample data features corresponding to the N sample audio data; wherein each sample audio data includes audio data to be evaluated and reference audio data; the sample data features corresponding to each sample audio data include data features to be evaluated corresponding to the audio data to be evaluated and reference audio features corresponding to the reference audio data; A model generation module for obtaining a score label corresponding to the original audio data, training an initial evaluation model based on the sample data features corresponding to the N sample audio data and the score label, and generating an audio evaluation model; the audio evaluation model is used to predict a target score corresponding to target audio data; A model adjustment module for annotating text data corresponding to low-score audio data or text data corresponding to high-score audio data in the target audio data to obtain score analysis data; sending the score analysis data to a target user, and the score analysis data is used for the target user to improve the target audio data; Wherein, the model generation module includes: A similarity determination unit for inputting the sample data features corresponding to the N sample audio data into the initial evaluation model, determining the audio similarity between the data features to be evaluated corresponding to each sample audio data and the reference audio features based on the initial evaluation model, and obtaining a sample prediction score according to the audio similarity; A model adjustment unit for adjusting the initial evaluation model based on the difference value between the score label and the sample prediction score to generate an audio evaluation model.

8. A computer device, characterized in that, it includes: A processor, a memory, and a network interface; The processor is connected to the memory and the network interface, wherein the network interface is used to provide a data communication function, the memory is used to store program codes, and the processor is used to call the program codes so that the computer device executes the method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded and executed by a processor so that a computer device with the processor executes the method according to any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes computer instructions, and when the computer instructions are executed by a processor, they are used to implement the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Calibration optimization method and system for speaking test evaluation

    CN104464423A

  • Audio classification method based on dual data enhancement strategy

    CN110808033A