Audio data processing methods, devices, media and equipment

By using automated processing methods to evaluate user pronunciation, the problems of inaccurate pronunciation evaluation and high labor costs in existing technologies are solved, achieving objective and comprehensive pronunciation feedback and saving labor costs.

CN114203199BActive Publication Date: 2025-10-31BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111512478.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-08
Publication Date
2025-10-31
Estimated Expiration
2041-12-08

AI Technical Summary

Technical Problem

In existing technologies, user pronunciation evaluation is limited by human factors, resulting in inaccurate evaluation results that cannot meet the needs of a large number of users, and also incurring high labor costs.

Method used

The system uses automated processing methods to acquire the audio to be evaluated and the text to be read aloud, extracts the speech content, scores it, and generates feedback text, enabling positive, negative, or neutral evaluations and generating feedback speech.

Benefits of technology

It achieves objective and comprehensive pronunciation evaluation, reduces the influence of human factors, saves labor costs, and allows users to intuitively understand pronunciation problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114203199B_ABST
    Figure CN114203199B_ABST
Patent Text Reader

Abstract

This disclosure relates to an audio data processing method, apparatus, medium, and device, comprising: acquiring an audio to be evaluated and a text to be read aloud; extracting the speech to be evaluated from the audio; scoring one or more segments of speech content corresponding to each word in the text to be evaluated, to obtain one or more scoring results; determining the text reading evaluation corresponding to the text to be read aloud based on the one or more scoring results, including positive evaluation, negative evaluation, and neutral evaluation; determining an evaluation template based on the text reading evaluation corresponding to the text to be read aloud and generating feedback text; and synthesizing the feedback text into feedback speech and sending it to the user corresponding to the audio to be evaluated. In this way, all of the user's pronunciations can be scored and comprehensively evaluated, resulting in a more objective, reasonable, and comprehensive text reading evaluation. Furthermore, the feedback evaluation speech can be automatically generated based on the evaluation, allowing users to more intuitively understand their pronunciation problems and reducing the influence of subjective human factors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of speech processing technology, and more specifically, to an audio data processing method, apparatus, medium, and device. Background Technology

[0002] In language learning or voice dubbing training, targeted evaluation and correction of a user's pronunciation can accelerate their learning progress. However, relying on live teachers for this service is limited by labor costs, making it difficult to meet the needs of a large number of users. Furthermore, the results of human evaluation and correction are also limited by the professional level of the teachers. For example, due to the different professional backgrounds of different teachers, the quality of student feedback will vary depending on the teacher's expertise. Each teacher's work status is different, and the service quality of each teacher is significantly affected by human factors when facing a large number of student evaluation requests. Current technologies for evaluating user pronunciation typically only provide a scoring function, and users' perception of the score can vary due to human factors, making it impossible to guarantee that users can accurately understand their own pronunciation status based on the score. Summary of the Invention

[0003] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] In a first aspect, this disclosure provides an audio data processing method, the method comprising:

[0005] Obtain the audio to be evaluated and the text to be read aloud;

[0006] Extract the speech to be evaluated from the audio to be evaluated;

[0007] Each segment or multiple segments of speech content corresponding to each word in the text to be read aloud are scored to obtain one or more scoring results.

[0008] The text reading evaluation corresponding to the text is determined by one or more scoring results, and the text reading evaluation includes positive evaluation, negative evaluation and neutral evaluation;

[0009] Based on the text reading evaluation corresponding to the text reading text, an evaluation template is determined and feedback text is generated;

[0010] The feedback text is synthesized into feedback speech and then sent to the user corresponding to the audio to be evaluated.

[0011] Secondly, this disclosure provides an audio data processing apparatus, the apparatus comprising:

[0012] The first acquisition module is used to acquire the audio to be evaluated and the text to be read aloud.

[0013] The extraction module is used to extract the speech to be evaluated from the audio to be evaluated;

[0014] The scoring module is used to score one or more segments of speech content that correspond to each word in the text to be read aloud in the speech to be evaluated, so as to obtain one or more scoring results.

[0015] The evaluation module is used to determine the text reading evaluation corresponding to the text reading based on the one or more scoring results. The text reading evaluation includes positive evaluation, negative evaluation, and neutral evaluation.

[0016] The feedback module is used to determine the evaluation template and generate feedback text based on the text reading evaluation corresponding to the reading text.

[0017] The synthesis module is used to synthesize the feedback text into feedback speech and send it to the user corresponding to the audio to be evaluated.

[0018] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described above.

[0019] Fourthly, this disclosure provides an electronic device, comprising:

[0020] A storage device on which computer programs are stored;

[0021] A processing device for executing the computer program in the storage device to implement the steps of the method described above.

[0022] The above technical solution processes the audio of the user's pronunciation practiced on the text, directly obtaining the user's pronunciation evaluation without human judgment. It also scores multiple pronunciations of the same word to determine the final evaluation. This comprehensive evaluation, based on scoring all of the user's pronunciations, results in a more objective, reasonable, and complete text practice evaluation. Furthermore, the pronunciation evaluation can be automatically classified as positive, negative, or neutral, and feedback audio can be automatically generated based on this evaluation. This allows users to more intuitively understand their pronunciation problems. This approach provides a highly objective evaluation of the user's pronunciation without any subjective human intervention, while also saving on labor costs.

[0023] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description

[0024] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:

[0025] Figure 1 This is a flowchart illustrating an audio data processing method according to an exemplary embodiment of the present disclosure.

[0026] Figure 2 This is a flowchart illustrating another exemplary audio data processing method according to the present disclosure.

[0027] Figure 3 This is a flowchart illustrating another exemplary audio data processing method according to the present disclosure.

[0028] Figure 4 This is a flowchart illustrating another exemplary audio data processing method according to the present disclosure.

[0029] Figure 5 This is a structural block diagram of an audio data processing apparatus according to an exemplary embodiment of the present disclosure.

[0030] Figure 6 This is a structural block diagram of an audio data processing apparatus according to yet another exemplary embodiment of the present disclosure.

[0031] Figure 7 A schematic diagram of the structure of an electronic device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation

[0032] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0033] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0034] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0035] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0036] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0037] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0038] Figure 1 This is a flowchart illustrating an audio data processing method according to an exemplary embodiment of the present disclosure. Figure 1 As shown, the method includes steps 101 to 106.

[0039] In step 101, the audio to be evaluated and the text to be read aloud are obtained.

[0040] The audio to be evaluated can be audio data recorded by any recording device. This disclosure does not limit the content of the audio. For example, it could be audio recorded during language teaching when a user reads a text aloud, or audio recorded during a user's dubbing learning process when a user reads a dubbing text aloud, as long as the audio to be evaluated is audio that the user reads aloud based on an existing text and requires evaluation feedback. The text to be evaluated is the text that can be manually confirmed by the user or automatically confirmed by an application when the audio to be evaluated is recorded.

[0041] In step 102, the speech to be evaluated is extracted from the audio to be evaluated.

[0042] Step 102, which involves extracting the speech to be evaluated from the audio to be evaluated, can be considered as denoising the audio to be evaluated, retaining only the speech of the target user. The noise to be removed can include noise caused by the user's recording environment, as well as speech from other people who are not associated with the text being read aloud. After identifying the speaker and other noise, the speech to be evaluated corresponding to the user is extracted from the audio to be evaluated through audio segmentation. Where the noise to be removed includes the speech of other people, the method for extracting the speech to be evaluated from the audio to be evaluated can be: extracting the speech to be evaluated from the audio to be evaluated based on the user's user information, which includes any one of the user's age, gender, and voice timbre. The user's age and gender can be obtained by the user during registration and authorization, while the user's voice timbre can be obtained by the user actively performing voice timbre recognition and authorizing the use of their voice timbre for extracting the speech to be evaluated.

[0043] In step 103, one or more segments of speech content corresponding to each word in the text to be read are scored to obtain one or more scoring results.

[0044] After extracting the speech to be evaluated from the audio to be evaluated, one or more segments of speech content corresponding to each word in the text to be read aloud are scored separately to obtain one or more scoring results. That is, the speech to be evaluated may contain instances where a certain word in the text is read aloud multiple times. For example, if a reading aloud exercise requires reading word A in the text three times, then the text only contains one word A, but the speech to be evaluated includes the user's reading of word A three times; or, if word A appears multiple times in the same sentence, then the speech to be evaluated will also include multiple readings of the word A with different pronunciations. In this disclosure, the pronunciation of each word in the speech to be evaluated is scored, which means that if the user reads a certain word in the text aloud multiple times, that word will correspond to multiple scoring results.

[0045] In step 104, the text reading evaluation corresponding to the text is determined through one or more scoring results. The text reading evaluation includes positive evaluation, negative evaluation and neutral evaluation.

[0046] After obtaining all the scoring results for the speech to be evaluated, the overall evaluation of the text to be read aloud can be determined, that is, the text reading evaluation. The specific method of determination is not limited in this embodiment; it can be based on the average score of all scoring results, or it can be determined by sampling from the scoring results, etc. The positive, negative, and neutral evaluations included in the text reading evaluation can be used to represent that the user read well, the user's pronunciation needs correction, and no feedback can be provided, respectively. Specifically, positive and negative evaluations can be determined based on the scoring results of all words corresponding to the speech to be evaluated, while neutral evaluations can be due to various reasons. For example, the evaluation of the text reading might be neutral because the text was omitted and could not be scored; or the scoring results might indicate that the text reading did not match the extracted speech to be evaluated, meaning there might be an error in the text reading, resulting in a neutral evaluation; or it might be due to audio recording or transmission issues, where the user's audio to be evaluated was not detected or the speech to be evaluated could not be extracted from the audio, resulting in a neutral evaluation.

[0047] In step 105, an evaluation template is determined based on the text reading evaluation corresponding to the reading text, and feedback text is generated.

[0048] After obtaining the text-based feedback for the given text, if the feedback is neutral, the reason for this neutral feedback can be further determined. Based on this reason, it can be determined whether an evaluation template needs to be created and feedback text generated. For example, if the reason is that the text was omitted, the user can be directly notified of the error, and processing of the audio to be evaluated can be stopped. If the reason is that the audio to be evaluated was not detected or the audio to be evaluated cannot be extracted, the corresponding explanatory text can be used as the feedback text, synthesized into speech, and fed back to the user without evaluating the user's pronunciation. Error reminders can be given to the user based on the reason. This explanatory text is pre-set.

[0049] If the text reading evaluation is positive or negative, then a corresponding evaluation template can be determined to generate the appropriate feedback text. This evaluation template is the text template used to evaluate the user. Positive and negative evaluations can each include multiple different templates, with different feedback text content in each template, but all corresponding to either positive or negative evaluation. When determining the evaluation template, one can be randomly selected from multiple templates corresponding to the text reading evaluation. For example, among evaluation templates corresponding to positive evaluations, one could include the feedback statement "Great reading, teacher gives you a thumbs up," while other templates could include the feedback statement "Very good, great progress."

[0050] In addition, the evaluation template may include text statements that need to be filled in by, for example, follow-along text or other information, as well as any combination of other types of statements used for user interaction or explanation. After the evaluation template is randomly determined based on the text follow-along evaluation, the required data can be obtained according to the template content included in the evaluation template to generate the feedback text.

[0051] In step 106, the feedback text is synthesized into feedback speech and then sent to the user corresponding to the audio to be evaluated.

[0052] When synthesizing the feedback text into feedback speech, it can also be done according to the speech template selected by the user. For example, the user can choose a female speech template or a male speech template, and the synthesized speech will be different according to the user's different choices.

[0053] The above technical solution processes the audio of the user's pronunciation practiced on the text, directly obtaining the user's pronunciation evaluation without human judgment. It also scores multiple pronunciations of the same word to determine the final evaluation. This comprehensive evaluation, based on scoring all of the user's pronunciations, results in a more objective, reasonable, and complete text practice evaluation. Furthermore, the pronunciation evaluation can be automatically classified as positive, negative, or neutral, and feedback audio can be automatically generated based on this evaluation. This allows users to more intuitively understand their pronunciation problems. This approach provides a highly objective evaluation of the user's pronunciation without any subjective human intervention, while also saving on labor costs.

[0054] Figure 2 This is a flowchart illustrating another exemplary audio data processing method according to this disclosure. For example... Figure 2 As shown, the method further includes steps 201 to 204.

[0055] In step 104 above, when determining the text reading evaluation corresponding to the reading text through one or more scoring results, it can be determined in the following way: Based on the scoring results of one or more audio segments corresponding to each word in the reading text, determine the word reading evaluation corresponding to each word in the reading text, where the word reading evaluation includes positive evaluation, negative evaluation, and neutral evaluation; determine the text reading evaluation corresponding to the reading text based on the word reading evaluation corresponding to each word in the reading text. That is, when determining the text reading evaluation, the word reading evaluation corresponding to each word in the reading text can be determined first, and then the text reading evaluation can be determined based on the word reading evaluation.

[0056] Specifically, the method for determining the evaluation of the word's pronunciation can be as shown in steps 201 and 202.

[0057] In step 201, the highest score among the scoring results of one or more audio segments corresponding to each word in the follow-up text is taken as the word score result corresponding to each word in the follow-up text.

[0058] In step 202, the word reading evaluation corresponding to each word in the reading text is determined according to the preset evaluation criteria and the word scoring results.

[0059] For example, in the aforementioned example, if word A appears multiple times in the same sentence, the speech to be evaluated will include multiple pronunciations of the same word A. After scoring the speech, the score result will also include multiple scores corresponding to word A. Finally, the highest score will be used as the score for word A. In this way, by scoring all of the user's speech pronunciations, it is possible to minimize the possibility that a user's accurate pronunciation of a word might be misidentified due to audio recording issues or the accuracy of the speaker's speech extraction, resulting in a negative evaluation and causing user confusion.

[0060] Furthermore, after determining the unique word reading evaluation corresponding to each word in the reading text, the specific method used to determine the text reading evaluation corresponding to the reading text can be as shown in step 203 or step 204.

[0061] In step 203, that is, when the word reading evaluations corresponding to each word in the reading text include either the positive evaluation or the negative evaluation, the word reading evaluation corresponding to the word with the longest text length and whose corresponding word reading evaluation is not the neutral evaluation is taken as the text reading evaluation corresponding to the reading text.

[0062] As described above, if any word in the text receives a positive or negative evaluation in its word reading evaluation, the text reading evaluation can be determined based on the word reading evaluation in the text.

[0063] In this embodiment, a method is adopted to use the word reading evaluation corresponding to the longest word in the reading text that is not a neutral evaluation as the text reading evaluation corresponding to the reading text. That is, the word reading evaluation of the longest word in the reading text is searched. If the word reading evaluation of the longest word is neither a positive nor a negative evaluation, the word reading evaluation of the second longest word is searched, and so on.

[0064] In another possible implementation, the method for determining the text reading evaluation of the text to be read aloud can be as follows: First, obtain the word reading evaluation of the central word included in the text to be read aloud. If the word reading evaluation of the central word is neither a positive nor a negative evaluation, then obtain the word reading evaluation corresponding to the longest word in the text to be read aloud and whose corresponding word reading evaluation is not the neutral evaluation. The central word can be pre-set and can be obtained at the same time as obtaining the text to be read aloud.

[0065] In step 204, that is, if the word reading evaluation corresponding to each word in the reading text does not include the positive evaluation or the negative evaluation, the text reading evaluation corresponding to the reading text is determined to be a neutral evaluation.

[0066] If the word-based reading evaluations in the text do not include either positive or negative evaluations, it means there is no text content in the entire text that can be used as evaluation feedback. In this case, the text-based reading evaluation for that text can be directly determined as a neutral evaluation. This allows for a more direct and faster way to obtain the text-based reading evaluation without further searching for the central word or the length of each word in the text.

[0067] Figure 3 This is a flowchart illustrating an audio data processing method according to yet another exemplary embodiment of this disclosure. For example... Figure 3 As shown, the method further includes steps 301 and 302.

[0068] In step 301, user data corresponding to the user and interpretation information and question type information corresponding to the text to be read aloud are obtained.

[0069] The aforementioned user data may include various user-authorized data such as user name, user gender, and user voice. The interpretation information corresponding to the aforementioned text to be read aloud can be Chinese when the text is in English, English when the text is in Chinese, and corresponding modern Chinese when the text is in classical Chinese. The specific interpretation information can be set according to the content of the text to be read aloud. The aforementioned question type information is information that can be determined when the user selects the text to be read aloud before the audio to be evaluated is recorded. The question type can be a regular sentence reading reminder or a phonics question, etc. Different question types can also correspond to different evaluation templates.

[0070] In step 302, the evaluation template is determined and the feedback text is generated based on the user data, the definition information and question type information corresponding to the reading text, and the text reading evaluation corresponding to the reading text.

[0071] When determining the evaluation template based on user data, the corresponding explanatory information and question type information of the text to be read aloud, and the text reading evaluation corresponding to the text to be read aloud, a hierarchical logical triggering method can be used. For example, firstly, the content of the text reading evaluation can be used to determine whether the feedback attitude to be triggered is positive or negative, thereby determining the first part of the evaluation template data corresponding to the text reading evaluation; then, it can be determined whether the explanatory information corresponding to the text to be read aloud is empty, thereby determining the second part of the evaluation template data corresponding to the explanatory information corresponding to the text to be read aloud. If the explanatory information is empty, the second part of the evaluation template data can include interactive templates, such as guiding the user to read aloud the text again (the explanatory information corresponding to phonics questions is usually empty). The specific number of templates and the sentences used in the templates are not limited in this disclosure. If the explanatory information is not empty, the second part of the evaluation template data can include text explanation templates, which are used to explain the explanatory information of the text to the user using different explanation methods. The user can then determine whether to add a third part of the evaluation template data based on the definition information. If the definition information is empty, the interactive template for guiding the user to reread the text has already been determined in the previous step, so the third part of the evaluation template data will not be added in this case. However, if the definition information is not empty, the third part of the evaluation template data can be added. Finally, the user can determine the fourth part of the evaluation template data based on whether the text reading evaluation corresponding to the text is positive or negative, and whether the final interactive content is a comforting inquiry or an encouraging message. For example, if the text reading evaluation corresponding to the text is negative, the template content included in the fourth part of the evaluation template data can be various templates such as "Don't be discouraged, we look forward to your even better performance in the future".

[0072] After obtaining the evaluation template data for the above-mentioned multiple parts step by step, one template can be randomly selected from each to form the final evaluation template. The evaluation template can be as follows:

[0073] (1) You need to practice "follow-along text" and "user name" more! (2) What does "follow-along text" mean? It means "interpretive information". (3) We need to pay attention to mouth shape and pronunciation. Let's listen to the correct pronunciation: "follow-along text", "follow-along text", "follow-along text". (4) Remember it? We look forward to your even better performance in the future.

[0074] The four sentences (1)-(4) in the above evaluation template are randomly selected and combined from the evaluation template data of the above four parts. The "follow-up text" part is used to fill in the actual follow-up text, and the "user name" part is used to fill in the actual user name. After filling in the actual follow-up text and user name, the complete feedback text can be generated.

[0075] In one possible implementation, when determining the evaluation template and generating the feedback text based on the user data, the interpretation information and question type information corresponding to the reading text, and the text reading evaluation corresponding to the reading text, some templates may require the user's name to be filled in. For example, in sentence (1) of the example evaluation template above, the "user's name" needs to be filled in based on the user data to generate the feedback text. At this time, it is necessary to judge the user's name in the user data. If the user's name in the user data is not in the target whitelist, it is determined that neither the obtained evaluation template nor the generated feedback text includes the user's name. That is, in order to ensure that there is no situation where the user's custom name may be a non-real name, such as using numbers, letters, or combinations of instruments as the user's name, or using abnormal names (such as the names of other objects) as the user's name, when determining the evaluation template, it is first determined whether the user's name is in the whitelist. Only if it is, the template that requires the "user's name" to be filled in will be determined in the final evaluation template. The names in the target whitelist are user names that have been pre-screened for reasonableness.

[0076] Figure 4 This is a flowchart illustrating an audio data processing method, as shown in another exemplary embodiment of this disclosure. Figure 3 As shown, the method further includes steps 401 and 402.

[0077] In step 401, if the question type to which the follow-up text belongs is a phonics question type, the text in the phonics part of the follow-up text is transcribed into phonetic symbols.

[0078] In step 402, the phonetic transcription of the follow-up text and the feedback text are combined to form the feedback speech, which is then sent to the user corresponding to the audio to be evaluated.

[0079] Since conventional speech synthesis methods cannot effectively synthesize the letters pronounced during spelling in the follow-up text included in phonics exercises, when the follow-up text belongs to the phonics exercise type, the follow-up text can first be transcribed into phonetic symbols. The transcribed follow-up text is then used as the "follow-up text" to be filled in the evaluation template. After determining the feedback text, speech synthesis technology can then be used to synthesize the feedback text, so that the synthesized speech can better pronounce the letters in the spelling.

[0080] Figure 5 This is a structural block diagram of an audio data processing apparatus according to an exemplary embodiment of the present disclosure. Figure 5 As shown, the device includes: a first acquisition module 10, used to acquire the audio to be evaluated and the text to be read aloud; an extraction module 20, used to extract the speech to be evaluated from the audio to be evaluated; a scoring module 30, used to score one or more segments of speech content corresponding to each word in the text to be read aloud in the speech to be evaluated, so as to obtain one or more scoring results; an evaluation module 40, used to determine the text reading evaluation corresponding to the text to be read aloud based on the one or more scoring results, wherein the text reading evaluation includes positive evaluation, negative evaluation and neutral evaluation; a feedback module 50, used to determine an evaluation template based on the text reading evaluation corresponding to the text to be read aloud and generate feedback text; and a synthesis module 60, used to synthesize the feedback text into feedback speech and send it to the user corresponding to the audio to be evaluated.

[0081] The above technical solution processes the audio of the user's pronunciation practiced on the text, directly obtaining the user's pronunciation evaluation without human judgment. It also scores multiple pronunciations of the same word to determine the final evaluation. This comprehensive evaluation, based on scoring all of the user's pronunciations, results in a more objective, reasonable, and complete text practice evaluation. Furthermore, the pronunciation evaluation can be automatically classified as positive, negative, or neutral, and feedback audio can be automatically generated based on this evaluation. This allows users to more intuitively understand their pronunciation problems. This approach provides a highly objective evaluation of the user's pronunciation without any subjective human intervention, while also saving on labor costs.

[0082] In one possible implementation, the evaluation module 40 is further configured to: determine the word reading evaluation corresponding to each word in the reading text based on the scoring results of one or more audio segments corresponding to each word in the reading text, wherein the word reading evaluation includes the positive evaluation, the negative evaluation, and the neutral evaluation; and determine the text reading evaluation corresponding to the reading text based on the word reading evaluation corresponding to each word in the reading text.

[0083] Figure 6 This is a structural block diagram of an audio data processing apparatus according to yet another exemplary embodiment of the present disclosure. Figure 6 As shown, the device further includes: a word scoring submodule 401, used to take the highest score among the scoring results of one or more audio segments corresponding to each word in the follow-up text as the word score result corresponding to each word in the follow-up text; a word evaluation submodule 402, used to determine the word follow-up evaluation corresponding to each word in the follow-up text according to a preset evaluation standard and the word scoring result; a first text evaluation submodule 403, used to take the word follow-up evaluation corresponding to the word with the longest text length and whose corresponding word follow-up evaluation is not the neutral evaluation as the text follow-up evaluation corresponding to the follow-up text when the word follow-up evaluation corresponding to each word in the follow-up text includes the positive evaluation or the negative evaluation; and a second text evaluation submodule 404, used to determine the text follow-up evaluation corresponding to the follow-up text as a neutral evaluation when the word follow-up evaluation corresponding to each word in the follow-up text does not include the positive evaluation or the negative evaluation.

[0084] In one possible implementation, such as Figure 6 As shown, the device further includes a second acquisition module 70, used to acquire user data corresponding to the user and the interpretation information and question type information corresponding to the reading text; the feedback module 50 is also used to: determine the evaluation template and generate the feedback text based on the user data, the interpretation information and question type information corresponding to the reading text, and the text reading evaluation corresponding to the reading text.

[0085] In one possible implementation, the feedback module 50 is further configured to: determine that neither the obtained evaluation template nor the generated feedback text contains the user name if the user name in the user data is not in the target whitelist, wherein the names in the target whitelist are user names that have been pre-filtered for reasonableness.

[0086] In one possible implementation, the synthesis module 60 is further configured to: when the question type to which the follow-up text belongs is a phonics question type, transcribing the text in the phonics part of the follow-up text into phonetic symbols; and synthesizing the phonetically transcribed follow-up text and the feedback text into the feedback speech and sending it to the user corresponding to the audio to be evaluated.

[0087] In one possible implementation, the extraction module 20 is further configured to: extract the voice to be evaluated from the audio to be evaluated based on the user's user information, wherein the user information includes any one of the user's age, user's gender, and user's voice timbre.

[0088] The following is for reference. Figure 7 The diagram illustrates a structural schematic of an electronic device 700 suitable for implementing embodiments of the present disclosure. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0089] like Figure 7 As shown, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0090] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0091] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the methods of embodiments of this disclosure.

[0092] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0093] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0094] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0095] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire an audio to be evaluated and a text to be read aloud; extract the speech to be evaluated from the audio to be evaluated; score one or more segments of speech content corresponding to each word in the text to be read aloud, to obtain one or more scoring results; determine a text reading evaluation corresponding to the text to be read aloud based on the one or more scoring results, the text reading evaluation including positive evaluation, negative evaluation, and neutral evaluation; determine an evaluation template based on the text reading evaluation corresponding to the text to be read aloud and generate feedback text; and synthesize the feedback text into feedback speech and send it to the user corresponding to the audio to be evaluated.

[0096] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0097] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0098] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules are not necessarily limiting in certain circumstances; for example, the first acquisition module can also be described as "a module for acquiring the audio to be evaluated and the text to be read aloud".

[0099] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0100] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0101] According to one or more embodiments of this disclosure, Example 1 provides an audio data processing method, the method comprising:

[0102] Obtain the audio to be evaluated and the text to be read aloud;

[0103] Extract the speech to be evaluated from the audio to be evaluated;

[0104] Each segment or multiple segments of speech content corresponding to each word in the text to be read aloud are scored to obtain one or more scoring results.

[0105] The text reading evaluation corresponding to the text is determined by one or more scoring results, and the text reading evaluation includes positive evaluation, negative evaluation and neutral evaluation;

[0106] Based on the text reading evaluation corresponding to the text reading text, an evaluation template is determined and feedback text is generated;

[0107] The feedback text is synthesized into feedback speech and then sent to the user corresponding to the audio to be evaluated.

[0108] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, wherein determining the text reading evaluation corresponding to the reading text through the one or more scoring results includes:

[0109] Based on the scoring results of one or more audio segments corresponding to each word in the text to be read aloud, the word reading evaluation corresponding to each word in the text to be read aloud is determined, and the word reading evaluation includes the positive evaluation, the negative evaluation and the neutral evaluation;

[0110] The text reading evaluation is determined based on the word reading evaluation corresponding to each word in the reading text.

[0111] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 2, wherein determining the word reading evaluation corresponding to each word in the reading text based on the scoring results of one or more audio segments corresponding to each word in the reading text includes:

[0112] The highest score among the scores of one or more audio segments corresponding to each word in the text to be read aloud is taken as the word score for each word in the text to be read aloud.

[0113] The word reading evaluation for each word in the reading text is determined based on the preset evaluation criteria and the word scoring results.

[0114] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 2, wherein determining the text reading evaluation corresponding to the reading text based on the word reading evaluation corresponding to each word in the reading text includes:

[0115] If the word reading evaluation corresponding to each word in the reading text includes either the positive evaluation or the negative evaluation, the word reading evaluation corresponding to the word with the longest text length and whose corresponding word reading evaluation is not the neutral evaluation shall be taken as the text reading evaluation corresponding to the reading text.

[0116] If the word reading evaluations corresponding to each word in the reading text do not include the positive or negative evaluations, the text reading evaluation corresponding to the reading text will be determined as a neutral evaluation.

[0117] According to one or more embodiments of this disclosure, Example 5 provides the method of Example 1, the method further comprising:

[0118] Obtain the user data corresponding to the user and the interpretation information and question type information corresponding to the text to be read aloud;

[0119] The step of determining the evaluation template and generating feedback text based on the text reading evaluation corresponding to the reading text includes:

[0120] The evaluation template is determined and the feedback text is generated based on the user data, the definition information and question type information corresponding to the reading text, and the text reading evaluation corresponding to the reading text.

[0121] According to one or more embodiments of this disclosure, Example 6 provides the method of Example 5, wherein determining the evaluation template and generating the feedback text based on the user data, the definition information and question type information corresponding to the follow-up text, and the text follow-up evaluation corresponding to the follow-up text includes:

[0122] If the user's name in the user data is not in the target whitelist, it is determined that neither the obtained evaluation template nor the generated feedback text includes the user's name, wherein the names in the target whitelist are user names that have been pre-screened for reasonableness.

[0123] According to one or more embodiments of this disclosure, Example 7 provides the method of Example 1, wherein the step of synthesizing the feedback text into feedback speech and sending it to the user corresponding to the audio to be evaluated includes:

[0124] If the question type to which the text to be read belongs is a phonics question type, the text in the phonics part of the text to be read will be transcribed into phonetic symbols;

[0125] The feedback speech, which is synthesized from the phonetic transcription of the follow-up text and the feedback text, is then sent to the user corresponding to the audio to be evaluated.

[0126] According to one or more embodiments of this disclosure, Example 8 provides the method of Example 1, wherein extracting the speech to be evaluated from the audio to be evaluated includes:

[0127] The voice to be evaluated is extracted from the audio to be evaluated based on the user's user information, which includes any one of the user's age, gender, and voice timbre.

[0128] According to one or more embodiments of this disclosure, Example 9 provides an audio data processing apparatus, the apparatus comprising:

[0129] The first acquisition module is used to acquire the audio to be evaluated and the text to be read aloud.

[0130] The extraction module is used to extract the speech to be evaluated from the audio to be evaluated;

[0131] The scoring module is used to score one or more segments of speech content that correspond to each word in the text to be read aloud in the speech to be evaluated, so as to obtain one or more scoring results.

[0132] The evaluation module is used to determine the text reading evaluation corresponding to the text reading based on the one or more scoring results. The text reading evaluation includes positive evaluation, negative evaluation, and neutral evaluation.

[0133] The feedback module is used to determine the evaluation template and generate feedback text based on the text reading evaluation corresponding to the reading text.

[0134] The synthesis module is used to synthesize the feedback text into feedback speech and send it to the user corresponding to the audio to be evaluated.

[0135] According to one or more embodiments of the present disclosure, Example 10 provides a computer-readable medium having a computer program stored thereon that, when executed by a processing device, implements the steps of the method described in any one of Examples 1-8.

[0136] According to one or more embodiments of this disclosure, Example 11 provides an electronic device, including:

[0137] A storage device on which computer programs are stored;

[0138] A processing device for executing the computer program in the storage device to implement the steps of any one of the methods in Examples 1-8.

[0139] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0140] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0141] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.

Claims

1. An audio data processing method, characterized in that, The method includes: Obtain the audio to be evaluated, the text to be read aloud, the corresponding explanation information and question type information of the text to be read aloud; Extract the speech to be evaluated from the audio to be evaluated; Each segment or multiple segments of speech content corresponding to each word in the text to be read aloud are scored to obtain one or more scoring results. The text reading evaluation corresponding to the text is determined by one or more scoring results, and the text reading evaluation includes positive evaluation, negative evaluation and neutral evaluation; Determine the first evaluation template data and / or the second evaluation template data for the text reading evaluation, and / or, based on the question type information, determine the third evaluation template data and / or the fourth evaluation template data for the interpretation information; Based on at least one of the first evaluation template data, the second evaluation template data, the third evaluation template data, and the fourth evaluation template data, an evaluation template for evaluating the user is determined, and feedback text for evaluating the user is generated based on the evaluation template. The feedback text is synthesized into feedback speech and then sent to the user corresponding to the audio to be evaluated.

2. The method according to claim 1, characterized in that, The process of determining the text reading evaluation corresponding to the reading text through one or more scoring results includes: Based on the scoring results of one or more audio segments corresponding to each word in the text to be read aloud, the word reading evaluation corresponding to each word in the text to be read aloud is determined, and the word reading evaluation includes the positive evaluation, the negative evaluation and the neutral evaluation; The text reading evaluation is determined based on the word reading evaluation corresponding to each word in the reading text.

3. The method according to claim 2, characterized in that, The step of determining the word-based pronunciation evaluation for each word in the text based on the scoring results of one or more audio segments corresponding to each word in the text includes: The highest score among the scores of one or more audio segments corresponding to each word in the text to be read aloud is taken as the word score for each word in the text to be read aloud. The word reading evaluation for each word in the reading text is determined based on the preset evaluation criteria and the word scoring results.

4. The method according to claim 2, characterized in that, The step of determining the text reading evaluation corresponding to the reading text based on the word reading evaluation corresponding to each word in the reading text includes: If the word reading evaluation corresponding to each word in the reading text includes either the positive evaluation or the negative evaluation, the word reading evaluation corresponding to the word with the longest text length and whose corresponding word reading evaluation is not the neutral evaluation shall be taken as the text reading evaluation corresponding to the reading text. If the word reading evaluations corresponding to each word in the reading text do not include the positive or negative evaluations, the text reading evaluation corresponding to the reading text will be determined as a neutral evaluation.

5. The method according to claim 1, characterized in that, The method further includes: Obtain the user data corresponding to the user; The step of generating feedback text for evaluating users based on the evaluation template includes: At least one of the user data, the reading text, and the explanatory information is filled into the evaluation template, and feedback text for evaluating the user is generated based on the filled evaluation template.

6. The method according to claim 5, characterized in that, The step of generating feedback text for evaluating users based on the evaluation template includes: If the user's name in the user data is not in the target whitelist, it is determined that neither the filled evaluation template nor the feedback text generated based on the filled evaluation template for evaluating the user includes the user's name. The names in the target whitelist are user names that have been pre-screened for reasonableness.

7. The method according to claim 1, characterized in that, The step of synthesizing the feedback text into feedback speech and sending it to the user corresponding to the audio to be evaluated includes: If the question type to which the text to be read belongs is a phonics question type, the text in the phonics part of the text to be read will be transcribed into phonetic symbols; The feedback speech, which is synthesized from the phonetic transcription of the reading text and the feedback text, is then sent to the user corresponding to the audio to be evaluated.

8. The method according to claim 1, characterized in that, The extraction of the speech to be evaluated from the audio to be evaluated includes: The voice to be evaluated is extracted from the audio to be evaluated based on the user's user information, which includes any one of the user's age, gender, and voice timbre.

9. An audio data processing device, characterized in that, The device includes: The first acquisition module is used to acquire the audio to be evaluated, the text to be read aloud, the corresponding explanation information and question type information of the text to be read aloud; The extraction module is used to extract the speech to be evaluated from the audio to be evaluated; The scoring module is used to score one or more segments of speech content that correspond to each word in the text to be read aloud in the speech to be evaluated, so as to obtain one or more scoring results. The evaluation module is used to determine the text reading evaluation corresponding to the text reading based on the one or more scoring results. The text reading evaluation includes positive evaluation, negative evaluation and neutral evaluation. The feedback module is used to determine first evaluation template data and / or second evaluation template data for the text reading evaluation, and / or, based on the question type information, determine third evaluation template data and / or fourth evaluation template data for the interpretation information, and determine an evaluation template for evaluating the user based on at least one of the first evaluation template data, the second evaluation template data, the third evaluation template data and the fourth evaluation template data, and generate feedback text for evaluating the user based on the evaluation template; The synthesis module is used to synthesize the feedback text into feedback speech and send it to the user corresponding to the audio to be evaluated.

10. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processing device, it implements the steps of the method described in any one of claims 1-8.

11. An electronic device, characterized in that, include: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Reading content display method and device, medium and computing equipment

    CN111694622A