Spoken language pronunciation evaluation feedback method, device and computer equipment

By acquiring the speech features of the speech to be evaluated and selecting the feedback processing mode, the problem of inaccurate evaluation results in computer-aided pronunciation training systems is solved, targeted pronunciation feedback is provided, and learners' learning motivation and efficiency are improved.

CN122435945APending Publication Date: 2026-07-21GUANGZHOU XIBEISI INTELLIGENT TECHNOLOGY CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU XIBEISI INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2025-01-20
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing computer-aided pronunciation training systems lack effective feedback on the reasons for learners' low scores, leading to inaccurate assessment results and affecting learners' motivation for oral practice.

Method used

By acquiring the speech features of the speech to be evaluated, selecting the feedback processing mode, and using a preset compensation algorithm to process the quality assessment information to correct non-pronunciation problems such as environmental noise, or by determining pronunciation defect information based on speech features and reference text, targeted feedback results are provided.

Benefits of technology

It improves the learning motivation and efficiency of oral practice learners, and helps learners identify and improve pronunciation defects through accurate feedback results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122435945A_ABST
    Figure CN122435945A_ABST
Patent Text Reader

Abstract

The application relates to a spoken language pronunciation evaluation feedback method and device and computer equipment, and relates to the technical field of data processing. The method comprises the following steps: obtaining first quality evaluation information of a to-be-evaluated voice and voice features of the to-be-evaluated voice; when the first quality evaluation information does not reach a pre-set quality evaluation condition, a feedback processing mode of the to-be-evaluated voice is obtained based on the voice features; when the feedback processing mode is a first feedback processing mode, the first quality evaluation information is processed based on the voice features and a pre-set compensation algorithm to obtain corresponding second quality evaluation information, and a spoken language pronunciation evaluation feedback result of the to-be-evaluated voice is obtained according to the second quality evaluation information; and when the feedback processing mode is a second feedback processing mode, defect information of the to-be-evaluated voice is determined based on the voice features and a corresponding reference text of the to-be-evaluated voice, and the defect information is taken as the spoken language pronunciation evaluation feedback result of the to-be-evaluated voice, so that the accuracy of the evaluation feedback result obtained by a spoken language learner is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus and computer device for evaluating and providing feedback on spoken pronunciation. Background Technology

[0002] With societal development, foreign language learning has become an urgent need for people. Oral communication skills are the most important aspect of foreign language proficiency, and the ability to pronounce accurately and express meaning precisely is a crucial indicator of foreign language ability. Traditional methods of learning spoken foreign languages ​​rely solely on imitation, reading aloud, and simulated dialogues with other learners, or on hiring a teacher for one-on-one oral instruction.

[0003] However, by practicing spoken English through imitation, reading aloud, and simulated dialogues with other students, it is difficult to identify problems and deficiencies in our spoken English skills, thus making it impossible to improve our spoken English skills in a targeted manner. In addition, hiring a teacher for one-on-one spoken English instruction is costly and highly dependent on the teacher's ability.

[0004] With the development of speech processing and natural language processing technologies, various computer devices are able to understand human language more accurately. Automatic pronunciation assessment in computer-aided pronunciation training primarily focuses on scoring the speaker's pronunciation quality—that is, evaluating the learner's speech—lacking feedback on the reasons for low scores. However, low scores can be related to environmental noise, speech rate, volume, and other factors. These non-pronunciation issues leading to low scores can negatively impact learners' motivation for oral practice and provide misleading feedback.

[0005] Therefore, this paper proposes a spoken pronunciation assessment feedback method to address the problems of inaccurate evaluation results and lack of analysis of reasons for low scores in automatic pronunciation assessment in computer-assisted pronunciation training. Summary of the Invention

[0006] Therefore, it is necessary to provide a method, device, computer equipment, computer-readable storage medium, and computer program product for oral pronunciation assessment and feedback that can accurately analyze and provide feedback on the practice of oral language learners, addressing the aforementioned technical problems.

[0007] Firstly, this application provides a method for assessing and providing feedback on spoken pronunciation, including:

[0008] Obtain the first quality assessment information of the speech to be evaluated, as well as the speech features of the speech to be evaluated;

[0009] If the first quality assessment information does not meet the preset quality assessment conditions, the feedback processing mode of the speech to be evaluated is obtained based on the speech features; the feedback processing mode includes a first feedback processing mode and a second feedback processing mode.

[0010] When the feedback processing mode is the first feedback processing mode, the first quality assessment information is processed based on the speech features and the preset compensation algorithm to obtain the corresponding second quality assessment information, and the spoken pronunciation assessment feedback result of the speech to be evaluated is obtained according to the second quality assessment information.

[0011] When the feedback processing mode is the second feedback processing mode, the defect information of the speech to be evaluated is determined based on the speech features and the reference text corresponding to the speech to be evaluated, and the defect information is used as the speech pronunciation evaluation feedback result of the speech to be evaluated.

[0012] In one embodiment, the step of obtaining the feedback processing mode of the speech to be evaluated based on the speech features includes:

[0013] Based on the speech features, the defect type of the speech to be evaluated and the degree of influence of the defect type on the first quality assessment information are determined; the defect type is used to characterize the reason why the first quality assessment information of the speech to be evaluated fails to meet the preset quality assessment conditions.

[0014] If the defect type belongs to a preset type and the impact value is greater than a preset threshold, the feedback processing mode of the speech to be evaluated is determined to be the first feedback processing mode.

[0015] If the defect type does not belong to the preset type, and / or the impact value is less than or equal to the preset threshold, the feedback processing mode of the speech to be evaluated is determined to be the second feedback processing mode.

[0016] In one embodiment, obtaining the spoken pronunciation evaluation feedback result of the speech to be evaluated based on the second quality assessment information includes:

[0017] The second quality assessment information and the defect type are used as the oral pronunciation assessment feedback result of the speech to be evaluated.

[0018] In one embodiment, determining the defect type of the speech to be evaluated based on the speech features, and the degree of influence of the defect type on the first quality assessment information, includes:

[0019] Based on the speech features, the phoneme-level Mel-spectral coefficients, phoneme-level posterior probabilities, and multidimensional sound features corresponding to the speech to be evaluated are determined; the multidimensional sound features are used to characterize at least one of the speech rate information, noise information, and volume information of the speech to be evaluated.

[0020] The phoneme-level Mel-spectral coefficients, the phoneme-level posterior probabilities, and the multidimensional sound features are input into a pre-trained factor feedback model to obtain the defect type of the speech to be evaluated and the degree of influence of the defect type on the first quality assessment information.

[0021] In one embodiment, the first quality assessment information includes a first assessment value, and the second quality assessment information includes a second assessment value;

[0022] The process of processing the first quality assessment information to obtain the corresponding second quality assessment information includes:

[0023] Based on the first evaluation value and the degree of influence value, a second evaluation value corresponding to the first evaluation value is obtained.

[0024] In one embodiment, determining the defect information of the speech to be evaluated based on the speech features and the reference text corresponding to the speech to be evaluated includes:

[0025] Based on the speech features, obtain the speech phoneme sequence corresponding to the speech to be evaluated, and obtain the text phoneme sequence of the reference text corresponding to the speech to be evaluated;

[0026] Based on the speech phoneme sequence and the text phoneme sequence, the defect information of the speech to be evaluated is determined.

[0027] In one embodiment, the step of obtaining the feedback processing mode of the speech to be evaluated based on the speech features when the first quality assessment information does not meet the preset quality assessment conditions includes:

[0028] If the first quality assessment information does not meet the preset quality assessment conditions, the abnormal pronunciation detection result of the speech to be evaluated is obtained; the abnormal pronunciation detection result is used to characterize the degree of matching between the speech to be evaluated and the corresponding reference text.

[0029] When the abnormal pronunciation detection result indicates that the matching degree value is greater than or equal to a preset matching value, the feedback processing mode of the speech to be evaluated is obtained based on the speech features.

[0030] In one embodiment, after obtaining the abnormal pronunciation detection result of the speech to be evaluated, the method further includes:

[0031] If the abnormal pronunciation detection result indicates that the matching degree value is less than the preset matching value, the first quality assessment information is used as the oral pronunciation evaluation feedback result of the speech to be evaluated.

[0032] Secondly, this application also provides a spoken pronunciation assessment and feedback device, comprising:

[0033] The data acquisition module is used to acquire the first quality assessment information of the speech to be evaluated, as well as the speech features of the speech to be evaluated;

[0034] The processing mode determination module is used to obtain the feedback processing mode of the speech to be evaluated based on the speech features when the first quality assessment information does not meet the preset quality assessment conditions; the feedback processing mode includes a first feedback processing mode and a second feedback processing mode.

[0035] The feedback result module is used to process the first quality assessment information based on the speech features and a preset compensation algorithm to obtain the corresponding second quality assessment information when the feedback processing mode is the first feedback processing mode, and to obtain the spoken pronunciation evaluation feedback result of the speech to be evaluated based on the second quality assessment information; it is also used to determine the defect information of the speech to be evaluated based on the speech features and the reference text corresponding to the speech to be evaluated when the feedback processing mode is the second feedback processing mode, and to use the defect information as the spoken pronunciation evaluation feedback result of the speech to be evaluated.

[0036] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0037] Obtain the first quality assessment information of the speech to be evaluated, as well as the speech features of the speech to be evaluated;

[0038] If the first quality assessment information does not meet the preset quality assessment conditions, the feedback processing mode of the speech to be evaluated is obtained based on the speech features; the feedback processing mode includes a first feedback processing mode and a second feedback processing mode.

[0039] When the feedback processing mode is the first feedback processing mode, the first quality assessment information is processed based on the speech features and the preset compensation algorithm to obtain the corresponding second quality assessment information, and the spoken pronunciation assessment feedback result of the speech to be evaluated is obtained according to the second quality assessment information.

[0040] When the feedback processing mode is the second feedback processing mode, the defect information of the speech to be evaluated is determined based on the speech features and the reference text corresponding to the speech to be evaluated, and the defect information is used as the speech pronunciation evaluation feedback result of the speech to be evaluated.

[0041] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:

[0042] Obtain the first quality assessment information of the speech to be evaluated, as well as the speech features of the speech to be evaluated;

[0043] If the first quality assessment information does not meet the preset quality assessment conditions, the feedback processing mode of the speech to be evaluated is obtained based on the speech features; the feedback processing mode includes a first feedback processing mode and a second feedback processing mode.

[0044] When the feedback processing mode is the first feedback processing mode, based on the speech features and the preset compensation algorithm, the second quality assessment information corresponding to the first quality assessment information is determined, and the spoken pronunciation assessment feedback result of the speech to be evaluated is obtained according to the second quality assessment information.

[0045] When the feedback processing mode is the second feedback processing mode, the defect information of the speech to be evaluated is determined based on the speech features and the reference text corresponding to the speech to be evaluated, and the defect information is used as the speech pronunciation evaluation feedback result of the speech to be evaluated.

[0046] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:

[0047] Obtain the first quality assessment information of the speech to be evaluated, as well as the speech features of the speech to be evaluated;

[0048] If the first quality assessment information does not meet the preset quality assessment conditions, the feedback processing mode of the speech to be evaluated is obtained based on the speech features; the feedback processing mode includes a first feedback processing mode and a second feedback processing mode.

[0049] When the feedback processing mode is the first feedback processing mode, the first quality assessment information is processed based on the speech features and the preset compensation algorithm to obtain the corresponding second quality assessment information, and the spoken pronunciation assessment feedback result of the speech to be evaluated is obtained according to the second quality assessment information.

[0050] When the feedback processing mode is the second feedback processing mode, the defect information of the speech to be evaluated is determined based on the speech features and the reference text corresponding to the speech to be evaluated, and the defect information is used as the speech pronunciation evaluation feedback result of the speech to be evaluated.

[0051] The present application provides a spoken pronunciation evaluation feedback method, apparatus, computer device, computer-readable storage medium, and computer program product. The spoken pronunciation evaluation feedback method acquires first quality assessment information of the speech to be evaluated, as well as the speech features of the speech to be evaluated. If the first quality assessment information does not meet preset quality assessment conditions, it indicates that the quality of the speech to be evaluated is relatively poor. In this case, a feedback processing mode for the speech to be evaluated is acquired based on the speech features. When the feedback processing mode is the first feedback processing mode, the first quality assessment information is processed based on the speech features and a preset compensation algorithm to obtain corresponding second quality assessment information, and a spoken pronunciation evaluation feedback result for the speech to be evaluated is obtained based on the second quality assessment information. When the feedback processing mode is the second feedback processing mode, defect information of the speech to be evaluated is determined based on the speech features and the reference text corresponding to the speech to be evaluated, and the defect information is used as the spoken pronunciation evaluation feedback result for the speech to be evaluated. In other words, when it is known that the quality of the speech to be evaluated is poor, this application processes and analyzes the speech features of the speech to be evaluated, and determines different content of the feedback information to be given to the oral practicer based on the different analysis results. This enables the oral practicer to be provided with more targeted evaluation feedback results for the speech to be evaluated, which helps to improve the accuracy of the evaluation feedback results obtained by the oral practicer.

[0052] Furthermore, when the feedback processing mode is determined to be the first feedback processing mode, this application compensates for the first quality assessment information of the speech to be evaluated based on the speech features of the speech to be evaluated and a preset compensation algorithm. This is to avoid the influence of environmental noise and other issues other than pronunciation problems on the quality assessment results of the speech to be evaluated, thereby improving the accuracy of the evaluation feedback results of the speech to be tested and enhancing the learning motivation of the speech learners. Moreover, when the feedback processing mode is determined to be the second feedback processing mode, this application reflects the specific defect information of the speech to be evaluated in the evaluation feedback results given to the speech learners. This allows the speech learners to know the specific defects in their pronunciation, enabling them to conduct targeted training based on these defects. This further enhances the learning motivation and accelerates the learning efficiency of the speech learners. Attached Figure Description

[0053] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0054] Figure 1 This is a diagram illustrating the application environment of a spoken pronunciation assessment and feedback method in one embodiment.

[0055] Figure 2 This is a flowchart illustrating a spoken pronunciation assessment and feedback method in one embodiment;

[0056] Figure 3 This is a schematic diagram of a refined feedback system for oral assessment in one embodiment;

[0057] Figure 4 This is a schematic diagram of the evaluation feedback system structure in one embodiment;

[0058] Figure 5 This is a structural block diagram of a spoken pronunciation assessment and feedback device in one embodiment;

[0059] Figure 6 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0061] Computer-Aided Pronunciation Training (CAPT) is an innovative self-directed language learning solution designed to help non-native speakers improve their foreign language pronunciation skills without relying on professional language teachers. Through a computerized learning system, CAPT offers a flexible and cost-effective learning approach. Automatic Pronunciation Assessment (APA) plays a central role in CAPT, providing learners with comprehensive, real-time feedback across multiple dimensions of pronunciation, helping them achieve significant progress in foreign language pronunciation. Compared to traditional classroom teaching methods, CAPT demonstrates greater flexibility and affordability.

[0062] APA (Pronunciation Assessment) technology has been extensively studied due to its broad application potential. Research focuses on how to evaluate pronunciation quality, specifically using APA technology to score users' pronunciation and provide scoring metrics such as accuracy, fluency, and completeness at different levels, including phonemes, words, sentences, and passages. The evaluation methods for these scoring metrics have evolved from traditional machine learning scoring models to employing advanced techniques such as deep neural networks (DNNs), recurrent neural networks (RNNs), and convolutional neural networks (CNNs). Recent research trends emphasize integrating information from different levels of detail and scoring dimensions into a single model for comprehensive evaluation. This approach fully leverages the complementarity of information across different levels and scoring dimensions, thereby continuously improving the performance and accuracy of the scoring model.

[0063] Current research on APA primarily focuses on scoring pronunciation quality, i.e., evaluating learners' speech, but lacks feedback on the reasons for low scores. A learner's low score may not be due to pronunciation problems, but rather to significant background noise in their environment, resulting in a noisy recording and causing significant bias in the scoring system. It could also be due to the learner speaking too fast or too slow, or speaking too loudly or too softly, leading to a low score.

[0064] Low scores not caused by pronunciation issues are common, especially in children's oral practice. This is a pain point and a source of complaints for learners using the CAPT system. These low scores, which are not due to pronunciation problems, will seriously affect learners' enthusiasm for oral practice. At the same time, the biased scoring will provide learners with misleading feedback.

[0065] To address the aforementioned problem that existing oral assessment systems lack effective analysis and feedback on the reasons for score deductions, i.e., inaccurate oral assessment feedback, this application provides an oral pronunciation assessment feedback method. This method involves simultaneously acquiring the speech features of the speech to be assessed upon obtaining the first quality assessment information. Then, if the first quality assessment information does not meet preset quality assessment conditions, a specific feedback processing mode is selected based on the speech features. If, based on the speech features, the first feedback processing mode is determined to be used, the first quality assessment information is processed based on the speech features and a preset compensation algorithm to obtain the corresponding second quality assessment information. The second quality assessment information is obtained by compensating and correcting the first quality assessment information. Based on the second quality assessment information, the oral pronunciation assessment feedback result is output to the oral practicer regarding the speech to be assessed. The corrected oral pronunciation assessment feedback result is conducive to improving the learning enthusiasm of the oral practicer. When the second feedback processing mode is determined based on the speech features, the defect information of the speech to be assessed is determined based on the speech features and the reference text corresponding to the speech to be assessed. The defect information is used as the oral pronunciation assessment feedback result of the speech to be assessed and output to the oral practicer so that the oral practicer can clearly understand what type of pronunciation defect problem he has.

[0066] The spoken pronunciation assessment and feedback method provided in this application embodiment can be applied to, for example, Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or located on the cloud or other network servers. The data storage system can be used to store data such as the spoken language learner's speech to be evaluated, the corresponding reference text, and preset compensation algorithms. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0067] In one exemplary embodiment, such as Figure 2 As shown, a method for assessing and providing feedback on spoken pronunciation is provided, which is then applied to... Figure 1Taking terminal 102 as an example, the explanation includes the following steps 201 to 204. Wherein:

[0068] Step 201: Obtain the first quality assessment information of the speech to be evaluated, as well as the speech features of the speech to be evaluated.

[0069] The speech to be evaluated can be the speech produced by the speaker during oral practice; the speech to be evaluated can be a sentence or a segment of speech, and this application does not make specific restrictions on the length of the speech to be evaluated.

[0070] Among them, speech features are the acoustic features of the speech to be evaluated. This application does not specifically limit the types of features included in the acoustic features.

[0071] The first quality assessment information can be the assessment result of a pronunciation evaluation system in related technologies on the speech of a spoken language learner; for example, it can be the pronunciation assessment result of the spoken language learner on the speech of a spoken language learner by CAPT.

[0072] For example, to improve the accuracy of pronunciation assessment results for spoken language learners' speech samples, firstly, the initial quality assessment information of the speech sample can be obtained, such as pronunciation assessment results related to the speech sample based on CAPT output, and the speech features of the speech sample can be obtained. The speech features of the speech sample are helpful for subsequent further analysis of the speech sample and can also improve the accuracy of the analysis.

[0073] This application does not specifically limit the order in which the first quality assessment information and speech features of the speech to be evaluated are obtained. One embodiment suggests that the speech features of the speech to be evaluated can be obtained after the first quality assessment information is acquired. In this way, if the first quality assessment information of the speech to be evaluated is good, it indicates that the speech produced by the oral practice student is relatively good, and further analysis of the reasons for score deductions may not be necessary, thus eliminating the need for the step of acquiring the speech features.

[0074] Step 202: If the first quality assessment information does not meet the preset quality assessment conditions, obtain the feedback processing mode of the speech to be evaluated based on the speech features; the feedback processing mode includes a first feedback processing mode and a second feedback processing mode.

[0075] The pre-set quality assessment conditions are reference values ​​or thresholds for quality assessment information about the speech to be evaluated, pre-set by relevant technical personnel based on requirements. For example, if the first quality assessment information includes a score, the pre-set quality assessment conditions will include a corresponding score threshold or an optimal range for the score. This application does not specifically limit the number and types of conditions included in the pre-set quality assessment conditions.

[0076] Feedback processing mode refers to the specific processing method of the first quality assessment information and speech features, as well as the specific information of the final output feedback result.

[0077] For example, after obtaining the first quality assessment information and speech features of the speech to be evaluated, the first quality assessment information of the speech to be evaluated can be compared with the corresponding preset quality assessment conditions. Specifically, each sub-information in the first quality assessment information can be compared with each corresponding sub-condition in the preset quality assessment conditions. If the comparison result shows that the first quality assessment information does not meet the preset quality assessment conditions, it indicates that the quality of the speech to be evaluated has not met the preset requirements, that is, the quality of the speech to be evaluated is not good enough. The fact that the first quality assessment information does not meet the preset quality assessment conditions can also be said to mean that the speech to be evaluated produced by the oral practicer differs significantly from the reference text they read, that is, the accuracy of the speech to be evaluated produced by the oral practicer is poor. At this time, based on the speech features of the speech to be evaluated, the feedback processing mode of the speech to be evaluated can be determined so that the determined feedback processing mode can be used to further process the speech features and the first quality assessment information of the speech to be evaluated, and obtain the feedback result information corresponding to the feedback processing mode.

[0078] Step 203: When the feedback processing mode is the first feedback processing mode, the first quality assessment information is processed based on the speech features and the preset compensation algorithm to obtain the corresponding second quality assessment information, and the spoken pronunciation assessment feedback result of the speech to be evaluated is obtained based on the second quality assessment information.

[0079] The second quality assessment information is obtained by performing compensation calculations on the first quality assessment information based on speech features and a preset compensation algorithm, and can be used as the final assessment result of the speech to be evaluated produced by the oral practicer.

[0080] The preset compensation algorithm is a pre-set algorithm by relevant technical personnel used to calculate the second quality assessment information based on the first quality assessment information.

[0081] For example, based on execution step 202, if it is learned that a first feedback processing mode is needed to further process the speech features and first quality assessment information of the speech to be evaluated, a pre-set compensation algorithm can be invoked to perform compensation calculation on the first quality assessment information to obtain the compensated second quality assessment information. The compensation operation performed here can improve the quality assessment information, for example, it can increase the relevant assessment scores in the quality assessment information. That is, the process of performing compensation calculation on the first quality assessment information is a process of removing unfavorable factors that affect the first quality assessment information of the speech to be evaluated, such as removing noise from environmental factors that affect the quality assessment information of the speech to be evaluated, thereby avoiding the adverse effects of factors other than non-pronunciation problems on the quality assessment results of the speech to be evaluated.

[0082] Then, the second quality assessment information obtained after the compensation calculation can be used as the spoken pronunciation evaluation feedback result of the speech to be evaluated. In addition to the second quality assessment information, the spoken pronunciation evaluation feedback result may also include other information as needed; this application does not impose specific limitations on this.

[0083] Step 204: When the feedback processing mode is the second feedback processing mode, the defect information of the speech to be evaluated is determined based on the speech features and the reference text corresponding to the speech to be evaluated, and the defect information is used as the speech pronunciation evaluation feedback result of the speech to be evaluated.

[0084] The reference text corresponding to the speech to be evaluated refers to either the baseline speech or the baseline speech text corresponding to the speech to be evaluated. That is, the reference text can be either audio text or text text.

[0085] Among them, defect information refers to the specific defects in the spoken language trainee's voice to be evaluated, such as pronunciation errors, unclear pronunciation, etc.

[0086] For example, if, based on execution step 202, it is learned that a second feedback processing mode is needed to further process the speech features and first quality assessment information of the speech to be evaluated, the reference text corresponding to the speech to be evaluated can be obtained first. Then, based on the speech features of the speech to be evaluated and the corresponding reference text, the specific defect information of the speech to be evaluated can be obtained to determine the reason for the poor quality assessment result of the speech to be evaluated. Subsequently, the defect information corresponding to the speech to be evaluated can be used as the oral pronunciation evaluation feedback result of the speech to be evaluated, so that the oral pronunciation learner can understand the reason for their poor pronunciation effect, and thus the oral pronunciation learner can carry out targeted practice based on the reason for the defect.

[0087] The spoken pronunciation evaluation feedback method provided in this application obtains first quality assessment information and speech features of the speech to be evaluated. If the first quality assessment information does not meet the preset quality assessment conditions, it indicates that the quality of the speech to be evaluated is relatively poor. At this time, a feedback processing mode for the speech to be evaluated is obtained based on the speech features. If the feedback processing mode is the first feedback processing mode, the first quality assessment information is processed based on the speech features and a preset compensation algorithm to obtain the corresponding second quality assessment information, and the spoken pronunciation evaluation feedback result of the speech to be evaluated is obtained based on the second quality assessment information. If the feedback processing mode is the second feedback processing mode, the defect information of the speech to be evaluated is determined based on the speech features and the reference text corresponding to the speech to be evaluated, and the defect information is used as the spoken pronunciation evaluation feedback result of the speech to be evaluated. In other words, when it is known that the quality of the speech to be evaluated is poor, this application processes and analyzes the speech features of the speech to be evaluated, and determines different content of the feedback information to be given to the oral practicer based on the different analysis results. This enables the oral practicer to be provided with more targeted evaluation feedback results for the speech to be evaluated, which helps to improve the accuracy of the evaluation feedback results obtained by the oral practicer.

[0088] Furthermore, when the feedback processing mode is determined to be the first feedback processing mode, this application compensates for the first quality assessment information of the speech to be evaluated based on the speech features of the speech to be evaluated and a preset compensation algorithm. This is to avoid the influence of environmental noise and other issues other than pronunciation problems on the quality assessment results of the speech to be evaluated, thereby improving the accuracy of the evaluation feedback results of the speech to be tested and enhancing the learning motivation of the speech learners. Moreover, when the feedback processing mode is determined to be the second feedback processing mode, this application reflects the specific defect information of the speech to be evaluated in the evaluation feedback results given to the speech learners. This allows the speech learners to know the specific defects in their pronunciation, enabling them to conduct targeted training based on these defects. This further enhances the learning motivation and accelerates the learning efficiency of the speech learners.

[0089] In an exemplary embodiment, the feedback processing mode for obtaining the speech to be evaluated based on speech features, performed in step 202 above, can be implemented by executing steps 221-223, wherein:

[0090] Step 221: Determine the defect type of the speech to be evaluated and the degree of influence of the defect type on the first quality assessment information based on the speech features; the defect type is used to characterize the defect reason why the first quality assessment information of the speech to be evaluated does not meet the preset quality assessment conditions.

[0091] Step 222: If the defect type belongs to a preset type and the impact value is greater than a preset threshold, determine the feedback processing mode of the speech to be evaluated as the first feedback processing mode.

[0092] Step 223: If the defect type does not belong to the preset type and / or the impact value is less than or equal to the preset threshold, determine the feedback processing mode of the speech to be evaluated as the second feedback processing mode.

[0093] These defect types include, for example, speech rate type, noise type, volume type, audio truncation type, etc. Speech rate type specifically includes excessively fast and excessively slow speech rates; volume type specifically includes excessively high and excessively low volume; noise type specifically includes human voice interference, wind noise interference, background music interference, animal sound interference, etc. These defect types refer to non-pronunciation problem influencing factors, which are different from the defect information mentioned above.

[0094] The influence level value is a percentage, used to characterize the type of defect in the speech being evaluated and its impact on the first quality assessment information of the speech being evaluated.

[0095] The preset types can be set according to needs, and can include at least the above-mentioned types such as too fast speech, too slow speech, too high volume, too low volume, human voice interference, wind noise interference, background music interference, animal sound interference, and audio truncation.

[0096] For example, the defect types that cause poor quality of the speech to be evaluated can be obtained through the speech features of the speech to be evaluated, and the influence degree value of the obtained defect type on the first quality assessment information can be obtained. Then, it can be selected to determine whether the obtained defect type of the speech to be evaluated belongs to a relevant preset type. If the determination result is that the defect type belongs to the preset type, the influence degree value and the size of the preset threshold corresponding to the influence degree value can be further determined. If the determination result is that the influence degree value is greater than the preset threshold, it is determined that the first feedback processing mode will be adopted to further process the speech features of the speech to be evaluated and the first quality assessment information.

[0097] If the determination result indicates that the defect type does not belong to the preset type, the second feedback processing mode can be selected to further process the speech features and first quality assessment information of the speech to be evaluated. Alternatively, if the determination result indicates that the defect type does not belong to the preset type, the magnitude of the impact value and the corresponding preset threshold can be further determined. Only if the magnitude of the impact value is less than or equal to the preset threshold will the second feedback processing mode be selected to further process the speech features and first quality assessment information of the speech to be evaluated.

[0098] Another alternative implementation is provided, in which, after determining the defect type of the speech to be evaluated and the degree of influence of the defect type on the first quality assessment information based on execution step 221, the degree of influence and the corresponding preset threshold can be directly judged. If the judgment result is that the degree of influence is less than or equal to the preset threshold, it is directly determined that the second feedback processing mode will be adopted to further process the speech features and the first quality assessment information of the speech to be evaluated.

[0099] The embodiment provides that by acquiring the defect types of the speech to be evaluated and the degree of influence of the defect types on the first quality assessment information, and through further analysis, the specific feedback processing mode for further processing of the speech features and the first quality assessment information of the speech to be evaluated can be determined. This helps to improve the accuracy of the selected feedback processing mode, thereby improving the accuracy of the final oral pronunciation evaluation feedback results provided to the oral practicer.

[0100] In an exemplary embodiment, for the step 203 above, which involves obtaining the spoken pronunciation evaluation feedback result of the speech to be evaluated based on the second quality assessment information, this application provides an optional implementation method in which the second quality assessment information and the defect type are used as the spoken pronunciation evaluation feedback result of the speech to be evaluated.

[0101] For example, when this application adopts the first feedback processing mode to further process the speech features and first quality assessment information of the speech to be evaluated, the final speech pronunciation evaluation feedback result output to the oral practicer may, in addition to the compensated second quality assessment information, further provide the defect type corresponding to the speech to be evaluated, so that the user knows the external factors that cause the poor quality assessment result of the speech to be evaluated, and what the specific type is, so that when outputting the speech to be evaluated in the next time, the external factors that cause the poor quality assessment result can be avoided as much as possible.

[0102] In an exemplary embodiment, the determination of the defect type of the speech to be evaluated based on speech features, and the degree of influence of the defect type on the first quality assessment information, performed in step 221 above, can be implemented by executing steps 213 and 214, wherein:

[0103] Step 213: Determine the phoneme-level Mel-spectral coefficients, phoneme-level posterior probabilities, and multidimensional sound features corresponding to the speech to be evaluated based on the speech features; the multidimensional sound features are used to characterize at least one of the speech rate information, noise information, and volume information of the speech to be evaluated.

[0104] Step 214: Input the phoneme-level Mel-Cepstral coefficients, phoneme-level posterior probabilities, and multidimensional sound features into the pre-trained factor feedback model to obtain the defect type of the speech to be evaluated and the degree of influence of the defect type on the first quality assessment information.

[0105] Mel-scale frequency cepstral coefficients (MFCCs), which are the coefficients that make up the Mel frequency cepstral spectrum, are a linear transformation of the logarithmic energy spectrum based on the nonlinear Mel scale of sound frequency. Posterior probability is an updated estimate of the probability of an event occurring after the data has been observed.

[0106] For example, speech features may include frame-level Mel-Cepstral coefficients, frame-level posterior probabilities, and timestamp information corresponding to the speech to be evaluated. By using some pre-provided feature extraction (conversion) algorithms, based on the processing of frame-level Mel-Cepstral coefficients, frame-level posterior probabilities, and timestamp information, phoneme-level Mel-Cepstral coefficients, phoneme-level posterior probabilities, and multidimensional sound features corresponding to the speech to be evaluated can be obtained. Then, the obtained phoneme-level Mel-Cepstral coefficients, phoneme-level posterior probabilities, and multidimensional sound features are input into a pre-trained factor feedback model to obtain the defect type of the speech to be evaluated and the degree of influence of the defect type on the first quality assessment information output by the factor feedback model.

[0107] By using the feature extraction algorithms and pre-trained factor feedback models provided in the above embodiments, the speech features of the speech to be evaluated can be processed, which helps to improve the accuracy of the defect types of the speech to be evaluated and the degree of influence of the defect types on the first quality assessment information, and also helps to improve the processing efficiency of related data.

[0108] In an exemplary embodiment, the first quality assessment information includes a first assessment value, and the second quality assessment information includes a second assessment value; the processing of the first quality assessment information in step 203 to obtain the corresponding second quality assessment information can be optionally performed as follows: based on the first assessment value and the degree of influence value, obtain the second assessment value corresponding to the first assessment value.

[0109] The first evaluation value is equivalent to the pronunciation score that CAPT gives to the spoken language learner's speech to be evaluated, while the second evaluation value is obtained by compensating for the first evaluation value.

[0110] For example, when performing compensation calculations on the second evaluation value corresponding to the first evaluation value, the calculation of the second evaluation value can be performed based on the first evaluation value and the degree of influence value, which helps to improve the rationality and accuracy of the calculated second evaluation value.

[0111] In an exemplary embodiment, the defect information of the speech to be evaluated, determined by step 204 based on speech features and the reference text corresponding to the speech to be evaluated, can be achieved by executing steps 241 and 242, wherein:

[0112] Step 241: Obtain the speech phoneme sequence corresponding to the speech to be evaluated based on speech features, and obtain the text phoneme sequence of the reference text corresponding to the speech to be evaluated;

[0113] Step 242: Based on the speech phoneme sequence and the text phoneme sequence, determine the defect information of the speech to be evaluated.

[0114] For example, the speech features of the speech to be evaluated can be input into a pre-trained pronunciation defect and error detection model to obtain the speech phoneme sequence corresponding to the speech output by the model. Then, the text phoneme sequence of the reference text corresponding to the speech to be evaluated can be obtained. Finally, the speech phoneme sequence corresponding to the speech to be evaluated and the text phoneme sequence of the reference text are aligned to obtain the defect information of the speech to be evaluated. This method can identify which factors in the speech to be evaluated contain pronunciation defects and errors, improving the accuracy and clarity of defect information. This helps oral language learners clearly understand their pronunciation problems and avoid them in future practice.

[0115] In an exemplary embodiment, step 202, which involves obtaining the feedback processing mode of the speech to be evaluated based on speech features when the first quality assessment information does not meet the preset quality assessment conditions, can be implemented by executing steps 311 and 322, wherein:

[0116] Step 311: If the first quality assessment information does not meet the preset quality assessment conditions, obtain the abnormal pronunciation detection result of the speech to be evaluated; the abnormal pronunciation detection result is used to characterize the degree of matching between the speech to be evaluated and the corresponding reference text.

[0117] Step 322: If the abnormal pronunciation detection result represents a matching degree value that is greater than or equal to a preset matching value, obtain the feedback processing mode of the speech to be evaluated based on the speech features.

[0118] For example, if it is learned that the first quality assessment information corresponding to the speech to be evaluated does not meet the preset quality assessment conditions, abnormal pronunciation detection (scrawl detection) can be performed on the speech to be evaluated first. For example, the speech to be evaluated and the corresponding reference text can be input together into a pre-trained abnormal pronunciation detection model to obtain the abnormal pronunciation detection result output by the abnormal pronunciation detection model. Based on the abnormal pronunciation detection result, it can be determined whether the speech to be evaluated is a sentence that the speaker is scrambling to read. Further, if the abnormal pronunciation detection result, which represents the degree of matching between the speech to be evaluated and the corresponding reference text, is greater than or equal to a preset matching value, it indicates that the speech to be evaluated is not scrambling to read. The speech to be evaluated can then undergo subsequent processing steps to output the speech pronunciation evaluation feedback result corresponding to the speech to be evaluated.

[0119] In this embodiment, if the first quality assessment information does not meet the preset quality assessment conditions, the system first detects whether the speech to be evaluated is garbled speech. This is to avoid performing subsequent processing steps on the speech to be evaluated if it is garbled speech, thus avoiding the waste of computing power, data storage space, and time of the relevant assessment system.

[0120] In an exemplary embodiment, after obtaining the abnormal pronunciation detection result of the speech to be evaluated in step 311, the method further includes: if the abnormal pronunciation detection result characterization matching degree value is less than a preset matching value, using the first quality assessment information as the spoken pronunciation evaluation feedback result of the speech to be evaluated.

[0121] If the abnormal pronunciation detection result, which indicates the degree of matching between the speech to be evaluated and the corresponding reference text, is less than the preset matching value, it means that the speech to be evaluated is being pronounced randomly by the speaker or someone else around the speaker, and is not being pronounced according to the reference text. In this case, it is not worthwhile to perform at least step 203 or step 204 above to obtain the relevant speech pronunciation evaluation feedback result. Therefore, an alternative implementation method is to directly use the first quality assessment information as the speech pronunciation evaluation feedback result of the speech to be evaluated.

[0122] Regarding the spoken pronunciation assessment feedback method provided in this application, such as Figure 3As shown, the spoken language learner's provided speech (voice) to be evaluated, and the corresponding reference text, will first be processed by any of the relevant spoken language assessment systems to obtain first quality assessment information for the speech. Then, the first quality assessment information and the speech features of the speech will be further processed by the assessment feedback system to obtain the assessment result and feedback result (spoken pronunciation assessment feedback result) for the speech. That is, this application can provide a refined spoken language assessment feedback system, which includes at least a spoken language assessment system and an assessment feedback system.

[0123] Furthermore, this application may also provide a specific structure for a refined feedback system for oral assessment, such as... Figure 4 As shown, it may include a feature extraction module (including multiple feature extraction algorithms), an abnormal pronunciation detection module (including an abnormal pronunciation detection model), a multi-factor feedback detection module (including a factor feedback model), a scoring compensation module (including a preset compensation algorithm), and a pronunciation defect and error detection module (including a pronunciation defect and error detection model). The abnormal pronunciation detection module is... Figure 4 The random read detection module in the middle.

[0124] For example, the spoken language assessment system mainly consists of three parts: an acoustic model, a forced alignment decoding network, and a scoring model. Among them:

[0125] The acoustic model is a neural network model, trained through frame-level classification, and is mainly used to predict the classification probability of each frame of audio based on acoustic features.

[0126] The forced alignment decoding network is constructed from the reference text phoneme sequence information and is mainly used to align audio with text in order to obtain the timestamp information corresponding to the phoneme sequence.

[0127] The scoring model is a neural network model, which is trained based on timestamp information, acoustic features, and posterior probability features output by the acoustic model, so that the output score approximates the expert scoring result.

[0128] For example, the feature extraction module is mainly used to extract the input features of the abnormal pronunciation detection module and the multi-factor feedback detection module. The input features are mainly divided into two categories: one is implicit features, which include phoneme-level posterior probability and phoneme-level acoustic features MFCC (Mel-frequency cepstral coefficients); the other is explicit features, which are mainly designed manually. The 13-dimensional artificial features (corresponding to the multi-dimensional sound features mentioned above) cover factors that can characterize speech rate, noise, volume, audio truncation, etc.

[0129] The feature extraction module can extract MFCC features from the audio; extract 13-dimensional artificial features based on the audio and timestamp information; extract GOP features (corresponding to the aforementioned posterior probabilities) representing the quality of pronunciation based on the posterior probability and timestamp information; and finally, extract phoneme-level features of these features based on the timestamp information. The GOP (Goodness of Pronunciation) feature is composed of the log phoneme posterior probability (LPP) and the log posterior ratio (LPR).

[0130] For example, the abnormal pronunciation detection module mainly consists of two parts: an ASR recognition model and an abnormal pronunciation detection model. The ASR recognition model can be a large-scale Encoder-Decoder audio recognition model based on Transformer. This model is pre-trained using a large amount of adult data to obtain a basic model, and then fine-tuned using a certain amount of children's data to make the model more suitable for children's recognition scenarios and improve the model's recognition accuracy. ASR stands for Automatic Speech Recognition, which is the process of converting sound into text.

[0131] The abnormal pronunciation detection model consists of two parallel neural networks. The first neural network calculates the similarity between the reference text and the recognized text output by the ASR recognition module. The second neural network is used to calculate the weight matrix between the reference text and the audio output by the ASR recognition model. The features of the last hidden layer of the two neural networks are fused and input into the fully connected layer for classification detection of whether there is mispronunciation relative to the reference text audio.

[0132] The training set for the nonsensical speech detection is divided into two categories. The first category consists of nonsensical speech data selected from reading scenarios based on the scoring and recognition models. Specifically, this involves selecting audio and text data with low word scores in the reference text and high word scores in the recognition text, and where words from the reference text appear less frequently in the recognition text. The second category consists of simulated text data. Specifically, this involves selecting reading audio with a high consistency rate between the reference text and the recognition text, and a high score between the reference text and the audio. This process yields data matching the reference text and the audio. Then, simulated nonsensical speech data is constructed based on the selected matching data. Specifically, the reference text of the matched data is randomly replaced with the reference text of other matched data. The consistency rate between the original reference text and the replaced text is kept low, ensuring that the replaced text does not match the audio, thus constructing nonsensical speech data.

[0133] For example, the multi-factor feedback detection module mainly consists of two parallel neural networks. The first neural network processes the phoneme-level MFCC features of the reference text, and the second neural network processes the phoneme-level GOP features of the reference text. The hidden layer features of the last layer of the two neural networks are concatenated and fused. The training objectives of the multi-factor feedback model are divided into two types: one is five-class prediction, predicting speech rate, noise, volume, audio truncation, and normal category; the other is prediction of the difference ratio value.

[0134] The training set for the multi-factor feedback detection module is primarily obtained through data simulation based on real-world scenario data. Specifically, simulated audio is constructed by adjusting the speech rate and volume of the original audio, and adding noise. The audio truncation data is constructed mainly by truncating portions of the original audio using timestamp information. Finally, the score difference ratio is calculated by comparing the scores of the original audio and the simulated audio, thus measuring the degree of influence of this factor on the scoring system's performance.

[0135] For example, the scoring compensation module mainly performs scoring compensation operations based on the score difference ratio, scoring results and classification categories output by the multi-factor feedback detection module. The specific process is as follows: when the classification category is one of speech rate, noise, volume, or truncation, and the score difference ratio is greater than the manually set threshold, scoring compensation can be performed based on the original scoring results to obtain a more accurate score after compensation. The specific calculation formula is shown in (1).

[0136] (1)

[0137] The pronunciation defect detection module, after the multi-factor feedback detection module rules out low scores due to non-human factors, uses a phoneme correction model to detect pronunciation defects and errors in the speaker's pronunciation, thus revealing the speaker's inherent pronunciation problems. For example, English oral learners commonly have inaccurate vowel pronunciation (e.g., / i: / and / ɪ / ), difficulty pronouncing consonants (e.g., / θ / and / ð / ), and incomplete pronunciation (e.g., incomplete pronunciation that sounds confused with other vowels, such as "than" and "then"). The phoneme correction model can detect and correct these pronunciation defects and errors.

[0138] The pronunciation defect error detection model combines a neural network model with a Connectionist Temporal Classification (CTC) structure. The model takes audio MFCC features as input and outputs a phoneme sequence of the audio content. By aligning this sequence with a reference text phoneme sequence, it can identify which phonemes contain pronunciation defects and errors. CTC, short for Connectionist Temporal Classification, is an algorithm used to solve classification problems for time-series data.

[0139] The oral pronunciation assessment feedback method provided in this application also provides an alternative implementation method, including the following steps S1-S5, wherein:

[0140] Step S1: The user audio and reference text input evaluation system outputs the scoring results, timestamp information, frame-level MFCC features, and frame-level posterior probability.

[0141] In step S2, the feature extraction module extracts phoneme-level GOP, phoneme-level MFCC, and artificial features used to characterize speech rate, noise, volume, and truncation information based on the input timestamp information, frame-level MFCC features, and frame-level posterior probability. These features are used as part of the input to the multi-factor feedback detection module.

[0142] Step S3: Whether the system performs multi-factor feedback depends on whether it performs random reading detection. Multi-factor feedback detection is performed only when the audio and text are determined to be non-random reading data by the abnormal pronunciation detection module; otherwise, the original scoring result is returned.

[0143] Step S4: The multi-factor feedback detection module predicts the category and difference ratio of multi-factor feedback based on the input phoneme-level GOP, phoneme-level MFCC, and artificial features.

[0144] Step S5: The feedback discrimination module determines whether to perform multi-factor feedback and score compensation based on the category and score difference ratio of the multi-factor feedback. If the category of the multi-factor feedback is one of speech rate, noise, volume, or truncation, and the score difference ratio is greater than the manually set threshold, then refined feedback on the reasons for the low score and score compensation based on the original score result are performed; otherwise, pronunciation defect and error detection are performed, and feedback is provided on the pronunciation defects and errors that exist in the user.

[0145] As can be seen, this method can provide detailed feedback on the reasons for users' pronunciation errors, guiding them to improve their pronunciation skills. At the same time, this method has a scoring compensation function, which can alleviate the problem of low scores caused by existing oral assessment methods due to issues such as excessively fast or slow speaking speed, background noise in the audio, excessively loud or soft volume, audio truncation, and human voice interference.

[0146] The refined feedback system for oral assessment provided in this application differs from traditional feedback that only scores. This method can provide the reasons for low scores given by the assessment system, and can also compensate for low scores caused by non-human factors, making the scoring feedback results more accurate and reliable. The refined feedback system for oral assessment can be applied to any existing oral pronunciation assessment system, and does not require additional manual annotation of training samples.

[0147] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0148] Based on the same inventive concept, this application also provides a spoken pronunciation evaluation feedback device for implementing the spoken pronunciation evaluation feedback method described above. The solution provided by this device is similar to the implementation scheme described in the above method; therefore, the specific limitations of one or more spoken pronunciation evaluation feedback device embodiments provided below can be found in the limitations of the spoken pronunciation evaluation feedback method described above, and will not be repeated here.

[0149] In one exemplary embodiment, such as Figure 5 As shown, a spoken pronunciation assessment feedback device 500 is provided, including: a data acquisition module 51, a processing mode determination module 52, and a feedback result module 53, wherein:

[0150] The data acquisition module 51 is used to acquire the first quality assessment information of the speech to be evaluated, as well as the speech features of the speech to be evaluated.

[0151] The processing mode determination module 52 is used to obtain the feedback processing mode of the speech to be evaluated based on speech features when the first quality assessment information does not meet the preset quality assessment conditions; the feedback processing mode includes a first feedback processing mode and a second feedback processing mode.

[0152] The feedback result module 53 is used to process the first quality assessment information based on speech features and a preset compensation algorithm to obtain the corresponding second quality assessment information when the feedback processing mode is the first feedback processing mode, and to obtain the spoken pronunciation evaluation feedback result of the speech to be evaluated based on the second quality assessment information; it is also used to determine the defect information of the speech to be evaluated based on speech features and the reference text corresponding to the speech to be evaluated when the feedback processing mode is the second feedback processing mode, and to use the defect information as the spoken pronunciation evaluation feedback result of the speech to be evaluated.

[0153] In an exemplary embodiment, the processing mode determination module 52 is used to obtain the feedback processing mode of the speech to be evaluated based on speech features. Specifically, it is used to: determine the defect type of the speech to be evaluated and the degree of influence of the defect type on the first quality assessment information based on the speech features; the defect type is used to characterize the defect reason why the first quality assessment information of the speech to be evaluated does not meet the preset quality assessment conditions; if the defect type belongs to the preset type and the degree of influence is greater than the preset threshold, the feedback processing mode of the speech to be evaluated is determined to be the first feedback processing mode; if the defect type does not belong to the preset type and / or the degree of influence is less than or equal to the preset threshold, the feedback processing mode of the speech to be evaluated is determined to be the second feedback processing mode.

[0154] In an exemplary embodiment, the feedback result module 53 is used to obtain the spoken pronunciation evaluation feedback result of the speech to be evaluated based on the second quality assessment information, specifically: using the second quality assessment information and the defect type as the spoken pronunciation evaluation feedback result of the speech to be evaluated.

[0155] In an exemplary embodiment, the processing mode determination module 52 is used to determine the defect type of the speech to be evaluated and the degree of influence of the defect type on the first quality assessment information based on speech features. Specifically, it is used to: determine the phoneme-level Mel-Cepstral coefficients, phoneme-level posterior probabilities, and multidimensional sound features corresponding to the speech to be evaluated based on speech features; the multidimensional sound features are used to characterize at least one of the speech rate information, noise information, and volume information of the speech to be evaluated; and input the phoneme-level Mel-Cepstral coefficients, phoneme-level posterior probabilities, and multidimensional sound features into a pre-trained factor feedback model to obtain the defect type of the speech to be evaluated and the degree of influence of the defect type on the first quality assessment information.

[0156] In an exemplary embodiment, the first quality assessment information includes a first assessment value, and the second quality assessment information includes a second assessment value; the feedback result module 53 is used to process the first quality assessment information to obtain the corresponding second quality assessment information, which can be specifically used to: obtain the second assessment value corresponding to the first assessment value based on the first assessment value and the degree of influence value.

[0157] In an exemplary embodiment, the feedback result module 53 is used to determine the defect information of the speech to be evaluated based on the speech features and the reference text corresponding to the speech to be evaluated. Specifically, it is used to: obtain the speech phoneme sequence corresponding to the speech to be evaluated based on the speech features, and obtain the text phoneme sequence of the reference text corresponding to the speech to be evaluated; and determine the defect information of the speech to be evaluated based on the speech phoneme sequence and the text phoneme sequence.

[0158] In an exemplary embodiment, the processing mode determination module 52 is used to obtain a feedback processing mode for the speech to be evaluated based on speech features when the first quality assessment information does not meet the preset quality assessment conditions. Specifically, it is used to: obtain abnormal pronunciation detection results of the speech to be evaluated when the first quality assessment information does not meet the preset quality assessment conditions; the abnormal pronunciation detection results are used to characterize the degree of matching between the speech to be evaluated and the corresponding reference text; and when the degree of matching characterized by the abnormal pronunciation detection results is greater than or equal to a preset matching value, obtain a feedback processing mode for the speech to be evaluated based on speech features.

[0159] In an exemplary embodiment, after the abnormal pronunciation detection result of the speech to be evaluated is obtained by the processing mode determination module 52, the feedback result module 53 is used to take the first quality assessment information as the oral pronunciation evaluation feedback result of the speech to be evaluated when the abnormal pronunciation detection result characterization matching degree value is less than the preset matching value.

[0160] Each module in the aforementioned spoken pronunciation assessment and feedback device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0161] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 6As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When executed by the processor, the computer program implements a spoken pronunciation evaluation feedback method. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0162] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0163] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0164] Obtain the first quality assessment information of the speech to be evaluated, as well as the speech features of the speech to be evaluated;

[0165] If the first quality assessment information does not meet the preset quality assessment conditions, the feedback processing mode of the speech to be evaluated is obtained based on the speech features; the feedback processing mode includes a first feedback processing mode and a second feedback processing mode.

[0166] When the feedback processing mode is the first feedback processing mode, the first quality assessment information is processed based on speech features and a preset compensation algorithm to obtain the corresponding second quality assessment information, and the spoken pronunciation assessment feedback result of the speech to be evaluated is obtained based on the second quality assessment information.

[0167] When the feedback processing mode is the second feedback processing mode, the defect information of the speech to be evaluated is determined based on the speech features and the reference text corresponding to the speech to be evaluated, and the defect information is used as the speech pronunciation evaluation feedback result of the speech to be evaluated.

[0168] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0169] Obtain the first quality assessment information of the speech to be evaluated, as well as the speech features of the speech to be evaluated;

[0170] If the first quality assessment information does not meet the preset quality assessment conditions, the feedback processing mode of the speech to be evaluated is obtained based on the speech features; the feedback processing mode includes a first feedback processing mode and a second feedback processing mode.

[0171] When the feedback processing mode is the first feedback processing mode, the first quality assessment information is processed based on speech features and a preset compensation algorithm to obtain the corresponding second quality assessment information, and the spoken pronunciation assessment feedback result of the speech to be evaluated is obtained based on the second quality assessment information.

[0172] When the feedback processing mode is the second feedback processing mode, the defect information of the speech to be evaluated is determined based on the speech features and the reference text corresponding to the speech to be evaluated, and the defect information is used as the speech pronunciation evaluation feedback result of the speech to be evaluated.

[0173] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:

[0174] Obtain the first quality assessment information of the speech to be evaluated, as well as the speech features of the speech to be evaluated;

[0175] If the first quality assessment information does not meet the preset quality assessment conditions, the feedback processing mode of the speech to be evaluated is obtained based on the speech features; the feedback processing mode includes a first feedback processing mode and a second feedback processing mode.

[0176] When the feedback processing mode is the first feedback processing mode, the first quality assessment information is processed based on speech features and a preset compensation algorithm to obtain the corresponding second quality assessment information, and the spoken pronunciation assessment feedback result of the speech to be evaluated is obtained based on the second quality assessment information.

[0177] When the feedback processing mode is the second feedback processing mode, the defect information of the speech to be evaluated is determined based on the speech features and the reference text corresponding to the speech to be evaluated, and the defect information is used as the speech pronunciation evaluation feedback result of the speech to be evaluated.

[0178] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0179] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0180] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0181] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for assessing and providing feedback on spoken pronunciation, characterized in that, The method includes: Obtain the first quality assessment information of the speech to be evaluated, as well as the speech features of the speech to be evaluated; If the first quality assessment information does not meet the preset quality assessment conditions, the feedback processing mode of the speech to be evaluated is obtained based on the speech features; the feedback processing mode includes a first feedback processing mode and a second feedback processing mode. When the feedback processing mode is the first feedback processing mode, the first quality assessment information is processed based on the speech features and the preset compensation algorithm to obtain the corresponding second quality assessment information, and the spoken pronunciation assessment feedback result of the speech to be evaluated is obtained according to the second quality assessment information. When the feedback processing mode is the second feedback processing mode, the defect information of the speech to be evaluated is determined based on the speech features and the reference text corresponding to the speech to be evaluated, and the defect information is used as the speech pronunciation evaluation feedback result of the speech to be evaluated.

2. The method according to claim 1, characterized in that, The feedback processing mode for obtaining the speech to be evaluated based on the speech features includes: Based on the speech features, the defect type of the speech to be evaluated and the degree of influence of the defect type on the first quality assessment information are determined; the defect type is used to characterize the reason why the first quality assessment information of the speech to be evaluated fails to meet the preset quality assessment conditions. If the defect type belongs to a preset type and the impact value is greater than a preset threshold, the feedback processing mode of the speech to be evaluated is determined to be the first feedback processing mode. If the defect type does not belong to the preset type, and / or the impact value is less than or equal to the preset threshold, the feedback processing mode of the speech to be evaluated is determined to be the second feedback processing mode.

3. The method according to claim 2, characterized in that, The step of obtaining the spoken pronunciation evaluation feedback result of the speech to be evaluated based on the second quality assessment information includes: The second quality assessment information and the defect type are used as the oral pronunciation assessment feedback result of the speech to be evaluated.

4. The method according to claim 2, characterized in that, The step of determining the defect type of the speech to be evaluated based on the speech features, and the degree of influence of the defect type on the first quality assessment information, includes: Based on the speech features, the phoneme-level Mel-spectral coefficients, phoneme-level posterior probabilities, and multidimensional sound features corresponding to the speech to be evaluated are determined; the multidimensional sound features are used to characterize at least one of the speech rate information, noise information, and volume information of the speech to be evaluated. The phoneme-level Mel-spectral coefficients, the phoneme-level posterior probabilities, and the multidimensional sound features are input into a pre-trained factor feedback model to obtain the defect type of the speech to be evaluated and the degree of influence of the defect type on the first quality assessment information.

5. The method according to claim 4, characterized in that, The first quality assessment information includes a first assessment value, and the second quality assessment information includes a second assessment value; The process of processing the first quality assessment information to obtain the corresponding second quality assessment information includes: Based on the first evaluation value and the degree of influence value, a second evaluation value corresponding to the first evaluation value is obtained.

6. The method according to claim 1, characterized in that, The step of determining the defect information of the speech to be evaluated based on the speech features and the reference text corresponding to the speech to be evaluated includes: Based on the speech features, obtain the speech phoneme sequence corresponding to the speech to be evaluated, and obtain the text phoneme sequence of the reference text corresponding to the speech to be evaluated; Based on the speech phoneme sequence and the text phoneme sequence, the defect information of the speech to be evaluated is determined.

7. The method according to claim 1, characterized in that, When the first quality assessment information does not meet the preset quality assessment conditions, the step of obtaining the feedback processing mode of the speech to be evaluated based on the speech features includes: If the first quality assessment information does not meet the preset quality assessment conditions, the abnormal pronunciation detection result of the speech to be evaluated is obtained; the abnormal pronunciation detection result is used to characterize the degree of matching between the speech to be evaluated and the corresponding reference text. When the abnormal pronunciation detection result indicates that the matching degree value is greater than or equal to a preset matching value, the feedback processing mode of the speech to be evaluated is obtained based on the speech features.

8. The method according to claim 7, characterized in that, After obtaining the abnormal pronunciation detection results of the speech to be evaluated, the method further includes: If the abnormal pronunciation detection result indicates that the matching degree value is less than the preset matching value, the first quality assessment information is used as the oral pronunciation evaluation feedback result of the speech to be evaluated.

9. A spoken pronunciation assessment and feedback device, characterized in that, The device includes: The data acquisition module is used to acquire the first quality assessment information of the speech to be evaluated, as well as the speech features of the speech to be evaluated; The processing mode determination module is used to obtain the feedback processing mode of the speech to be evaluated based on the speech features when the first quality assessment information does not meet the preset quality assessment conditions; the feedback processing mode includes a first feedback processing mode and a second feedback processing mode. The feedback result module is used to process the first quality assessment information based on the speech features and a preset compensation algorithm to obtain the corresponding second quality assessment information when the feedback processing mode is the first feedback processing mode, and to obtain the spoken pronunciation evaluation feedback result of the speech to be evaluated based on the second quality assessment information; it is also used to determine the defect information of the speech to be evaluated based on the speech features and the reference text corresponding to the speech to be evaluated when the feedback processing mode is the second feedback processing mode, and to use the defect information as the spoken pronunciation evaluation feedback result of the speech to be evaluated.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.