Machine-implemented oral training method, device, and readable storage medium
The oral training method implemented by machines, using training modes of different difficulty levels and real-time evaluation, solves the problems of large human and financial investment and content incompatibility in existing technologies, and achieves personalized and flexible oral training effects.
Patent Information
- Application Number
- CN202111516695.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-06
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2041-12-06
AI Technical Summary
Existing oral training methods require a lot of manpower and financial resources, are subject to time and location constraints, and the content provided by existing APPs is not colloquial and systematic enough, making it difficult to match the oral proficiency of different users, making it difficult for users to master oral skills in actual scenarios.
A machine-implemented oral training method is provided. By setting multiple oral training modes, personalized training is carried out based on the different difficulty levels of the same conversation content and the user's oral proficiency, including follow-up, challenge and difficulty training modes, and real-time evaluation and feedback are carried out using natural language processing and automatic speech recognition technology.
It realizes step-by-step training for users on the same conversation content, improves the effect of oral training and users' self-confidence, adapts to the learning needs of different users, reduces human and financial investment, and improves the flexibility and systematicness of oral training.
Smart Images

Figure CN114255759B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of data processing technology. More specifically, the embodiments of the present invention relate to a machine-implemented spoken language training method, a device for implementing spoken language training, and a computer-readable storage medium. Background Art
[0002] This section is intended to provide background or context for embodiments of the present invention as recited in the claims. The description herein may include concepts that could be explored, but not necessarily concepts that have been previously conceived or explored. Therefore, unless otherwise indicated herein, the material described in this section is not prior art with respect to the specification and claims of this application and is not admitted to be prior art by inclusion in this section.
[0003] Currently, there are two main types of oral training methods: real-person oral instruction and imitative machine oral learning. Real-person oral instruction typically involves face-to-face instruction between a teacher and a student. Because real-person oral instruction allows for real-time conversation and feedback, conversational training is relatively flexible and more suitable for practical oral application scenarios. Existing imitative machine oral learning is typically implemented using oral learning applications (APPs), which provide oral practice sentences and allow users to imitate them for oral training. Summary of the Invention
[0004] However, live oral instruction requires significant human and financial resources, and is subject to constraints such as time and location. Each student has limited time to practice speaking, making it difficult to guarantee the effectiveness of oral training. Existing oral learning apps often lack colloquial and systematic content, and struggle to match the spoken language proficiency of different users, making it difficult for users to persist in oral learning. Furthermore, even with trained oral content, users often find it difficult to truly master and apply it in actual exams or oral application scenarios, failing to truly help users escape the dilemma of "dumb English."
[0005] Therefore, there is a great need for an improved oral training method that can not only reduce the investment of manpower and financial resources, but also provide a personalized oral training method that conforms to oral application scenarios and is suitable for user difficulty.
[0006] In this context, embodiments of the present invention are intended to provide a machine-implemented spoken language training method, a device for implementing spoken language training, and a computer-readable storage medium.
[0007] In a first aspect of an embodiment of the present invention, a machine-implemented oral training method is provided, comprising: setting a plurality of oral training modes with different levels of difficulty based on the same conversation content; and entering a corresponding oral training mode based on the difficulty of the plurality of oral training modes or a user selection.
[0008] In one embodiment of the present invention, the oral training method further includes: grading the difficulty of each candidate conversation content among multiple candidate conversation contents in terms of content difficulty; and determining the range of candidate conversation contents that the user can select based on the difficulty level corresponding to the user's oral proficiency, so that the user can select the conversation content within the range.
[0009] In another embodiment of the present invention, the content difficulty is determined based on at least one of the following: the topic of the candidate conversation content; the vocabulary of the candidate conversation content; the grammar of the candidate conversation content; and the sentence length of the candidate conversation content.
[0010] In another embodiment of the present invention, entering a corresponding oral training mode based on the difficulty levels of multiple oral training modes includes: entering the corresponding oral training modes in order from easy to difficult according to the multiple oral training modes; and determining whether to enter the next oral training mode based on the user's overall evaluation results in the current oral training mode.
[0011] In another embodiment of the present invention, the oral training mode includes a follow-up training mode, and the oral training method further includes: in response to entering the follow-up training mode, determining the first role of the user in the conversation content and the target sentence corresponding to the first role; outputting the conversation content, and when outputting the target sentence of the first role, receiving a first voice of the user following the target sentence; and based on the target sentence, performing an oral evaluation on the first voice to determine whether to output the conversation content for the next round.
[0012] In one embodiment of the present invention, the oral training method further includes: in response to the end of each round of conversation for the first role, determining the second role of the user in the conversation content that is different from the first role to continue the shadowing training; and in response to the end of each round of conversation for each role in the conversation content, determining the overall evaluation result based on the oral evaluation results of each round of conversation.
[0013] In another embodiment of the present invention, the oral training mode includes a challenge training mode, and the oral training method further includes: in response to entering the challenge training mode, outputting questions in the conversation content and outputting a first answer prompt corresponding to the question; receiving a second voice in which the user responds based on the first answer prompt; and based on the first answer prompt, performing an oral evaluation on the second voice to determine whether to output the conversation content for the next round.
[0014] In another embodiment of the present invention, determining whether to output the content of the next round of conversation includes: in response to the oral evaluation result of the second voice being higher than or equal to a first threshold, outputting questions for the next round of conversation; or in response to the oral evaluation result of the second voice being lower than the first threshold, classifying the second voice, and performing a corresponding first operation based on a first category obtained from the classification.
[0015] In another embodiment of the present invention, the first category includes one or more of the following: semantic irrelevance, inaccurate pronunciation, and incomplete answer; and the corresponding first operation includes: when the first category is semantic irrelevance, determining a second answer prompt with different degrees of completeness based on the number of semantic irrelevances in the current round of conversation; when the first category is inaccurate pronunciation, outputting pronunciation prompt information for prompting re-pronunciation; and / or when the first category is incomplete answer, outputting a third answer prompt regarding the incomplete part.
[0016] In one embodiment of the present invention, the oral training mode includes a difficult training mode, and the oral training method further includes: in response to entering the difficult training mode, outputting questions in the conversation content; receiving a third voice in which the user responds to the question; and performing an oral evaluation on the third voice to determine whether to output the conversation content for the next round.
[0017] In another embodiment of the present invention, determining whether to output the content of the next round of conversation includes: in response to the oral evaluation result of the third voice being higher than or equal to a second threshold, outputting questions for the next round of conversation; or in response to the oral evaluation result of the third voice being lower than the second threshold, classifying the third voice, and performing a corresponding second operation based on a second category obtained from the classification.
[0018] In another embodiment of the present invention, a corresponding second operation is performed based on the second category, including: when the second category is semantically irrelevant, the corresponding second operation includes repeatedly outputting the question; and / or when the second category is a category other than the semantically irrelevant category, the corresponding second operation includes outputting recommendation information related to the question.
[0019] In another embodiment of the present invention, before outputting questions for the next round of conversation, the oral training method further includes: in response to receiving other questions other than the conversation content, performing any one of the following operations: skipping the current round of conversation; or outputting response information related to the other questions.
[0020] In a second aspect of an embodiment of the present invention, a device for implementing oral training is provided, comprising: a processor configured to execute program instructions; and a memory configured to store the program instructions, wherein when the program instructions are loaded and executed by the processor, the device performs the oral training method according to any one of the first aspects of the embodiment of the present invention.
[0021] In a third aspect of the embodiments of the present invention, a computer-readable storage medium is provided, in which program instructions are stored. When the program instructions are loaded and executed by a processor, the processor executes the oral training method according to any one of the first aspects of the embodiments of the present invention.
[0022] According to the oral training method implemented by a machine in an embodiment of the present invention, multiple oral training modes with different levels of difficulty can be set based on the same conversation content, thereby providing a step-by-step training mode in the training method, which is conducive to users gradually conducting multiple trainings of different levels of difficulty for the same conversation, thereby helping users to truly master the training content and bringing users a better experience and training effect.
[0023] Furthermore, in some embodiments, the user's spoken language proficiency can be matched based on the content difficulty of each candidate conversation content among multiple candidate conversation contents, thereby implementing a spoken language training method that combines content difficulty with training mode difficulty. In other embodiments, by setting a challenge training mode and outputting a first response prompt corresponding to the question in the challenge training mode to guide the user to practice answering under the first response prompt, this can help improve the user's answer accuracy while guiding the user's answer, thereby increasing the user's confidence and sense of accomplishment in spoken language training, and helping the user to challenge and train in more difficult training modes. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily apparent by reading the following detailed description with reference to the accompanying drawings, in which several embodiments of the present invention are shown by way of example and not limitation, in which:
[0025] Figure 1 A block diagram schematically illustrates an exemplary computing system 100 suitable for implementing embodiments of the present invention;
[0026] Figure 2 Schematically shows a flow chart of a spoken language training method according to an embodiment of the present invention;
[0027] Figure 3 Schematically shows a flow chart of a spoken language training method including content difficulty grading according to another embodiment of the present invention;
[0028] Figure 4 A flowchart of a spoken language training method for entering a follow-up reading training mode according to an embodiment of the present invention is schematically shown;
[0029] Figure 5 A flowchart of a spoken language training method for entering a challenge training mode according to an embodiment of the present invention is schematically shown;
[0030] Figure 6 Schematically illustrates a conversation flow chart including outputting a second answer prompt according to an embodiment of the present invention; and
[0031] Figure 7 The flowchart of the oral training method for entering the difficult training mode according to an embodiment of the present invention is schematically shown.
[0032] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts. DETAILED DESCRIPTION
[0033] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided solely to enable those skilled in the art to better understand and implement the present invention, and are not intended to limit the scope of the present invention in any way. Rather, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.
[0034] Figure 1 Schematically illustrates a block diagram of an exemplary computing system 100 suitable for implementing embodiments of the present invention. Figure 1As shown, the computing system 100 may include a central processing unit (CPU) 101, a random access memory (RAM) 102, a read-only memory (ROM) 103, a system bus 104, a hard disk controller 105, a keyboard controller 106, a serial interface controller 107, a parallel interface controller 108, a display controller 109, a hard disk 110, a keyboard 111, a serial peripheral device 112, a parallel peripheral device 113, and a display 114. Of these devices, the CPU 101, RAM 102, ROM 103, hard disk controller 105, keyboard controller 106, serial controller 107, parallel controller 108, and display controller 109 are coupled to the system bus 104. The hard disk 110 is coupled to the hard disk controller 105, the keyboard 111 is coupled to the keyboard controller 106, the serial peripheral device 112 is coupled to the serial interface controller 107, the parallel peripheral device 113 is coupled to the parallel interface controller 108, and the display 114 is coupled to the display controller 109. It should be understood that Figure 1 The structured block diagram is only for the purpose of illustration, rather than for limiting the scope of the present invention. In some cases, some devices may be added or reduced according to specific circumstances.
[0035] Those skilled in the art will appreciate that embodiments of the present invention may be implemented as a system, method, or computer program product. Accordingly, the present invention may be implemented in the following forms: entirely in hardware, entirely in software (including firmware, resident software, microcode, etc.), or in a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, the present invention may be implemented in the form of a computer program product embodied in one or more computer-readable media containing computer-readable program code.
[0036] Any combination of one or more computer-readable media can be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (non-exhaustive examples) of computer-readable storage media can include, for example: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device.
[0037] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0038] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0039] The computer program code for performing the operations of the present invention can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0040] The following will describe the embodiments of the present invention with reference to the flowcharts of the methods and block diagrams of the devices (or apparatuses or systems) according to the embodiments of the present invention. It should be understood that each block in the flowcharts and / or block diagrams, as well as the combination of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, and these computer program instructions are executed by the computer or other programmable data processing device to produce a device that implements the functions / operations specified in the blocks in the flowcharts and / or block diagrams.
[0041] These computer program instructions can also be stored in a computer-readable medium that enables a computer or other programmable data processing device to operate in a specific manner. In this way, the instructions stored in the computer-readable medium produce a product that includes an instruction device that implements the functions / operations specified in the blocks in the flowchart and / or block diagram.
[0042] Computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer or other programmable apparatus provide a process that implements the functions / operations specified in the blocks in the flowchart and / or block diagram.
[0043] According to an embodiment of the present invention, a machine-implemented spoken language training method, a device for implementing spoken language training, and a computer-readable storage medium are proposed.
[0044] In this article, it is to be understood that the terms involved include the following:
[0045] NLP: Natural language processing, natural language processing technology, which studies various theories and methods that can achieve effective communication between humans and computers using natural language, and is mainly used in machine translation, public opinion monitoring, automatic summarization, opinion extraction, text classification, question answering, text semantic comparison, speech recognition, text recognition (OCR), etc.
[0046] ASR: Automatic Speech Recognition, automatic speech recognition technology, can convert speech into text.
[0047] CAPT: Computer Aided Pronunciation Training, machine-assisted pronunciation guidance, allows the machine to evaluate and score based on the text provided by the user and the pronunciation of the text.
[0048] Second Language Acquisition: Second Language Acquisition, referred to as SLA or second language acquisition, usually refers to any other language learning after mother tongue acquisition.
[0049] Keyword extraction technology: also known as keyword extraction technology, is a technology that can automatically extract key meaning groups, keywords and / or keyword groups that reflect the text.
[0050] A phrase refers to the components of a sentence divided by meaning and structure. Each component is called a phrase, and the words in a phrase are closely related. A phrase can be a meaningful chunk or a block that summarizes the key points of a sentence.
[0051] Keywords: refers to words or phrases that can reflect the theme or core idea of the text, or can be understood as words that have practical meaning or can summarize the key points of a sentence.
[0052] In addition, any number of elements in the drawings is for illustration and not limitation, and any naming is for distinction only and does not have any limiting meaning.
[0053] The principles and spirit of the present invention are explained in detail below with reference to several representative embodiments of the present invention. SUMMARY OF THE INVENTION
[0055] The inventors have discovered that in actual communication, second language learners already possess a high level of proficiency in their native language and rely to a certain extent on native language thinking. Therefore, second language learners are more likely to seek translation when speaking. Secondly, most people can understand complex English novels, but they don't know how to start extremely simple daily conversations. For example, a Level 6 English user tested by the inventors was stunned when he heard common greetings like "what's up" and didn't know how to respond for a while. Some second language learners, while able to express themselves in spoken language and their expressions are semantically sound, have higher expectations for themselves, such as pronunciation accuracy and grammatical details. Therefore, they prefer to have a real-time error correction partner for practice.
[0056] In response to the above case, the inventors also found that the use of artificial intelligence (AI) technology can help users use the most suitable dialogue content and practice methods for themselves, and conduct repeated training for the user's personalized problems and weaknesses. For example, for words with inaccurate pronunciation, CAPT automatic evaluation technology can be used to tell the user which specific inaccurate words and phonemes are, as well as the correct pronunciation method. Furthermore, for the same question, different users can have a variety of ways to answer due to differences in age and scenario. In other words, oral dialogue should not be a rigid and single text, so the inventors considered using ASR technology and NLP technology to analyze what the user answered and whether the semantics of the answer sentence are consistent with the current scenario.
[0057] After introducing the basic principles of the present invention, various non-limiting embodiments of the present invention are described in detail below.
[0058] Application Scenario Overview
[0059] The oral training method of the embodiment of the present invention can be implemented by an application running on a machine. Such an application can be, for example, a language training APP, in particular a spoken language training APP. The language can be various existing languages, including but not limited to English, French, German, Spanish, Korean, Japanese, Chinese, etc. The user group can be, for example, second language learners. The user group can also be adults, teenagers, young children, etc. Usually, in such a language training APP, oral training can be performed on the user based on the oral training content selected by the user or set by the system. In other application scenarios, the oral training content set by the system can be selected based on the oral level that matches the user's previous oral training results. Furthermore, a speaker can usually be set on the machine that implements the language training APP to play the spoken training content, and / or a recording device can be set to receive the user's response voice, etc.
[0060] Exemplary Methods
[0061] In combination with the above application scenarios, Figure 2 The following describes a machine-implemented spoken language training method according to an exemplary embodiment of the present invention. It should be noted that the above application scenarios are merely provided to facilitate understanding of the spirit and principles of the present invention, and the embodiments of the present invention are not limited in this respect. Rather, the embodiments of the present invention can be applied to any applicable scenario.
[0062] like Figure 2 As shown in , the oral training method 200 may include: in step 210, based on the same conversation content, setting multiple oral training modes with different levels of difficulty. Setting multiple oral training modes based on the same conversation content can be understood as repeatedly performing oral training on the same conversation content through multiple different training methods. The conversation content may include questions and possible responses (or reference responses) associated therewith. In some embodiments, the conversation content may include one round of conversation, that is, one question and one or more corresponding reference responses. In other embodiments, the conversation content may include multiple rounds of conversation (or dialogues), that is, multiple questions and at least one possible response associated with each question. In yet other embodiments, the conversation content may be determined by the user selecting from multiple candidate conversation contents provided by the machine, or may be determined by the machine by obtaining the user's oral proficiency.
[0063] The conversation content described above can be stored in a variety of forms. For example, in one embodiment of the present invention, the conversation content described above can be pre-stored in a machine or other accessible media in the form of text, and when the conversation content needs to be output, the text can be converted into voice and output to the user. The above operations can be performed using existing technologies such as text-to-speech (TTS) technology, or various text-to-speech technologies developed in the future. In another embodiment of the present invention, the conversation content described above can be pre-stored in a machine or other accessible media in the form of voice, and when the conversation content needs to be output, the stored voice can be directly output.
[0064] In yet other embodiments, the spoken language training method 200 may further include determining, based on the user-selected conversation scenario, conversation content related to the scenario. The machine may provide multiple conversation scenarios for the user to choose from, such as ordering food at a restaurant, shopping at a supermarket, greeting someone, asking for directions, etc. The user may select the conversation scenario they wish to train in, and the machine may determine the characters and relevant conversation content in the conversation scenario based on the conversation scenario.
[0065] The multiple oral training modes described above can be set to have different training difficulties in training methods, so that users can conduct multi-dimensional training of different difficulty levels for the same conversation content, so that users can more comprehensively master and apply the conversation content. Compared with a single training mode, or using different training modes for different conversation contents, according to an embodiment of the present invention, oral training modes of different difficulty levels are used for the same conversation content, so that users can conduct step-by-step training and learning for the same conversation content, which is conducive to improving the user's oral training effect. In some embodiments, the multiple oral training modes may include at least two of the following training modes, challenge training modes, and difficult training modes with increasing difficulty. The specific implementation method of each mode will be described in detail below. Figure 4-Figure 7 A detailed description is given and will not be repeated here.
[0066] Next, in step 220, the corresponding speaking training mode can be entered based on the difficulty level of the multiple speaking training modes or user selection. In some embodiments, the machine can provide multiple different speaking training modes for the same conversation content, and the user can select at least one of the speaking training modes to train based on personal needs. In other embodiments, entering the corresponding speaking training mode based on the difficulty level of the multiple speaking training modes can be automatically determined and implemented by the machine. Specific implementation methods may include, for example, steps 221 and 222, which will be described in detail below.
[0067] like Figure 2As further shown in FIG, in step 221 (shown by the dotted box), the corresponding oral training mode can be entered in sequence according to the order of the multiple oral training modes from easy to difficult. The order from easy to difficult can be understood as the order from easy to difficult in terms of training methods.
[0068] Further, in step 222, it is possible to determine whether to enter the next spoken language training mode based on the total evaluation result of the user under the current spoken language training mode. The difficulty of the next spoken language training mode may be greater than the current spoken language training mode. Under the current spoken language training mode, the spoken language training effect of the user is evaluated orally, and the various spoken language evaluation technologies of existing (such as CAPT pronunciation evaluation technology, spoken language evaluation model, etc.) or future developments may be used to perform the above operations. In certain embodiments, the total evaluation result may include the comprehensive results of the spoken language evaluation of multiple rounds of conversations under the current spoken language training mode. In certain embodiments, the total evaluation result may include the comprehensive results of the multi-dimensional (such as pronunciation, fluency, semantics, completeness, etc.) spoken language evaluation of the conversation content under the current spoken language training mode.
[0069] In some embodiments, in response to the user's total evaluation result in the current speaking training mode being greater than or equal to a preset threshold, the user may be directed to enter the next speaking training mode with a higher difficulty level. In other embodiments, in response to the user's total evaluation result in the current speaking training mode being less than a preset threshold, the user may be directed to return to the current speaking training mode and resume training.
[0070] Combination of the above Figure 2 In general, an exemplary description of the oral training method implemented by a machine according to an embodiment of the present invention has been given. It will be understood by those skilled in the art that the above description is exemplary and not restrictive. For example, in step 220, entering the corresponding oral training mode based on the difficulty of multiple oral training modes may not be limited to the order from easy to difficult in step 221, but may also be in an alternating order of easy and difficult as needed. In other application scenarios, when it is necessary to test the user's oral proficiency, the corresponding oral training mode may also be entered in an order from difficult to easy. When the user can successfully pass the oral training mode with a higher difficulty, the user's oral proficiency can be quickly determined. For example, the oral training method according to an embodiment of the present invention may not be limited to setting different difficulty levels only in the training method, but may also match the user's training difficulty in combination with the content difficulty of the conversation content. This will be described below in combination with Figure 3 An exemplary description is given.
[0071] Figure 3 The flowchart of the oral training method including content difficulty grading according to another embodiment of the present invention is schematically shown. Figure 3As shown in , the spoken language training method 300 may include: in step 310, each candidate conversation content in the plurality of candidate conversation content may be graded in terms of content difficulty. In some embodiments, step 310 may further include: using natural language processing (NLP) technology and keyword extraction technology to grade the content difficulty of each candidate conversation content in the plurality of candidate conversation content.
[0072] Specifically, in another embodiment of the present invention, content difficulty can be determined based on at least one of the following: the topic of the candidate conversation content; the vocabulary of the candidate conversation content; the grammar of the candidate conversation content; and the sentence length of the candidate conversation content. The topic of the candidate conversation content can be extracted using natural language processing (NLP) technology, the vocabulary of the candidate conversation content can be extracted using keyword extraction technology, the grammar of the candidate conversation content can be analyzed using natural language processing (NLP) technology, and the sentence length of the candidate conversation content can also be analyzed using natural language processing (NLP) technology. In some embodiments, the vocabulary of the candidate conversation content can include key meaning groups, keywords, and / or keyword phrases.
[0073] The difficulty grading described above can be achieved by comprehensively evaluating the topic difficulty, vocabulary difficulty, grammatical difficulty, and / or sentence length of each candidate conversation content to achieve a content difficulty grading for each candidate conversation content. In some embodiments, the corresponding difficulty can be determined by matching the topic, vocabulary, grammar, and / or sentence length of the conversation content with different levels of language proficiency standards. For example, keywords extracted using keyword extraction technology can be matched with vocabulary lists from standards for IELTS, CET-4, CET-6, high school English, and junior high school English to determine the difficulty level of the keywords, and thus the difficulty level of the conversation content. In other embodiments, the difficulty of multiple candidate conversation contents can be ranked based on the topic difficulty, vocabulary difficulty, grammatical difficulty, and / or sentence length of each candidate conversation content, and then divided into multiple difficulty levels, such as Level 1 to Level 5, in order of difficulty.
[0074] Next, in step 320, a range of candidate conversation content selectable by the user may be determined based on the difficulty level corresponding to the user's spoken language proficiency, so that the user can select conversation content within this range. For example, in some application scenarios, if the user's spoken language proficiency is at a high school level, the range of candidate conversation content selectable by the user may be determined to be at the high school level, so that the user can select the conversation content they want to practice speaking from multiple candidate conversation content within the high school level range, without recommending candidate conversation content of a higher difficulty level (e.g., IELTS) or a lower difficulty level (e.g., elementary school level) to the user.
[0075] In some embodiments, the user's speaking level can be set by the user themselves, or it can be determined by a machine based on a comprehensive assessment of the user's speaking training history. For example, in other embodiments, the speaking training method 300 can further include: determining whether to output candidate conversation content for the next content difficulty level based on the user's overall assessment results of speaking training at the current content difficulty level. For example, in other application scenarios, if the user's overall assessment results of speaking training for conversation content at level one meet a preset standard, it can be determined that the user's speaking level has exceeded level one, and candidate conversation content within level two can be presented for the user to select. The overall assessment results here can include the overall results of the user completing all speaking training modes at the current content difficulty level, or can include the overall results of some speaking training modes. For example, in some embodiments, the speaking training method 300 can further include: determining whether to output candidate conversation content for the next content difficulty level based on the overall assessment results of the user's speaking training at the current content difficulty level in the current speaking training mode.
[0076] Then, the process can proceed to step 330, where multiple oral training modes with different levels of difficulty can be set based on the same conversation content set by the machine or selected by the user. Further, in step 340, the corresponding oral training mode can be entered based on the difficulty level of the multiple oral training modes or the user's selection. It can be understood that steps 330 and 340 have been combined in the previous text. Figure 2 Steps 210 and 220 are described in detail and will not be repeated here.
[0077] Combination of the above Figure 3 An exemplary description is given of the oral training method including content difficulty grading according to an embodiment of the present invention. It can be understood that according to the oral training method of this embodiment, both the content difficulty of the conversation content and the difficulty of the training method can be graded to better match different user types and different training needs of users, and a systematic oral training method can be provided by combining content difficulty and training method difficulty from easy to difficult.
[0078] It is also necessary to understand that by combining the content difficulty of the conversation content and the difficulty of the training method, it can be achieved that each conversation content can correspond to multiple oral training modes of different difficulty levels, so that users can conduct advanced repeated training of different difficulty levels on the same conversation content; it can also be achieved that for each difficulty level of oral training mode, multiple candidate conversation content ranges of different content difficulty levels can be provided, so that users can conduct oral training of different content difficulty levels in each oral training mode. According to such a setting, the learning methods and user bases of different users can be taken into account, making oral training easier to use, which is conducive to improving the flexibility and diversity of users' oral training, and is also conducive to setting personalized oral training according to user characteristics. The following will be combined with Figure 4-Figure 7 Examples are given of multiple oral training modes with different levels of difficulty.
[0079] Figure 4 The flowchart of the oral training method for entering the follow-up training mode according to an embodiment of the present invention is schematically shown. It should be understood that the follow-up training mode can be one of multiple oral training modes, and the method 400 for entering the follow-up training mode for oral training can be a specific manifestation of the oral training method 200 or the oral training method 300. Therefore, the above description of Figure 2 and Figure 3 The description can also be applied to the following Figure 4 's description.
[0080] like Figure 4 As shown in , method 400 may include: in step 410, in response to entering the follow-up training mode, the first role of the user in the conversation content and the target sentence corresponding to the first role may be determined. In some embodiments, the conversation content may include multiple roles for dialogue, and each role has words that need to be expressed (i.e., corresponding target sentences) to form complete and authentic conversation content. For example, in some application scenarios, the conversation content may include a conversation between a customer and a salesperson. In other application scenarios, the conversation content may include a conversation between a father, a mother, a grandfather, a grandmother, and a child. The target sentence for each role may include at least one type of question and answer.
[0081] In some embodiments, the first role may be determined based on user selection or randomly assigned by a machine. In other embodiments, the target sentence for each role may be stored in text form (e.g., stored in a memory), and when output is required, text-to-speech technology may be used to convert the target text of the target sentence into a corresponding target speech for output.
[0082] Next, in step 420, the conversation content may be output, and when the target sentence of the first character is output, a first voice message from the user following the target sentence may be received. Specifically, after the first character to be played by the user is determined, the conversational voice message of the conversation content may be output, and when the target sentence of the first character is output, a first voice message from the user imitating the target sentence may be received.
[0083] Then, the process can proceed to step 430, and the first speech can be evaluated for oral expression based on the target sentence to determine whether to output the content of the next round of conversation. The oral evaluation can include at least one of pronunciation, fluency, completeness, error rate, etc. The pronunciation evaluation can include evaluating the pronunciation accuracy of each sentence and each word in the sentence. The fluency evaluation can include evaluating whether the overall oral expression of the first speech is stuck. The completeness evaluation can include evaluating whether there are unpronounced words (or missing words) in the first speech. The error rate evaluation can include evaluating whether there are grammatical errors, word usage errors, etc. in the first speech.
[0084] In some embodiments, oral evaluation can be achieved by comparing the received first speech with the target sentence and using speech scoring technology. Comparing the first speech with the target sentence may include at least one of the following: comparing the digital representation of the first speech with the digital representation of the target sentence; and comparing the first text converted from the first speech with the target text of the target sentence. The conversion of the first speech into the first text can be achieved by adopting existing technology such as ASR or various speech-to-text technologies to be developed in the future. In other embodiments, method 400 may also include: the oral evaluation results of the first speech can be presented through, for example, a human-computer interaction interface.
[0085] In yet other embodiments, method 400 may further include: in response to the spoken language evaluation result of the first speech being greater than or equal to a third threshold, outputting the next round of conversation content, wherein the next round of conversation content may include the next round of target sentences; and in response to the spoken language evaluation result of the first speech being less than the third threshold, repeatedly outputting the target sentences of the current round. In some embodiments, the user may also choose to repeat the target sentences of the current round or proceed to the next round of conversation content based on the presented spoken language evaluation results.
[0086] According to this setting, in the follow-up training mode, oral evaluation can be performed on each sentence read by the user and the oral evaluation results can be displayed, so that the user can understand the dialogue scenario and imitate the original sound to train the sense of language in the process of imitation learning. In addition, the user can grasp his or her own oral expression in real time and conduct targeted repeated training of single sentences, so that the user can develop good oral expression habits and lay a solid foundation for oral expression in the early stages of oral training.
[0087] like Figure 4 As further shown in, optionally or additionally, method 400 may further include: in step 440 (shown in a dotted box), in response to the end of each round of conversation for the first role, determining the second role of the user in the conversation content that is different from the first role, so as to continue the follow-up training. This step can also be understood as a human-computer role swap, that is, after the user completes all the target sentences of the first role in the follow-up training, he can select other roles in the conversation content to continue the follow-up training, and can follow each target sentence of the current role according to the process of steps 410-430 above. The process is still implemented in the form of human-computer conversation until all conversations for the current role are completed. In another embodiment, if the conversation content includes at least three roles, the user can continue to select the third role for follow-up training after completing the conversation of two of the roles.
[0088] Alternatively or additionally, in step 450 (shown in a dotted box), in response to the completion of each round of conversation for each character in the conversation content, a total evaluation result is determined based on the oral evaluation results of each round of conversation. In some embodiments, the conversation content may include multiple characters and multiple rounds of conversation, wherein each round of conversation may include dialogue statements of all or part of the multiple characters. The completion of each round of conversation for each character in the conversation content can be understood as the user performing follow-up training for all characters in the conversation content, and the follow-up training for all rounds of conversation for each character in the conversation content has ended.
[0089] In other embodiments, the overall evaluation result can be a comprehensive score of the oral evaluation results of each round of conversation. In still other embodiments, the oral training method 400 can further include determining whether to enter the next oral training mode with a higher difficulty level based on the overall evaluation result of the user in the follow-up training mode. In yet another embodiment, the oral training method can further include determining whether to enter the next conversation content with a higher difficulty level in the follow-up training mode based on the overall evaluation result of the user's current conversation content in the follow-up training mode.
[0090] In one embodiment, the user can also choose to retrain the current conversation content in the follow-up training mode or choose to enter the next more difficult oral training mode based on the overall evaluation results presented. In another embodiment, the user can also choose to retrain the current conversation content in the follow-up training mode or choose to enter the next more difficult conversation content based on the overall evaluation results presented.
[0091] Combination of the above Figure 4 An exemplary description of the oral training method in the shadowing training mode according to an embodiment of the present invention has been provided. It will be understood by those skilled in the art that the above description is exemplary and not restrictive. For example, steps 440 and 450 are exemplary and may not be performed in actual applications as needed. For another example, the first role may not be limited to one role in the conversation content. In another embodiment, the first role may be used to represent multiple roles in the conversation content, i.e., the user may choose to shadow target sentences of multiple roles in the conversation content simultaneously, thereby improving the training efficiency of a single shadowing training session.
[0092] Figure 5 The flowchart of the oral training method for entering the challenge training mode according to an embodiment of the present invention is schematically shown. It should be understood that the challenge training mode can be one of multiple oral training modes, and the method 500 for entering the challenge training mode for oral training can be a specific manifestation of the oral training method 200 or the oral training method 300. Therefore, the above description of the challenge training mode can be omitted. Figure 2 and Figure 3 The description can also be applied to the following Figure 5 's description.
[0093] like Figure 5 As shown, method 500 may include: in step 510, in response to entering the challenge training mode, outputting the question in the conversation content and outputting the first answer prompt corresponding to the question. The purpose of setting the challenge training mode is to guide the user to speak the conversation content independently and to try to organize the language to express it. In some embodiments, in the challenge training mode, the question in the conversation content can be output in the form of voice, and at least one recommended answer (or reference answer) corresponding to the question in the conversation content can be stored in the form of text. The first answer prompt can be output before the user answers, so as to prompt the user to start answering and prompt the content of the answer.
[0094] In other embodiments, step 510 may further include: generating a first response prompt using NLP technology. In some embodiments, the first response prompt may include the overall sentence meaning of the reference response. In yet other embodiments, the first response prompt may be output in the form of the user's native language. For example, if the user's native language is Chinese, when the user is performing oral English training, the machine first outputs the English voice of the question "How are you?", and then the machine may continue to output the first response prompt in Chinese "Answer prompt: I'm doing OK."
[0095] In some further embodiments, the first response prompt may be output in an audible and / or visual form. For example, the first response prompt "Response prompt: I'm doing OK" may be directly presented in text form, or the first response prompt "Response prompt: I'm doing OK" may be converted into speech using text-to-speech technology and output.
[0096] Next, in step 520, a second voice message may be received in which the user responds based on the first response prompt. In some embodiments, step 520 may further include determining whether the received second voice message is a response based on the first response prompt. For example, this determination may be based on the time interval between the received second voice message and the output of the first response prompt.
[0097] Then, the process can proceed to step 530, and based on the first answer prompt, the second voice can be evaluated for oral language to determine whether to output the content of the next round of conversation. The oral language evaluation can include at least one of semantic relevance, pronunciation, fluency, completeness, error rate, etc. The semantic relevance evaluation can include a semantic relevance score between the second voice and the corresponding question, which can be implemented using existing or future semantic analysis technology. The specific contents of pronunciation evaluation, fluency evaluation, completeness evaluation and error rate evaluation are combined with the above. Figure 4 The oral evaluation in step 430 is the same or similar and will not be described again here.
[0098] In some embodiments, oral evaluation can be performed by comparing the received second speech with a reference response and utilizing speech scoring technology. Comparing the second speech with the reference response can include at least one of: comparing a digital representation of the second speech with a digital representation of the reference response; and comparing a second text converted from the second speech with the reference response. Converting the second speech to the second text can be performed using speech-to-text technology.
[0099] In other embodiments, method 500 may further include presenting a spoken language evaluation result of the second voice through, for example, a human-computer interaction interface. In still other embodiments, the user may choose to retake the current round of questions or proceed to the next round of conversation based on the presented spoken language evaluation result.
[0100] According to one embodiment of the present invention, Figure 5 As further shown in FIG, step 530 may include step 531 (shown in a dotted box) or step 532 (shown in a dotted box). Specifically, in step 531, in response to the oral evaluation result of the second voice being greater than or equal to a first threshold, a question for the next round of conversation may be output. The first threshold may be set as needed. In other embodiments, after outputting the question for the next round of conversation in step 531, an answer prompt corresponding to the question may be output.
[0101] Optionally, in step 532, in response to the oral evaluation result of the second speech being lower than the first threshold, the second speech may be classified, and a corresponding first operation may be performed based on the first classification obtained. With this configuration, whether to provide further prompt feedback can be determined based on the oral evaluation result, and personalized further response prompts can be provided for different errors of different users.
[0102] In some embodiments, classifying the second speech may include classifying the reasons why the oral evaluation result of the second speech is lower than a first threshold, that is, the first category of the second speech can be determined based on the dimensions with lower scores in the oral evaluation result of the second speech (such as semantic relevance, pronunciation, fluency, completeness or error rate, etc.).
[0103] In another embodiment of the present invention, the above-mentioned first category may include one or more of the following: semantic irrelevance, inaccurate pronunciation, incomplete answer, etc.; and its corresponding first operation may include: when the first category is semantic irrelevance, based on the number of semantic irrelevances in the current round of conversation, determining a second answer prompt with different degrees of completeness; when the first category is inaccurate pronunciation, outputting pronunciation prompt information for prompting re-pronunciation; and / or when the first category is incomplete answer, outputting a third answer prompt about the incomplete part. In some embodiments, the second answer prompt, pronunciation prompt information and / or the third answer prompt can be output in a visual and / or audible form. The first category and the corresponding first operation will be exemplarily described below.
[0104] For example, in some application scenarios, the machine outputs the question "What's your favorite color?" and the first response prompt "Answer prompt: I like pink best." For user A, the oral assessment results show that the user's pronunciation of "color" scores below 40 points, while the overall sentence is complete and has no grammatical errors. This means that the first category is inaccurate pronunciation. In this case, the machine's first action is to output a pronunciation prompt, such as "Pronunciation reminder: The pronunciation of "color" is not good enough. Please try again." In other application scenarios, the machine outputs the question "What's your favorite color?" and the first response prompt "Answer prompt: I like pink best." For user B, user B's answer does not mention the word "pink", but there are no obvious problems with other parts. This means that the first category is incomplete. In this case, the machine's first action is to output a third response prompt for the incomplete part, such as "Keyword prompt: pink means pink."
[0105] In other embodiments, for example, when the second voice received based on the current round of conversation is semantically unrelated to the question for the first time, the second answer prompt may include a keyword prompt and may be output in the second language or in a combination of the native language and the second language. When the second voice received based on the current round of conversation is semantically unrelated to the question for the second time, the second answer prompt may include a recommended answer and may be output in the second language or in a combination of the native language and the second language. In order to facilitate understanding of the first operation when the first category is semantically unrelated, an exemplary description will be given below in conjunction with 6.
[0106] Figure 6 Schematically shows a conversation flow chart including outputting a second answer prompt according to an embodiment of the present invention. Figure 6 As shown in the figure, the numbers 1 and 2 in the circles represent the questions of the adjacent rounds of conversation, and the Ⅰ, Ⅱ, and Ⅲ in the square boxes are used to represent the number of times the second voice for the question of the current round of conversation (① in the figure) is received. Specifically, when the current round of conversation begins, the machine outputs the question "what's your favorite color?" and the first answer prompt "Answer prompt: I like pink the most", and then the machine receives the first second voice Ⅰ in which the user responds based on the first answer prompt. In response to the oral evaluation result of the first second voice Ⅰ being higher than or equal to the first threshold, that is, the first second voice Ⅰ is semantically relevant to the question, the question for the next round of conversation is output (shown as ② in the figure). In response to the first category of the first second voice Ⅰ being semantically irrelevant, that is, the user's answer completely deviates from the meaning of the recommended answer, the machine can output a second answer prompt including a keyword prompt "Keyword prompt: I'm fine is often used to express that it is okay".
[0107] Next, the machine receives the second second voice II in which the user responds based on the second response prompt. In response to the oral evaluation result of the second second voice II being higher than or equal to the first threshold, that is, the second second voice II is semantically relevant to the question, the machine outputs the question for the next round of conversation (indicated by ② in the diagram). In response to the first category of the second second voice II still being semantically irrelevant, at this time in the current round of conversation, the two second voices received are both semantically irrelevant (i.e., the number of semantic irrelevances is 2 times), that is, the user's response completely deviates from the meaning of the recommended response for the second time, then the machine can output a second response prompt including the recommended response, "You can say: I'm fine."
[0108] The machine then receives a third second voice III from the user in response to the second response prompt including the recommended response. In response to the oral evaluation result of the third second voice III being higher than or equal to the first threshold, i.e., indicating that the third second voice III is semantically relevant to the question, the machine outputs the question for the next round of conversation (indicated by ② in the diagram). In further embodiments, in response to the first category of the third second voice III still being semantically irrelevant, a more complete guided second response prompt may be output, such as "Please say after me: I'm fine," or the output speed of the second response prompt may be reduced simultaneously.
[0109] Combination of the above Figure 5 and Figure 6 The challenge training mode according to the embodiment of the present invention is described in detail. It should be understood that in the challenge training mode, for each prompt link, it is still acceptable for the user to answer with a semantically similar sentence. Specifically, no matter which of the above-mentioned answer prompts the current response is, similar sentences can be accepted for answering. For example, the first response prompt is "Answer prompt: I'm doing well", even if the recommended response is "I'm fine", responses such as "I'm OK" or "I'm good" that mean "I'm doing well" can also be accepted by the machine, that is, the machine can store multiple recommended responses with similar semantics.
[0110] It should also be understood that the above description is exemplary and not restrictive. For example, in other embodiments, method 500 may further include: in response to the end of each round of conversation for the conversation content, determining a total evaluation result based on the oral evaluation results of each round of conversation. The total evaluation result is used to evaluate the user's overall performance in the challenge training mode. The total evaluation result can be displayed to the user in a visual and / or audible form, and the specific details of each response of the user can be displayed. The user can choose to re-enter the challenge training mode for training, or enter the next more difficult oral training mode based on the total evaluation result.
[0111] Figure 7The flowchart of the oral training method for entering the difficult training mode according to an embodiment of the present invention is schematically shown. It should be understood that the difficult training mode can be one of multiple oral training modes, and the method 700 for entering the difficult training mode for oral training can be a specific manifestation of the oral training method 200 or the oral training method 300. Therefore, the above description of the difficult training mode can be omitted. Figure 2 and Figure 3 The description can also be applied to the following Figure 7 's description.
[0112] like Figure 7 As shown, method 700 may include: in step 710, in response to entering the difficult training mode, questions in the conversation content may be output. The purpose of setting the difficult training mode is to enable the user to complete the conversation content in a free dialogue manner, and to be able to try to get rid of external help and achieve the effect of speaking fluently based on the training in the previous simpler training mode. In some embodiments, in the difficult training mode, questions in the conversation content can be output in the form of voice, and at least one recommended answer (or reference answer) corresponding to the question in the conversation content can be stored in the form of text.
[0113] Next, in step 720, a third voice response from the user to the question may be received. In the difficult training mode, after the question is output, no response prompt is output. Instead, the user's third voice response is directly received to train the user's free conversation ability. In some application scenarios, the user can recall possible responses based on the machine-generated questions and the training process in the aforementioned follow-up training mode and / or challenge training mode.
[0114] Then, the process can proceed to step 730, where a spoken language evaluation can be performed on the third speech to determine whether to output the next round of conversation content. Here, the spoken language evaluation can include at least one of semantic relevance, pronunciation, fluency, completeness, error rate, etc. The semantic relevance evaluation can include a semantic relevance score between the third speech and the corresponding question, which can be implemented using existing or future semantic analysis technology. The specific contents of the pronunciation evaluation, fluency evaluation, completeness evaluation and error rate evaluation are combined with the above. Figure 4 The oral evaluation in step 430 is the same or similar and will not be described again here.
[0115] In some embodiments, oral evaluation can be achieved by comparing the received third voice with a reference answer and using voice scoring technology. Comparing the third voice with the reference answer may include at least one of the following: comparing the digital representation of the third voice with the digital representation of the reference answer; and comparing the third text converted from the third voice with the reference answer. The conversion of the third voice into the third text can be achieved by voice-to-text technology. In other embodiments, in order to ensure that the user and the machine can have a smooth conversation and simulate the conversation process of a real scene, the oral evaluation results may not be presented before the end of all the conversation content.
[0116] like Figure 7 As further shown in FIG, optionally or additionally, step 730 may include step 731 (shown in a dotted box) or step 732 (shown in a dotted box), wherein in step 731, in response to the spoken language evaluation result of the third speech being higher than or equal to a second threshold, a question for the next round of conversation is output. The second threshold can be set as needed.
[0117] In another embodiment of the present invention, before outputting questions for the next round of conversation, the spoken language training method 700 may further include: in response to receiving questions other than the conversation content, performing any of the following operations: skipping the current round of conversation; or outputting response information related to the other questions. The other questions here are not part of the originally set conversation content.
[0118] In some embodiments, NLP technology can be used to perform semantic analysis on other questions received, so that the machine can determine the response information that is semantically related to the other questions based on the semantic analysis results, and can output it in an audible and / or visual form, and then pull the conversation back to the originally set conversation content to continue the conversation. For example, in some application scenarios, the machine outputs the question "Where are you from?", and the machine receives the user's third voice as "I'm from China. Do you like China?". Obviously, the second half of the sentence "Do you like China?" does not belong to the scope of this conversation content. The machine uses semantic analysis technology to understand and judge the user's speech content and intention, and can output corresponding response information, such as "I like China.", and then continue to output the next round of questions in the originally set conversation content, such as "Then, tell me about your hometown."
[0119] In other embodiments, when other received questions cannot be understood or recognized, such as due to noise or background sound, the current round of conversation can be skipped and the next round of questions in the originally set conversation content can be directly output.
[0120] By skipping the current conversation or outputting responses related to other questions, appropriate responses can be given to different user responses, ensuring the overall conversation stays on topic. This ensures the continuity of the scenario conversation and addresses the problem of inefficient teaching caused by the uncertain direction of free conversation topics.
[0121] Optionally, in step 732, in response to the spoken language evaluation result of the third speech being lower than the second threshold, the third speech may be classified, and a corresponding second operation may be performed based on the second category obtained by classification. In some embodiments, classifying the third speech may include classifying the reason why the spoken language evaluation result of the third speech is lower than the second threshold, that is, the second category of the third speech may be determined based on the dimensions (e.g., semantic relevance, grammar, pronunciation, fluency, completeness, or error rate) with lower scores in the spoken language evaluation result of the third speech.
[0122] In another embodiment of the present invention, performing a corresponding second operation based on the second category may include: when the second category is semantically irrelevant, the corresponding second operation may include repeatedly outputting the question; and / or when the second category is a category other than semantically irrelevant, the corresponding second operation may include outputting recommendation information related to the question.
[0123] Specifically, when the second category is semantically irrelevant, meaning the third speech completely deviates from the semantic meaning, resulting in a spoken language assessment result below the second threshold, the machine can be controlled to repeat the current question. During execution, the machine can be controlled to simulate a human conversation. For example, in some application scenarios, if the user's response to the third speech is semantically deviated due to incomprehension, the machine can repeat the current question using phrases typically used by humans, such as "I mean, how have you been lately?" or "I just said, are you doing well?", rather than mechanically repeating the entire question. For example, if the machine outputs the question "Where are you from?" and the received third speech completely deviates from the meaning, the machine will repeat the current question in this manner, such as "I said, where are you from?" This setting can simulate a more realistic and humanized spoken conversation environment, helping to improve users' application and adaptability in real-world contexts.
[0124] The categories other than semantic irrelevance mentioned above may include at least one of pronunciation, fluency, completeness, and error rate. This indicates that the user understands the question but lacks oral expression skills. In this case, recommended information related to the question may be output. This recommended information may include recommended answers or keywords within the recommended answers.
[0125] In other embodiments, method 700 may further include: in response to the completion of each round of conversation regarding the conversation content, determining an overall evaluation result based on the oral evaluation results of each round of conversation. The overall evaluation result is used to evaluate the user's overall performance in the difficult training mode. The overall evaluation result can be displayed to the user in a visual and / or audible form and can display specific details of each user's response, such as the user's pronunciation and text information obtained through voice recognition of the user's response. One or more recommended responses can also be output for each question or conversation round with a low score.
[0126] Combination of the above Figure 7 An exemplary description of the difficult training mode according to an embodiment of the present invention is provided. It can be understood that in the difficult training mode, prompt information can be output only in the round of conversations in which the oral evaluation results are unqualified, while the oral evaluation results are not displayed in other rounds of conversations, and no prompt information is output. This can more realistically simulate the actual oral conversation scene, which is conducive to improving the oral training effect for users, and truly helping users to achieve free conversation in a real context, thereby getting rid of the dilemma of "dumb English".
[0127] Through the above description of the technical solution and multiple embodiments of the present invention in combination with the accompanying drawings, those skilled in the art can understand that by setting multiple oral training modes of different difficulty levels for the same conversation content, a step-by-step oral training method (for example, a process from imitation to attempting free oral expression) can be achieved, and the oral training method implemented by the machine according to the present invention can significantly reduce the investment in manpower and financial resources. A large number of training operations can be carried out anytime and anywhere according to user needs, so users can use fragmented time or after-school time for supplementary oral training.
[0128] In some embodiments, by setting different response prompts based on the oral evaluation results of each answer and the number of answers to the same question in the challenge training mode, in order to control the direction of the conversation during the training process, users can gradually adapt to the conversation process and obtain targeted feedback and response guidance. In other embodiments, by presenting the oral evaluation results and / or the overall evaluation results, the training results are presented in a visual manner, allowing users to see their own answer errors during the conversation and the gap between their answers and the recommended answers, which can help users perceive the effectiveness of their oral training for further improvement and perfection.
[0129] Furthermore, although the operations of the present method are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in that particular order, or that all of the operations shown must be performed to achieve the desired results. Rather, the steps depicted in the flowcharts may be performed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into a single step, and / or a single step may be broken down into multiple steps.
[0130] The use of the verbs "comprise," "include," and their conjugations in the application documents does not exclude the presence of elements or steps other than those recited in the application documents. The article "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. It should also be understood that the terms "first," "second," "third," and "fourth," etc. in the claims, description, and drawings of the present invention are used to distinguish between different objects, rather than to describe a particular order.
[0131] Although the spirit and principles of the present invention have been described with reference to several specific embodiments, it should be understood that the present invention is not limited to the specific embodiments disclosed, and the division into various aspects does not mean that the features of these aspects cannot be combined to benefit. Such division is merely for the convenience of expression. The present invention is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims. The scope of the appended claims is to be accorded the broadest interpretation so as to encompass all such modifications and equivalent structures and functions.
Claims
1. A method for oral language training implemented by a machine, comprising: Based on the same conversation content, multiple oral training modes with different levels of difficulty are set, wherein the multiple oral training modes include at least two of a follow-up training mode, a challenge training mode, and a difficult training mode with increasing difficulty; as well as Based on the difficulty levels of the multiple oral training modes or user selections, entering a corresponding oral training mode to conduct multi-dimensional training of different difficulty levels for the same conversation content; Based on the difficulty of multiple oral training modes, you can enter the corresponding oral training mode, including: Enter the corresponding oral training mode in order from easy to difficult according to the order of the multiple oral training modes; and Determine whether to enter the next speaking training mode based on the user's overall evaluation results in the current speaking training mode; In response to entering the challenge training mode, outputting a question in the conversation content and a first answer prompt corresponding to the question; In response to entering the difficult training mode, questions in the conversation content are output, and after the questions are output, no answer prompts are output.
2. The spoken language training method according to claim 1, further comprising: grading the difficulty of each candidate conversation content in terms of content difficulty; as well as Based on the difficulty level corresponding to the user's spoken language proficiency, a range of candidate conversation contents selectable by the user is determined, so that the user can select the conversation content within the range.
3. The oral language training method according to claim 2, wherein the difficulty of the content is determined based on at least one of the following: The topic of the candidate conversation content; Vocabulary of candidate conversation content; the grammar of the candidate conversation content; and The sentence length of the candidate conversation content.
4. The spoken language training method according to any one of claims 1 to 3, wherein: The spoken language training method also includes: In response to entering the follow-up reading training mode, determining a first role of the user in the conversation content and a target sentence corresponding to the first role; Outputting the conversation content, and when outputting a target sentence of the first character, receiving a first voice of a user reading the target sentence; and Based on the target sentence, a spoken language evaluation is performed on the first speech to determine whether to output the next round of conversation content.
5. The spoken language training method according to claim 4, further comprising: In response to the completion of each round of conversation for the first role, determining a second role of the user in the conversation content that is different from the first role, so as to continue the shadowing training; as well as In response to the completion of each round of conversation for each role in the conversation content, a total evaluation result is determined based on the oral evaluation results of each round of conversation.
6. The oral training method according to any one of claims 1 to 3, wherein: The spoken language training method also includes: In the challenge training mode, receiving a second voice message in which the user responds based on the first response prompt; and Based on the first response prompt, an oral evaluation is performed on the second voice to determine whether to output the next round of conversation content.
7. The spoken language training method according to claim 6, wherein determining whether to output the next round of conversation content comprises: In response to a spoken language evaluation result of the second speech being higher than or equal to a first threshold, outputting a question for a next round of conversation; or In response to a spoken language evaluation result of the second speech being lower than a first threshold, the second speech is classified, and a corresponding first operation is performed based on a first category obtained by the classification.
8. The spoken language training method according to claim 7, wherein The first category includes one or more of the following: semantic irrelevant, inaccurate pronunciation, incomplete response; and The corresponding first operation includes: When the first category is semantically irrelevant, determining second response prompts with different degrees of completeness based on the number of semantically irrelevant responses in the current round of conversation; When the first category is inaccurate pronunciation, outputting pronunciation prompt information for prompting re-pronunciation; and / or When the first category is an incomplete answer, a third answer prompt regarding the incomplete portion is output.
9. The spoken language training method according to any one of claims 1 to 3, wherein: The spoken language training method also includes: In the difficult training mode, receiving a third voice message in which the user responds to the question; and Performing an oral evaluation on the third speech to determine whether to output the next round of conversation content.
10. The spoken language training method according to claim 9, wherein determining whether to output the next round of conversation content comprises: In response to a spoken language evaluation result of the third speech being higher than or equal to a second threshold, outputting a question for a next round of conversation; or In response to the spoken language evaluation result of the third speech being lower than a second threshold, the third speech is classified, and a corresponding second operation is performed based on a second category obtained by the classification.
11. The spoken language training method according to claim 10, wherein performing the corresponding second operation based on the second category comprises: When the second category is semantically irrelevant, the corresponding second operation includes repeatedly outputting the question; and / or When the second category is a category other than the semantically irrelevant category, the corresponding second operation includes outputting recommendation information related to the question.
12. The oral training method according to claim 10, before outputting questions for the next round of conversation, further comprising: In response to receiving other questions other than the conversation content, perform any one of the following operations: skipping the current round of conversation; or Response information related to the other questions is output.
13. A device for implementing spoken language training, comprising: a processor configured to execute program instructions; as well as A memory configured to store the program instructions, which, when loaded and executed by the processor, enables the device to perform the spoken language training method according to any one of claims 1 to 12.
14. A computer-readable storage medium storing program instructions, wherein when the program instructions are loaded and executed by a processor, the processor is caused to execute the spoken language training method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Network multi-element intelligent promotion system
CN101261692A
Foreign language learning system and method incorporated with traditional culture
CN106898166A
Spoken language training interaction method and terminal equipment
CN111639218A
Spoken language practice method and device
CN112116832A
Spoken language evaluation method and device, and related product
CN112951207A