Human-computer dialogue oral evaluation system

Through the conversational human-computer interaction oral assessment system, combined with voice recognition and situational dialogue assessment, the problems of rigid English oral assessment and emotional interference in existing technologies are solved, multi-dimensional feedback and personalized guidance are achieved, learning interest is stimulated, and students' practical language application ability is improved.

CN112232083BActive Publication Date: 2025-09-16SHANGHAI SQUIRREL CLASSROOM ARTIFICIAL INTELLIGENCE TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202011100849.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-08-23
Publication Date
2025-09-16
Estimated Expiration
2039-08-23

AI Technical Summary

Technical Problem

Existing English oral assessment technology is rigid in form, unable to fully reflect students' actual language ability, and unable to provide personalized guidance. Traditional exams are easily disturbed by emotions, lack scenario-driven and task-oriented approaches, and cannot stimulate learning interest.

Method used

A conversational human-computer interaction oral evaluation system is designed. Through speech recognition, intent understanding, dialogue management and language generation modules, combined with situational dialogue speech and semantic evaluation, it provides multi-dimensional evaluation and feedback, supporting natural language communication and task-oriented dialogue.

Benefits of technology

It improves the authenticity and feedback effect of English oral assessment, reduces students' tension, stimulates their interest in learning, provides personalized guidance, comprehensively reflects the learning process and content, and improves their comprehensive English communication skills.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112232083B_ABST
    Figure CN112232083B_ABST
Patent Text Reader

Abstract

This application relates to a human-computer dialogue oral assessment system, which is a scenario-driven, task-oriented dialogue system for oral assessment based on human-computer dialogue and voice assessment technologies. The assessment system of this application has three main features: conversational, scenario-driven, and task-oriented. Through the task-oriented dialogue system that communicates with the user's natural language, it can understand the student user's actual language ability and ability to comprehensively use English for communication, which has a feedback effect on the student user's oral learning and the teacher's oral teaching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of human-computer interaction technology, and in particular to a conversational human-computer interaction spoken language assessment system. Background Art

[0002] There are two main types of oral tests: interviews and recorded oral exams. Interviews are more effective but time-consuming and labor-intensive to organize. Large-scale oral exams utilize a human-computer interaction approach. Test-takers simply use a computer and headset to complete their listening and speaking test responses. These responses are then automatically scored intelligently, assessing sentence rhythm, completeness, accuracy, and other factors. A report on the responses is also generated.

[0003] In online language training products, the use of speech recognition technology and speech assessment technology has become quite common. Through the method of "listening to the original sound - following / retelling - system scoring - multi-color visual feedback - adjustment", the pronunciation of students is compared with the pronunciation of the machine and scored. Through repeated practice, students can achieve the goal of improving English listening and pronunciation. Summary of the Invention

[0004] After long-term observation and research, the inventors discovered that spoken English is different from other courses. Its primary purpose is not to impart knowledge. English is a carrier of knowledge and culture. Students need to use language to express their thoughts and communicate with others to achieve the purpose of true cultivation. Cultivating students' ability to use language in practice and improving their ability to communicate in English has become the main teaching task of spoken English. Examinations and assessments should serve teaching. However, English assessment technology applied to human-computer interaction still has the following shortcomings:

[0005] First, assessing students' oral proficiency through pre-recorded audio tests is a rigid and monotonous format. Not only are the topics predetermined, but the content is also prescriptive, leaving students passively accepting the questions and scores. Test-oriented oral exams typically involve students speaking, examiners listening, and then assigning a score, which far from fully reflects the teaching and learning process. Furthermore, during interviews, the emotional interaction between examiners and students can also distort the assessment results.

[0006] 2. Traditional classroom or online oral assessment is a terminal assessment of test-based examinations. It is a question-driven assessment experience. It determines the students' learning results for a semester through a one-time final exam, or through a diagnostic test before the start of the semester to determine the course level of the students. Students are then upgraded one by one.

[0007] 3. During learning, through shadowing / retelling activities, students compare their own pronunciation with that of the machine, and repeatedly correct their pronunciation based on the scoring feedback. This practice is helpful for English listening and pronunciation, but it is impossible to assess students' actual level of ability to use the language and their ability to communicate in English comprehensively through existing technology, and it cannot provide any inspiration for oral English learning.

[0008] In view of the above-mentioned shortcomings of the existing technology, the present application provides a conversational human-computer interactive oral assessment system. It is a task-oriented dialogue system based on human-computer dialogue and voice assessment technologies, and is applied to oral assessment. The assessment system of the present application has three main features: conversational, scenario-driven, and task-oriented. Through the task-oriented dialogue system that communicates with the user's natural language, it can understand the student user's actual language ability and ability to comprehensively use English for communication, which has a feedback effect on the student user's oral learning and the teacher's oral teaching.

[0009] The present application provides a conversational human-computer interaction oral assessment system, including a dialogue system, which includes: a speech recognition module, which is configured to recognize the user's speech input and convert it into text; an intention understanding module, which is configured to perform semantic understanding on the converted text to recognize the user's intention; a dialogue management module, which is configured to generate corresponding system actions based on the understanding results of the intention understanding module; a language generation module, which is configured to convert the system actions generated by the dialogue management module into natural language; and a language synthesis module, which is configured to convert the natural language into speech and feed it back to the user.

[0010] In some embodiments, the intent understanding module is optionally further configured to be capable of slot filling, where a slot is information that needs to be completed to convert the user intent into a clear user instruction during the conversation.

[0011] In some embodiments, optionally, the intent understanding module is further configured to be able to understand the user intent based on the user portrait and / or scenario information.

[0012] In some embodiments, optionally, the dialogue management module further includes a dialogue state tracking module, which is configured to indicate the stage of the dialogue and integrate contextual information of the dialogue process.

[0013] In some embodiments, optionally, the dialogue management module further includes a dialogue strategy learning module, which is configured to generate the next operation of the system according to the current dialogue state.

[0014] In some embodiments, optionally, an evaluation system is also included, which includes: a situational dialogue speech and semantic evaluation module, which is configured to perform similarity comparison on the text converted from the user's speech based on standard speech and semantic content, and obtain a speech evaluation score and a semantic evaluation score; a grammar evaluation and error checking module, which is configured to perform grammar checking on the text converted from the user's speech and obtain a grammar evaluation score; and an easy-to-mix evaluation module, which is configured to mark easy-to-mix errors in the text converted from the user's speech for easy-to-mix evaluation.

[0015] In some embodiments, optionally, the dialogue management module is further configured to generate corresponding system actions according to the evaluation results of the evaluation system.

[0016] In some embodiments, optionally, the higher the similarity between the user's voice and the phonemes of the standard voice, the higher the voice evaluation score; and the higher the similarity between the content expressed by the user and the comparison reference answer, the higher the semantic evaluation score.

[0017] In some embodiments, optionally, the grammar evaluation and error checking module is further configured to examine the logical relationships in the sentence, where the logical relationships include one or more of the following relationships: subject-verb collocation, tense expression, syntactic structure, singular and plural.

[0018] In some embodiments, optionally, the conversational human-computer interaction oral assessment system is based on a stand-alone and / or online configured computer system to carry out assessment of language content.

[0019] Compared with the prior art, the present invention has the following advantages:

[0020] First, this application is a conversational human-computer interactive oral assessment system. Through human-computer dialogue, it provides numerous opportunities for interaction with different virtual humans, creating communication scenarios. Through repeated communication and practice, it can have a positive impact on student learning and teaching. This washback effect can change students' learning attitudes and stimulate their enthusiasm for learning and using spoken English. Furthermore, this conversational human-computer interactive oral assessment system can also avoid emotional interference between human examiners and test takers.

[0021] Second, this application is a scenario-driven oral assessment system, which is a technology that is meaningful and can reflect the content being taught, while also reflecting the learning content and learning process. Not only can detailed assessment feedback be obtained in the process of completing learning tasks, including: discovering problems in student users' voice, intonation, communication, and expression, and analyzing the causes of the problems, but also rich student user voices and adopted communication strategies can be collected, which is very meaningful for subsequent teachers to provide personalized guidance to student users. Furthermore, scenario-driven assessments can reduce student users' tension and anxiety, and more realistically reflect the student users' true level and performance.

[0022] Third, this application is a task-oriented oral assessment system. Task-based oral communication activities focus on the expression of meaning rather than the standardized form of language, which makes it easy for students to experience success and a sense of accomplishment, thereby stimulating their intrinsic interest and desire in learning and achieving better performance. Communicative English speaking emphasizes providing students with opportunities for personal experience, through participating in real, natural and communicative activities, to find knowledge, discover problems, construct their own communication models, concepts and strategies, and achieve the learning purpose of conveying information and expressing ideas by completing tasks.

[0023] The concept, specific structure and technical effects of this application will be further explained below in conjunction with the accompanying drawings to fully understand the purpose, characteristics and effects of this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The present application will become more readily understood when the following detailed description is read in conjunction with the accompanying drawings, in which like reference numerals represent like parts throughout, and in which:

[0025] Figure 1 This is a schematic diagram of the functional module structure of an embodiment of the present application.

[0026] Figure 2 This is a schematic diagram of the program module structure of an embodiment of the present application. DETAILED DESCRIPTION

[0027] The technical solutions in the embodiments of the present application will be described clearly and completely below. Obviously, the embodiments described are part of the embodiments of the present application, rather than all of the embodiments. The present application can be embodied through many different forms of embodiments, and the scope of protection of the present application is not limited to the embodiments mentioned in the text. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work should fall within the scope of protection of the present application.

[0028] The ordinal numbers used in this application, such as "first" and "second," are used solely for distinction and identification purposes and do not have any other meaning. Unless otherwise specified, they do not imply a specific order or relationship. For example, the term "first component" does not imply the existence of a "second component," nor does the term "second component" imply the existence of a "first component."

[0029] Figure 1 This is a schematic diagram of the functional module structure of an embodiment of the present application. Figure 1 As shown, the conversational human-computer interaction oral evaluation system can be based on a stand-alone and / or online configured computer system to carry out language content evaluation, including a dialogue system and an evaluation system.

[0030] The dialogue system includes a speech recognition module, an intent understanding module, a dialogue management module, a language generation module, and a language synthesis module. The speech recognition module recognizes user speech input and converts it into text; the intent understanding module performs semantic interpretation on the converted text to identify the user's intent; the dialogue management module generates corresponding system actions based on the intent understanding module's understanding results; the language generation module converts the system actions generated by the dialogue management module into natural language; and the language synthesis module converts natural language into speech and feeds it back to the user.

[0031] In some embodiments, the speech recognition module is responsible for recognizing the student user's voice input and converting it into text; the intent understanding module is responsible for semantic understanding of the text converted from the student user's speech, including user intent recognition and slot filling, where slots are the information required to convert user intent into clear user instructions during the conversation; the dialogue management module is responsible for overall dialogue management, including dialogue state tracking and dialogue strategy learning; the language generation module is responsible for converting the system actions selected by the dialogue strategy module into natural language; and the language synthesis module is responsible for converting text into speech and ultimately feeding it back to the student user. The intent understanding module can also understand user intent based on user profiles and / or scenario information.

[0032] Intent can be viewed as a text-based multi-classification problem, where the corresponding category is determined based on the user's expression. Intent can be understood as an application function or process that primarily fulfills the user's request and purpose. When a student user expresses "My name is Carol" or "This is Carol," both expressions may trigger the intention of introducing themselves. A slot is the information required to translate initial user intent into a clear user instruction during a multi-round conversation. A slot corresponds to a piece of information required to process a task. In the student user's expression "My name is Carol," "Carol" represents the name slot. In addition to voice input, the intent understanding module also considers user profiles and contextual information. This more comprehensive context improves the accuracy of intent understanding.

[0033] The user profile can include: the student's name, grade, location, spoken language proficiency (such as pitch accuracy, completeness, and fluency), behavioral traits, personality traits, and hobbies. Each round of conversation updates the user profile in real time, influencing contextual information in the next round. Combined with this contextual information, the virtual persona has a memory function. As conversations increase in frequency, the system gains a deeper understanding of the student, and the virtual persona's responses to the student become more fluent.

[0034] The dialogue management module may also include a dialogue state tracking module and / or a dialogue strategy learning module. The dialogue state tracking module can indicate the current stage of the dialogue and incorporate contextual information from the dialogue process. The dialogue strategy learning module can generate the system's next action based on the current dialogue state. In some embodiments, the dialogue state tracking module is used to represent the current dialogue state information, representing the current stage of the entire dialogue within the dialogue system and incorporating contextual information from the dialogue process. The dialogue strategy learning module is used to generate the system's next action based on the current dialogue state.

[0035] The evaluation system may include a scenario dialogue speech and semantics evaluation module, a grammar evaluation and error checking module, and an easy-to-confuse evaluation module. The scenario dialogue speech and semantics evaluation module can compare the similarity of the text converted from the user's speech based on the standard content of speech and semantics, and obtain speech evaluation scores and semantic evaluation scores; the grammar evaluation and error checking module can perform grammar checking on the text converted from the user's speech and obtain a grammar evaluation score; and the easy-to-confuse evaluation module can mark easy-to-confuse errors in the text converted from the user's speech to evaluate the easy-to-confuse.

[0036] In some embodiments, the evaluation system may include three modules: speech and semantic evaluation of situational dialogues, grammar evaluation and error checking, and easy-to-confuse evaluation. The speech and semantic evaluation module of situational dialogues is responsible for comparing the similarity of the text converted from the student user's voice with the standard content of speech and semantics. The higher the similarity between the user's voice and the standard voice phonemes, the higher the speech evaluation score. The higher the similarity between the user's expression and the reference answer, the higher the semantic evaluation score. The grammar evaluation and error checking is responsible for scoring and pointing out grammatical errors in the text converted from the student user's voice. It mainly examines the logical relationships in the sentence, including singular and plural, subject-verb collocation, tense expression, and the use of syntactic structures. The fewer grammatical errors, the higher the evaluation score. The easy-to-confuse evaluation module is responsible for marking easy-to-confuse errors in the text converted from the student user's voice. To achieve easy-to-confuse evaluation, it is necessary to incorporate common mistakes made by Chinese students into the training corpus of the model in the speech recognition module to avoid the speech recognition module actively correcting errors.

[0037] The dialogue management module can generate corresponding system actions based on the evaluation results of the evaluation system. In some embodiments, the evaluation results of the three modules of the evaluation system will be input into the dialogue management module of the dialogue system. After receiving the evaluation results of the evaluation system on the user's voice, the dialogue management module can respond based on the evaluation goals and strategies.

[0038] Figure 2 This is a schematic diagram of the program module structure of an embodiment of this application. Figure 2 As shown, the system first takes out the first test point, which corresponds to a task that needs to be completed in a scene, and the student user sees the task description on the front-end interface.

[0039] In some embodiments, in a conversational human-computer interaction oral assessment system: the description of the task, for student users, has conversation background and scene information, and student users are to complete a task-based activity that is real, natural, and communicative. When the front-end system is virtual reality, student users can also obtain the same experience as having a real conversation with a person from the rich three-dimensional information.

[0040] By adopting this technical solution: the system starts a conversation based on the contextual information. Depending on the needs of different test points, both the user and the system may start to ask questions or send questions. When the student user's voice is converted into text through voice recognition and the intention is recognized by the intent recognition module, the text will be evaluated by the evaluation module to obtain scores and error content in multiple dimensions such as voice, semantics, grammar, and easy-to-mix. This new information will be updated to the user portrait.

[0041] In some embodiments, the assessment module in a conversational human-computer interaction oral language assessment system includes: speech and semantic assessment of the scenario dialogue, grammatical assessment and error checking, and ambiguity assessment. In addition to providing a post-assessment report, this assessment also serves as information for the virtual human's responses, enabling automatic adjustments to language complexity, speech speed, and comprehension based on the conversation partner.

[0042] By adopting this technical solution: after the student user's voice is converted into text through voice recognition, the text obtains the intention of the conversation through intent recognition, and extracts slots based on the student user's expression, thereby understanding the student user's voice and deciding the content of the next conversation. Through language generation, the virtual person speaks it out. The whole process cycles through multiple test points until the end of the assessment and generates an assessment report.

[0043] In some embodiments, in the above-mentioned conversational human-computer interaction oral assessment system: the assessment report includes: student basic information, assessment results of the oral proficiency process, and can point out student user voice and grammatical errors, such as non-standard voice, inaccurate intonation, common grammatical errors, etc., and can further analyze the student user's ability to comprehensively use language and the communication strategies used from the behavioral characteristics of the student user.

[0044] In some embodiments, the conversational human-computer interaction spoken language assessment system may include two major parts: a dialogue system and an assessment system. In practice, as an example, its working process is as follows:

[0045] The system first takes out the first test point, which corresponds to a task that needs to be completed in a scenario. The student user sees the task description on the front-end interface. For example, the test point is to get to know strangers through English expressions. The system can display appropriate dialogue scenarios through rich text or virtual reality. The student user sees the task description as follows: Get to know new friends, greet them politely, and ask the other person's name and where they are from.

[0046] The system starts a conversation based on contextual information. The setting of this test point is to let the user start asking questions. When the student user says "Hello, I'm Ray. What's your name?", the student user's voice is converted into text through speech recognition. The text uses intent recognition to determine that the intention of the conversation is a greeting. The evaluation module then obtains scores in multiple dimensions including speech, semantics, grammar, and ease of mixing, and updates the scores to the user portrait.

[0047] Intent recognition determines that the intention of the conversation is to say hello, and extracts the slot based on the student user's expression, that is, the slot is extracted as the name, and the parameter value is Ray. After understanding the student user's voice, it is necessary to decide the content of the next conversation, and let the virtual person speak it through language generation. The whole process loops and extracts multiple test points until the evaluation is completed, and an evaluation report is generated.

[0048] In some instances, the system asked, "Where do you come from?" and the student user responded with the name of a small city in their hometown, which was beyond the system's comprehension. The system's dialogue state tracking module incorporated contextual information based on the current stage of the conversation. The dialogue strategy learning module adopted a universal response strategy, allowing the system to keep the conversation going by using a virtual person to respond with "Wow! That is a nice place!"

[0049] In some instances, this may also include: when a student user says "I want to make a phone call" in an airplane scenario, the system learns from the scenario information module that using a mobile phone on an airplane is not allowed, and learns from the user portrait that the student user has a low social interaction norm score, and will give priority to a serious and persuasive response in the dialogue strategy selection.

[0050] In some embodiments, the various methods, processes, modules, devices, equipment or systems described above may be implemented or executed in one or more processing devices (e.g., digital processors, analog processors, digital circuits designed for processing information, analog circuits designed for processing information, state machines, computing devices, computers and / or other mechanisms for electronically processing information). The one or more processing devices may include one or more devices that perform some or all operations of the method in response to instructions stored electronically on an electronic storage medium. The one or more processing devices may include one or more devices that are configured through hardware, firmware and / or software and are specifically designed to perform one or more operations of the method. The above is only a preferred specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with the technical field of the present invention, within the technical scope disclosed in the present application, shall make equivalent substitutions or changes based on the technical solution and inventive concept of the present application, which shall be covered within the scope of protection of the present application.

[0051] The embodiments of the present application can be carried out in hardware, firmware, software or its various combinations. The present application can also be implemented as an instruction stored on a machine-readable medium and that can be read and executed using one or more processing devices. In one embodiment, a machine-readable medium may include various mechanisms for storing and / or transmitting information in a machine (e.g., computing device) readable form. For example, a machine-readable storage medium may include a read-only memory, a random access memory, a disk storage medium, an optical storage medium, a flash memory device and other media for storing information, and a machine-readable transmission medium may include various forms of propagation signals (including carrier waves, infrared signals, digital signals) and other media for transmitting information. Although firmware, software, routines or instructions may be described in the above disclosure in terms of the specific exemplary aspects and embodiments of performing certain actions, it will be apparent that such descriptions are only for convenience purposes and such actions are actually generated by machine equipment, computing devices, processing devices, processors, controllers, or other devices or machines that perform firmware, software, routines or instructions.

[0052] This specification uses examples to disclose the present application, one or more of which are described or illustrated in the specification and its drawings. Each example is provided for the purpose of explaining the present application and is not intended to limit the present application. In fact, it is obvious to those skilled in the art that various modifications and variations can be made to the present application without departing from the scope or spirit of the present application. For example, the features illustrated or described as part of one embodiment can be used together with another embodiment to obtain a further embodiment. Therefore, it is intended that the present application covers modifications and variations made within the scope of the attached claims and their equivalents. The above is only a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with the technology in the field within the technical scope disclosed in the present application should be included in the scope of protection of the present application.

Claims

1. A human-computer dialogue oral evaluation system, characterized by include: A speech recognition module, wherein the speech recognition module is configured to recognize the speech input of the student user and convert it into text; An intention understanding module configured to perform semantic understanding on the converted text in combination with a user profile and contextual information to identify the user intention of the student user in the spoken conversation, wherein the user profile includes the user's spoken language proficiency dimension, and the contextual information includes a virtual scene in which the current conversation occurs; A dialogue management module, wherein the dialogue management module is configured to make a corresponding voice response based on the understanding result of the intention understanding module; a language generation module, the language generation module being configured to convert the system actions generated by the dialogue management module into natural language; and a language synthesis module, wherein the language synthesis module is configured to convert natural language into speech and feed the speech back to the student user; Among them, the dialogue management module includes a dialogue state tracking module and a dialogue strategy learning module. When the student user's response exceeds the scope that the intention understanding module can understand, the dialogue state tracking module represents the current dialogue state according to the current stage of the entire dialogue and integrates the context information of the dialogue process. The dialogue strategy learning module adopts a general response strategy according to the current dialogue state and keeps the conversation going by using a virtual person to respond to general statements.

2. The human-computer dialogue oral evaluation system according to claim 1, characterized in that: The intent understanding module is further configured to be able to fill in slots, wherein the slots are information that needs to be completed to convert the user intent into clear user instructions during the conversation.

3. The human-computer dialogue oral evaluation system according to claim 1, characterized in that: The dialogue management module further includes a dialogue state tracking module, which is configured to indicate the stage of the dialogue and integrate context information of the dialogue process.

4. The human-computer dialogue oral evaluation system according to claim 1, characterized in that: It also includes an evaluation system, which includes: a situational dialogue voice and semantic evaluation module, which is configured to be able to perform similarity comparison on the text converted from the user's voice based on the standard content of voice and semantics, and obtain a voice evaluation score and a semantic evaluation score.

5. The human-computer dialogue oral evaluation system according to claim 4, characterized in that: The evaluation system further includes a grammar evaluation and error checking module, which is configured to perform grammar checking on the text converted from the user's speech and obtain a grammar evaluation score.

6. The human-computer dialogue oral evaluation system according to claim 4, characterized in that: The evaluation system further includes an easy-to-mix evaluation module configured to mark easy-to-mix errors in the text converted from the user's voice, so as to perform an easy-to-mix evaluation.

7. The human-computer dialogue oral evaluation system according to claim 4, characterized in that: The dialogue management module is further configured to generate corresponding system actions according to the evaluation results of the evaluation system.

8. The human-computer dialogue oral evaluation system according to claim 4, characterized in that: The higher the similarity between the user's voice and the phonemes of the standard voice, the higher the voice evaluation score; and the higher the similarity between the content expressed by the user and the comparison reference answer, the higher the semantic evaluation score.

9. The human-computer dialogue oral evaluation system according to claim 5, characterized in that: The grammar evaluation and error checking module is further configured to examine the logical relationships in the sentence, wherein the logical relationships include one or more of the following relationships: subject-verb collocation, tense expression, syntactic structure, singular and plural.

10. The human-computer dialogue oral evaluation system according to claim 1, characterized in that: The human-computer dialogue oral evaluation system is a computer system based on a stand-alone and / or online configuration to carry out the evaluation of language content.

Citation Information

Patent Citations

  • Method and apparatus for smart man-machine chat based on artificial intelligence

    CN105094315A

  • Intelligent human-computer interaction method drove by voice

    CN105513593A

  • Data processing method and device for dialogue interaction system

    CN106557464A

  • Method and device used for oral proficiency assessment, electronic device and medium

    CN109785698A