Response state identification method and device, equipment, storage medium and product

By obtaining multimodal data during user learning and multi-dimensional response status recognition, the problem that interactive learning devices cannot accurately judge the response status is solved, and the learning quality and efficiency are improved.

CN120470249APending Publication Date: 2025-08-12HEFEI IFLYTEK TOYCLOUD TECH
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510409105.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Existing interactive learning devices cannot accurately judge the user's response status, which makes interactive learning difficult to carry out and affects the quality of learning.

Method used

By obtaining the image and audio data of the user during the learning process, extracting the response delay time, mouth, hand, eye movement data and response voice data, combining historical interactive learning data and the question types of the current learning content, multi-dimensional response status recognition is performed, and the simulation teacher performs accurate response status detection.

Benefits of technology

Real-time, effective and accurate response status recognition is achieved, the quality and efficiency of interactive learning is improved, and the consistency and pertinence of the learning process is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470249A_ABST
    Figure CN120470249A_ABST
Patent Text Reader

Abstract

The invention provides a response state recognition method and device, equipment, a storage medium and a product, and relates to the technical field of data processing, and the method comprises the steps: obtaining feedback data which is generated after a target object learns the current learning content; performing response data extraction according to the feedback data to obtain target response data; and performing response state identification according to the target response data to obtain a response state identification result of the target object. According to the method, the target response data capable of comprehensively reflecting the response state characteristics is acquired in real time to simulate a teacher to recognize the response state, so that the response state of the user is effectively and accurately detected in real time, the user is better guided to learn, and the interaction quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a response status identification method, device, equipment, storage medium and product. Background Art

[0002] With the rapid development of artificial intelligence technology, interactive learning methods are constantly emerging. These methods use human-computer interaction devices to guide users in interactive learning, using a teacher-led approach to guide students. The simple presentation of content and the use of objects to guide learning are instrumental in helping beginner English learners quickly and easily begin learning.

[0003] However, these devices often need to accurately determine the user's response to the displayed content in order to timely update the teaching guidance tasks and better guide students' learning. If the device cannot accurately determine the user's response status, interactive learning will be difficult to carry out. Therefore, how to effectively and accurately identify the user's response status is an important topic that needs to be studied urgently. Summary of the Invention

[0004] The present invention provides a response status identification method, device, equipment, storage medium and product to solve the defects in the prior art.

[0005] The present invention provides a response status identification method, comprising: Acquiring feedback data, wherein the feedback data is generated by the target subject after learning the current learning content; Extracting response data based on the feedback data to obtain target response data; Response status recognition is performed according to the target response data to obtain a response status recognition result of the target object.

[0006] According to a response status identification method provided by the present invention, the feedback data includes image data and audio data; Extracting response data based on the feedback data to obtain target response data includes: extracting a response delay time based on the image data and the audio data; Extracting response action data based on the image data, wherein the response action data includes at least one of mouth action data, hand action data, and eye action data; Extracting answer voice data based on the audio data; The target response data is acquired according to the response delay time, the response action data and the response voice data.

[0007] According to a response status recognition method provided by the present invention, extracting the response delay time based on the image data and the audio data includes: Perform response start time detection according to the image data to obtain a first response start time; Performing response start time detection according to the audio data to obtain a second response start time; The response delay time is determined according to the first response start time, the second response start time, and the response request time corresponding to the current learning content.

[0008] According to a response status recognition method provided by the present invention, performing response status recognition based on the target response data to obtain a response status recognition result of the target object includes: Obtaining the priority of each dimension of the target response data according to the historical interactive learning data of the target object and / or the topic type of the current learning content; According to the priority level, the response status of the response data of each dimension is sequentially identified to obtain the response status identification result.

[0009] According to a response status identification method provided by the present invention, obtaining the priority level of response data of each dimension in the target response data based on the historical interactive learning data of the target object and / or the topic type of the current learning content includes: Acquiring the response preference characteristics of the target object based on the historical interactive learning data; updating the initial level configuration mode corresponding to the target response data according to the response preference feature and / or the question type to obtain an updated level configuration mode; According to the updated level configuration mode, priority level configuration is performed on the response data of each dimension respectively to obtain the priority level of the response data of each dimension.

[0010] According to a response status identification method provided by the present invention, the response status identification is performed on the response data of each dimension in sequence according to the priority level to obtain the response status identification result, including: Determining, according to the modality type of the response data in each dimension and the question type, standard response data corresponding to the response data in each dimension in the standard response set corresponding to the current learning content; According to the priority level, the response data of each dimension is matched with the standard response data corresponding to the response data of each dimension in turn until all response data are matched, or the response data of any dimension is matched with the standard response data corresponding to the response data of any dimension; Response status recognition is performed according to the matching result to obtain the response status recognition result.

[0011] According to a response state identification method provided by the present invention, determining, based on the modality type of the response data in each dimension and the question type, standard response data corresponding to the response data in each dimension in a standard response set corresponding to the current learning content, includes: According to the modality type of the response data in each dimension, respectively matching the standard response set to obtain a standard response subset of the response data in each dimension; According to the question type, standard response data corresponding to the response data in each dimension are obtained by matching the standard response subsets of the response data in each dimension.

[0012] The present invention also provides a response status identification device, comprising: A data collection unit is used to obtain feedback data, wherein the feedback data is generated by the target subject after learning the current learning content; a processing unit, configured to extract response data based on the feedback data to obtain target response data; The identification unit is used to perform response status identification according to the target response data to obtain a response status identification result of the target object.

[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any one of the response status identification methods described above is implemented.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements any of the above-mentioned response status identification methods when executed by a processor.

[0015] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned response status identification methods.

[0016] The response status recognition method, device, equipment, storage medium and product provided by the present invention obtain the feedback data generated by the target object after learning the current learning content in real time, and extract the response data based on the feedback data to obtain target response data that can fully reflect the response status characteristics. On the basis of the target response data, the teacher is simulated to perform response status recognition, so as to perform real-time, effective and accurate user response status detection, so as to better reflect the accuracy of the user's response to the learning content that he does not know, thereby better guiding the user's learning and improving the quality of interaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 This is one of the flow charts of the response status identification method provided by the present invention.

[0019] Figure 2 This is one of the flow charts of a specific example of the response status identification method provided by the present invention.

[0020] Figure 3 It is a schematic diagram of the process of obtaining target response data provided by the present invention.

[0021] Figure 4 This is the second flow chart of a specific example of the response status identification method provided by the present invention.

[0022] Figure 5 It is a schematic diagram of the process of obtaining the response status recognition result provided by the present invention.

[0023] Figure 6 It is a structural diagram of the response status identification device provided by the present invention.

[0024] Figure 7 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0025] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0026] In current interactive learning scenarios, such as English interactive learning scenarios, English teaching hardware and application programs (APPs) are usually used as important tools to assist English learning. These technical products generally adopt an interactive learning model in which users actively memorize and speak to complete user guided learning. For example, they use memory curves and listening and speaking to help users memorize words, play English sentences and words, and receive users' spoken audio to guide users to speak, and guide users to express themselves correctly in English. However, they usually lack effective guidance and interaction, which makes it difficult for users to learn, thereby affecting the quality of interactive learning.

[0027] In this regard, in order to avoid learning difficulties caused by the lack of effective guidance and interaction among students and to further improve the quality of interactive learning, relevant technologies propose to guide users to conduct interactive learning through human-computer interactive devices (or intelligent bodies). Specifically, the device simulates the teacher, the user represents the student, and the device screen simulates the blackboard, so as to guide teaching and students' learning in the way of a teacher. The content display is relatively simple, and learning is guided by objects, which helps zero-based English learners to quickly and easily enroll.

[0028] However, such devices often need to accurately determine the user's response to the displayed content in order to timely update the teaching guidance tasks and better guide students' learning. If the device cannot accurately determine the user's response status, interactive learning will be difficult to carry out, which will in turn affect the quality of interactive learning. Therefore, how to effectively and accurately identify the user's response status to improve the quality of interactive learning is an important topic that needs to be studied urgently.

[0029] To this end, this application provides a response state recognition method that effectively and accurately identifies the user's response state, thereby improving the quality of interactive learning. This method can be applied to various interactive learning scenarios, including but not limited to language teaching (such as English and Chinese), subject teaching (such as mathematics and biology), and vocational training. This embodiment does not specifically limit this.

[0030] In addition, it should be noted that the execution subject of the response status recognition method provided in this application can be an interactive learning device (hereinafter also referred to as an interactive learning agent or device), such as a repeater, listening and speaking treasure, word treasure, etc., and this application does not make specific limitations on this.

[0031] Figure 1 This is one of the flow charts of the response status identification method provided by the present invention, such as Figure 1 As shown, the method includes step 110 , step 120 and step 130 .

[0032] Step 110 : Acquire feedback data, where the feedback data is generated by the target object after learning the current learning content.

[0033] The target audience here is users who need to learn interactively.

[0034] The current learning content herein may be the knowledge points that the target subject needs to master, as currently displayed on the interactive learning device's interface. This includes, but is not limited to, language instruction, subject instruction, or vocational training content, which is not specifically limited in this embodiment. The current learning content may be automatically generated by the interactive learning device according to a preset teaching plan, selected by the target subject, or automatically generated by the interactive learning device based on response state recognition results corresponding to feedback data generated by the target subject after a previous round of learning, etc., which is not specifically limited in this embodiment.

[0035] Optionally, during the interactive learning process, after the interactive learning device displays the current learning content, feedback data generated by the target subject after learning the current learning content can be collected in real time by a data collection device built into or external to the interactive learning device. This feedback data is generated by the target subject after learning the current learning content. For example, this data can be single-modal data, specifically image data, audio data, or text data, or multi-modal mixed data, specifically multiple of image data, audio data, and text data, etc. This embodiment does not specifically limit this.

[0036] Step 120: Extract response data based on the feedback data to obtain target response data.

[0037] Optionally, after obtaining the feedback data, the interactive learning device will extract effective information that can characterize the response state. These feedback data come in various forms, among which image data includes the target object's eye, mouth, hand and head movement data. For example, eye tracking technology can accurately record the eye's gaze point, scanning path and blinking frequency. These indicators can reflect the user's concentration and cognitive processing load. Oral movement data assists in judging the user's fluency and accuracy in language expression by analyzing features such as mouth opening and closing and lip movement. Hand movement data involves micro-movements of fingers, frequency and amplitude of gestures. This information is particularly important when users perform operational tasks and can reflect their proficiency and accuracy in operation. Head movement data includes movements such as nodding, shaking head, and head tilting. These movements are often related to the user's engagement and emotional response.

[0038] Furthermore, the response voice data within the audio data is also a key component. It includes characteristics such as pitch, intensity, timbre, and speaking speed. These acoustic parameters can reveal the user's emotional state and the confidence level of their response. By extracting and integrating this multimodal data, the system can construct a target data model that comprehensively reflects the response status from multiple levels and dimensions. This not only helps accurately detect the user's immediate reaction state, but also enables in-depth analysis of the user's mastery and depth of understanding of difficult content.

[0039] In practical applications, for example, in interactive learning scenarios, when a user faces a difficult math problem, the system can comprehensively analyze the user's various movement and voice data during the problem-solving process to determine whether the user is thinking, confused, or has mastered the solution. Based on these judgments, the interactive learning device can intelligently adjust the learning progress and provide users with more targeted learning resources and guidance, significantly improving learning efficiency and quality, ensuring the consistency and effectiveness of the learning process, and enabling users to steadily progress at a pace that suits them.

[0040] Therefore, after obtaining the feedback data, the eye movement data, mouth movement data, hand movement data, head movement data, etc. of the target object presented in the image data in the feedback data, as well as the response voice data of the target object presented in the audio data, etc. can be extracted, so as to extract response data of multiple dimensions including eye movement data, mouth movement data, hand movement data, head movement data and response voice data from the image data level and the audio data level, thereby obtaining target response data that can comprehensively reflect the response status, so as to better detect the user's response status (reaction status) in the future, thereby better reflecting the accuracy of the user's response to the content he does not know, so that the interactive learning device can smoothly advance the interactive learning progress and improve the quality of interactive learning.

[0041] Step 130: Perform response status recognition based on the target response data to obtain a response status recognition result of the target object.

[0042] Optionally, after obtaining the target response data, response status identification can be performed based on the target response data to determine whether the target object has effectively responded to the current learning content, that is, whether it has completely learned (mastered) the current learning content, thereby obtaining the target object's response status identification result, so as to subsequently help the interactive learning device better judge the target object's learning status and understanding level, thereby adjusting the teaching strategy to improve the quality of interactive learning.

[0043] The response status recognition results here include a valid response status or an invalid response status; the valid response status is status data used to represent that the target object can correctly and promptly answer questions raised by the device or complete corresponding tasks, that is, status data used to represent that the target object has completely learned (mastered) the current learning content; the invalid response status is status data used to represent that the target object gives an incorrect answer, an incomplete answer, has no response for a long time, or shows hesitation, that is, status data used to represent that the target object has not completely learned the current learning content.

[0044] Among them, when performing response status identification, the response data of each dimension in the target response data can be scored, and the response scores of the multi-dimensional response data can be summarized to obtain the total response score. The total response score is then compared with the score range corresponding to the valid response state or the invalid response state to determine whether the response status identification result of the target object is a valid response state or an invalid response state; or, according to the target priority level, the response status of each dimension response data in the target response data is identified in turn. If any dimension response data is obtained to match its corresponding standard response data during the matching process, the response status identification is stopped, and the response status identification result of the target object is determined to be a valid response state. If all dimension response data are matched, and it is detected that all dimension response data do not match their corresponding standard response data, the response status identification result of the target object is determined to be an invalid response state.

[0045] The target priority level can be generated by fixedly configuring the response data of each dimension in the target response data based on the initial level configuration mode configured by default in the interactive learning device, or it can be generated by dynamically configuring the response data of each dimension in the target response data based on the target level configuration mode determined by the historical interactive learning data of the target object and / or the question type of the current learning content, etc. This embodiment does not make specific limitations on this.

[0046] The method provided in this embodiment obtains the feedback data generated by the target object after learning the current learning content in real time, and extracts the response data based on the feedback data to obtain target response data that can comprehensively reflect the response status characteristics. On the basis of the target response data, the method simulates the teacher to perform response status recognition, so as to perform real-time, effective and accurate user response status detection, so as to better reflect the accuracy of the user's response to the learning content that he does not know, thereby better guiding the user's learning and improving the quality of interaction.

[0047] Based on the response status identification method provided in the above embodiment, a specific embodiment of the response status identification method is given below. Figure 2 This is one of the flow charts of a specific example of the response status identification method provided in this application; Figure 2 As shown, this example specifically includes step 210 , step 220 and step 230 .

[0048] Step 210 : Acquire feedback data generated by the target object after learning the current learning content, where the feedback data includes image data and audio data.

[0049] Optionally, during the feedback data collection process, an image collector and an audio collector can be used to synchronously collect feedback data generated by the target subject after learning the current learning content to obtain corresponding image data and audio data, and the image data and audio data are integrated to form feedback data. The image data here is formed by collecting images of the target subject's face and body movements, etc. The audio data here is formed by collecting the target subject's voice.

[0050] The image collector herein may be a webcam, a camera, or the like, which may be installed on the interactive learning device or may be external to the interactive learning device, as long as it maintains a communication connection with the interactive learning device. The audio collector may be a sound pickup, such as a microphone, which may be installed on the interactive learning device or on headphones connected to the interactive learning device, and this embodiment does not specifically limit this.

[0051] Step 220: Perform multimodal response data extraction on the feedback data to obtain target response data.

[0052] Figure 3 FIG. 1 is a flow chart of target response data acquisition provided by the present invention; FIG. Figure 3 As shown, in a possible implementation, the steps of performing multimodal response data extraction on feedback data to obtain target response data specifically include: step 211 , step 212 , step 213 and step 214 .

[0053] Step 211: Extract the response delay time based on the image data and the audio data.

[0054] Optionally, after obtaining feedback data containing image data and audio data, the response delay time can be extracted based on the image data and audio data. The response delay time here refers to the length of the time interval from allowing the target object to respond (that is, issuing a response request) to the target object officially starting to respond.

[0055] It should be noted that in the response delay time extraction process, the sample image data and sample audio data generated after the sample object learns the sample learning content, as well as the response delay time label corresponding to the sample object, can be used to perform supervised training on the machine learning model to construct a delay time extraction model, so that the image data and audio data can be used to perform corresponding response delay time extraction through the delay time extraction model; or, the response start time of the image data and audio data can be directly detected, so that the response delay time can be obtained by statistically calculating the detected response start time and the response request time corresponding to the current learning content, etc. This embodiment does not specifically limit this.

[0056] In one possible implementation, the step of extracting the response delay time based on the image data and the audio data specifically includes: detecting the response start time based on the image data to obtain a first response start time; detecting the response start time based on the audio data to obtain a second response start time; and determining the response delay time based on the first response start time and the second response start time, as well as the response request time corresponding to the current learning content.

[0057] Optionally, when extracting the response delay time, the image data may be detected to determine the start time of the target subject's response, thereby obtaining the first response start time. For example, when detecting the image data, the timestamp corresponding to the first frame of the image in which the target subject's lips are detected to open, eyes are detected to move to the answer confirmation area, or a response gesture is detected is recorded as the first response start time.

[0058] In addition, the audio data is detected to determine the start time of the target object's response, thereby obtaining a second response start time. For example, when detecting the audio data, the timestamp corresponding to the first frame of audio signal detected in the audio data exceeding a preset volume threshold is recorded as the second response start time.

[0059] After obtaining the first and second response start times, the time interval between the earliest of the first and second response start times and the response request time can be calculated to obtain the response delay time. Thus, by combining image and audio data for comprehensive response delay time detection, the accuracy of response delay time detection is improved, allowing subsequent auxiliary devices to generate more accurate response status recognition results.

[0060] Step 212: extracting response action data based on the image data, wherein the response action data includes at least one of mouth action data, hand action data, and eye action data.

[0061] Furthermore, the target subject's mouth, hand, and eye movement data can be extracted from the image data to obtain features at multiple levels of movement, thereby forming the target subject's response movement data. This allows the auxiliary device to fully understand the target subject's behavioral characteristics associated with their response state after learning the content, thereby better determining their response state. The mouth movement data includes at least data on changes in the target user's mouth shape; the hand movement data includes at least data on changes in the target subject's hands in front of their face; and the eye movement data includes at least data on changes in the target subject's eyes.

[0062] Step 213: extracting response voice data based on the audio data.

[0063] In addition, the response voice data of the target object can also be extracted from the audio data to obtain the response voice data of the target object. The response voice data here refers to voice data containing at least one keyword; the keyword here is a word associated with the response status, such as "um", "hum", etc., and this embodiment does not make specific limitations on this.

[0064] Step 214: Acquire the target response data according to the response delay time, the response action data, and the response voice data.

[0065] After obtaining response data at multiple levels, namely response delay time, response action data and response voice data, the target response data can be determined jointly with the response delay time, response action data and response voice data, thereby ensuring that the target response data contains comprehensive and diverse feature data that can characterize the response status, thereby ensuring the comprehensiveness and accuracy of response status recognition.

[0066] Step 230 , performing response status recognition based on the target response data to determine whether the target object has made a valid response to the current learning content, thereby obtaining a response status recognition result of the target object.

[0067] Optionally, after obtaining the target response data, response status identification can be performed based on the target response data to determine whether the target object has effectively responded to the current learning content, that is, whether it has completely learned the current learning content, thereby obtaining the target object's response status identification result, so as to subsequently help the interactive learning device better judge the target object's learning status and understanding level, thereby adjusting the teaching strategy to improve the quality of interactive learning.

[0068] Among them, when performing response status identification, the response data of each dimension in the target response data can be scored, and the response scores of the multi-dimensional response data can be summarized to obtain the total response score. The total response score is then compared with the score range corresponding to the valid response state or the invalid response state to determine whether the response status identification result of the target object is a valid response state or an invalid response state; or, according to the target priority level, the response status of each dimension response data in the target response data is identified in turn. If any dimension response data is obtained during the matching process and matches its corresponding standard response data, the response status identification is stopped, and the response status identification result of the target object is determined to be a valid response state. If all dimension response data are matched, and it is detected that all dimension response data do not match their corresponding standard response data, the response status identification result of the target object is determined to be an invalid response state, etc. This embodiment does not specifically limit this.

[0069] In this embodiment, the image data and audio data generated by the target object after learning the current learning content are obtained in real time, and the response delay time is extracted in combination with the image data and audio data, the response action data is extracted based on the image data, and the response voice data is extracted based on the audio data. The target response data is obtained by integrating the response delay time, response action data and response voice data, thereby ensuring that the target response data contains comprehensive and diverse feature data that can characterize the response status. Therefore, on the basis of the target response data, the teacher can be simulated to perform comprehensive and accurate response status recognition, so as to better reflect the accuracy of the user's response to the learning content that he does not know, better guide user learning, and improve the quality of interaction.

[0070] Figure 4 This is a flow chart of a specific example of the response status identification method provided by the present invention; Figure 4 As shown, in order to improve the response status recognition effect, in another specific embodiment, the response status of each dimension of the target response data can be recognized in combination with the priority level of each dimension of the target response data. Figure 4 As shown, this example specifically includes step 410 , step 420 and step 430 .

[0071] Step 410: Obtain feedback data generated by the target object after learning the current learning content.

[0072] Optionally, during the interactive learning process, after the interactive learning device displays the current learning content, the data acquisition device can be used to collect in real time the feedback data generated by the target object after learning the current learning content. The data acquisition device here may include an image collector and / or an audio collector; the image collector may be a camera, etc., which may be installed on the interactive learning device or may be external to the interactive learning device and set up separately, and it only needs to maintain a communication connection with the interactive learning device. The audio collector may be a pickup, such as a microphone, which may be installed on the interactive learning device or on headphones connected to the interactive learning device, etc., and this embodiment does not make specific restrictions on this. Accordingly, the feedback data may be single-modal data, specifically image data or audio data, or multi-modal mixed data, specifically mixed data of image data and audio data, etc., and this embodiment does not make specific restrictions on this.

[0073] Step 420: Extract multimodal response data based on the feedback data to obtain target response data.

[0074] After obtaining the feedback data, the eye movement data, mouth movement data, hand movement data, head movement data, etc. of the target object presented in the image data in the feedback data, as well as the response voice data of the target object presented in the audio data, etc. can be extracted. The response delay time can also be extracted from the time information presented in the image data and the time information presented in the audio data. Thus, response data of multiple dimensions including eye movement data, mouth movement data, hand movement data, head movement data, response voice data and response delay time are extracted from the image data level and the audio data level, that is, target response data that can fully reflect the response status is obtained, so as to better detect the user's response status in the future, thereby better reflecting the accuracy of the user's response to the content he does not know, so that the interactive learning device can smoothly advance the interactive learning progress and improve the quality of interactive learning.

[0075] Step 430 , performing response status recognition on the target response data to determine whether the target object has effectively responded to the current learning content, thereby obtaining a response status recognition result of the target object.

[0076] Figure 5 FIG. 1 is a flow chart of obtaining the response status recognition result provided by the present invention; FIG. Figure 5 As shown, in a possible implementation, the specific steps of performing multimodal response data extraction on feedback data to obtain target response data include: step 431 and step 432.

[0077] Step 431 : Obtain the priority level of each dimension of the target response data according to the historical interactive learning data of the target object and / or the topic type of the current learning content.

[0078] The historical interactive learning data here refers to the data generated by the target object and the device during the historical interactive learning process, including but not limited to historical learning content, historical response data and historical response status recognition results, as well as historical response patterns, to assist the device in better learning the response preference characteristics of the target object, such as whether the target object prefers to use response voice data to respond to learning content, or prefers to use mouth movement data, hand movement data, or eye movement data to respond to learning content during the historical interactive learning process.

[0079] The question type here is used to distinguish the category of the question in the current learning content, including but not limited to multiple-choice questions, oral questions, etc., which are not specifically limited in this embodiment. The response data of different dimensions have different importance for the responses corresponding to different question types.

[0080] Therefore, when performing response status recognition, the priority of each dimension of response data in the target response data can be determined based on the historical interactive learning data of the target object and / or the question type of the current learning content, so that the response status of each dimension of response data can be recognized in turn according to the execution order of the determined priority to obtain the response status recognition result; for example, when the response voice data, response delay time, oral movement data, hand movement data and eye movement data are prioritized in a mode of descending priority, in the response status recognition process, the response voice data, response delay time, oral movement data, hand movement data and eye movement data can be recognized in turn in the order of descending priority to obtain the response status recognition result.

[0081] It should be noted that when determining the priority, the priority of the response data of each dimension can be configured according to the priority configuration mode directly determined based on the historical interactive learning data of the target object and / or the topic type of the current learning content, or the priority of the response data of each dimension can be configured after updating part of the configuration content or all of the configuration content in the preset priority configuration mode (such as the priority configuration mode determined in the last interaction, or the default priority configuration mode of the device) based on the historical interactive learning data of the target object and / or the topic type of the current learning content, etc. This implementation does not make specific restrictions on this.

[0082] In one possible implementation, the step of obtaining the priority level of the response data of each dimension in the target response data based on the historical interactive learning data of the target object and / or the question type of the current learning content specifically includes: obtaining the response preference characteristics of the target object based on the historical interactive learning data; updating the initial level configuration mode corresponding to the target response data based on the response preference characteristics and / or the question type to obtain an updated level configuration mode; and configuring the priority level of the response data of each dimension based on the updated level configuration mode to obtain the priority level of the response data of each dimension.

[0083] The initial level configuration mode here is the default priority level configuration mode set by the interactive learning device for response data in different modes. For example, in the initial level configuration mode, the response voice data is configured as the highest priority by default, and the response delay time is configured as the second highest priority by default.

[0084] Optionally, when performing priority configuration, it can be specifically determined that the historical interactive learning data is non-empty data, that is, after the target object is determined to be non-empty data, that is, it is not the first time for the target object to interact with the interactive device for learning, then the historical interactive learning data of the target object is analyzed to dig out the target object's response preference characteristics, that is, in the historical interactive learning process, which modal type of response data is used to respond, or which modal type of response data is used to respond to the learning content of the corresponding question type. And, based on the question type, it is determined which modal type of response data is more important for responding to the current learning content; then, the importance of the response data of each modal type can be analyzed based on the response preference characteristics and / or question type, and based on the importance results obtained from the analysis, the initial level configuration mode is updated to obtain a level configuration mode that is adapted to the target object's response habits and / or matches the question type of the current learning content. During the update process, the priority configuration of the response data of some or all modal types in the initial level configuration mode can be modified.

[0085] After obtaining the updated level configuration mode, the priority level of the response data of each dimension can be configured according to the priority level configuration value corresponding to the response data of each dimension in the updated level configuration mode to obtain the priority level of the response data of each dimension. The priority level of the response data of each dimension is dynamically updated based on the historical interactive learning data of the target object and / or the question type of the current learning content, so that the response status recognition step is adapted to the response habits of the target object and / or the question type of the current learning content, effectively realizing efficient, orderly and targeted recognition processing of the response data, so as to improve the accuracy and adaptability of response recognition, and enhance the interaction efficiency and user experience.

[0086] Step 432: According to the priority level, the response status of the response data of each dimension is sequentially identified to obtain the response status identification result.

[0087] Optionally, when identifying the response status of the response data of each dimension, it can be achieved by matching the response data of each dimension with the corresponding standard response data, or it can be achieved by extracting the response features of at least one level of the response data of each dimension or performing response scoring on the response data of each dimension, and then matching the response features or response scores of each level of the response data of each dimension with the corresponding standard response features or standard response score thresholds of each level. This embodiment does not specifically limit this. For example, for the response voice data, sound features can be extracted from the time domain and frequency domain dimensions to match the time domain sound features (such as sound wave features) and frequency domain sound features (spectral features) of the response voice data with the standard time domain response features and standard frequency domain response features of the standard response voice data corresponding to the response voice data, respectively, to achieve response status identification of the response voice data.

[0088] In one possible implementation, according to the priority level, the response data of each dimension are sequentially identified in response status, and the steps of obtaining the response status identification result specifically include: determining the standard response data corresponding to the response data of each dimension in the standard response set corresponding to the current learning content according to the modal type of the response data of each dimension and the question type; according to the priority level, matching the response data of each dimension with the standard response data corresponding to the response data of each dimension in turn until all response data are matched, or the response data of any dimension is matched with the standard response data corresponding to the response data of any dimension; performing response status identification based on the matching result to obtain the response status identification result.

[0089] The modality types here can be specifically divided according to the data dimensions of the response data, such as delay modality, mouth movement modality, hand movement modality, eye movement modality and voice modality.

[0090] The standard response set corresponding to the current learning content here includes standard response data associated with different modal types and different question types; the standard response data can be understood as a reference benchmark or correct version of the response data generated by the user's response, that is, the standard response data shows the standard data in the valid response state corresponding to the response data generated by the user's response.

[0091] Optionally, during the response state identification process, the standard response data corresponding to each dimension response data can be jointly determined in the standard response set corresponding to the current learning content based on the modal type and question type of each dimension response data. Exemplarily, for each dimension response data, the identifier corresponding to the modal type of the dimension response data and the identifier corresponding to the question type can be integrated and encoded to match the standard response data that matches the integrated coding identifier in the standard response set as the standard response data corresponding to the dimension response data; or a preliminary matching search can be performed in the standard response set based on the modal type of the dimension response data to obtain multiple standard response data that match the modal type of the dimension response data, and then a matching search can be performed again in the multiple matching standard response data based on the question type to use the standard response data that further matches the question type as the standard response data corresponding to the dimension response data, etc. This embodiment does not specifically limit this.

[0092] In one possible implementation, the step of determining the standard response data corresponding to the response data in each dimension in the standard response set corresponding to the current learning content according to the modal type of the response data in each dimension and the question type specifically includes: according to the modal type of the response data in each dimension, matching the standard response subset of the response data in each dimension in the standard response set; according to the question type, matching the standard response data corresponding to the response data in each dimension in the standard response subset of the response data in each dimension.

[0093] Optionally, in the process of acquiring standard response data, a preliminary matching search can be performed in the standard response set based on the association relationship between the modal type of each dimension response data and the standard response data to obtain multiple standard response data that match the modal type of each dimension response data, thereby forming a standard response subset of each dimension response data; then, based on the association relationship between the question type and the standard response data, a matching search is performed again in the standard response subset of each dimension response data to further match the standard response data of each dimension response data with the question type as the standard response data corresponding to each dimension response data, thereby accurately screening the standard response of each dimension response data through the dual matching of the modal type and the question type, so as to more accurately identify the user's response status, thereby facilitating the interactive device to adjust the difficulty and progress of the learning content in real time according to the user's response status, so as to provide a personalized learning experience that better meets the user's needs and improve the quality of interactive learning. After obtaining the standard response data corresponding to each dimension response data, the execution order of each dimension response data can be determined according to the priority level, and the following response status identification operations are performed on each dimension response data in sequence according to the execution order: The current dimension response data is matched with its corresponding standard response data. If the current dimension response data matches its corresponding standard response data, the response state identification is stopped, and the response state identification result is determined as a valid response state.

[0094] If the current dimension response data does not match its corresponding standard response data, the matching of the next dimension response data will continue until any response data is obtained that matches its corresponding standard response data during the iterative matching process, or all response data are matched, then the response state identification will be stopped, and it will be determined whether there is any response data that matches its corresponding standard response data during each matching process. If so, the response state identification result will be determined as a valid response state. If not, that is, all response data do not match their corresponding standard response data, then the response state identification result will be determined as an invalid response state.

[0095] It should be noted that, during the matching process, the corresponding matching mode can be selected according to the data type of the response data of each dimension for matching. For example, for the response voice data, which is voice data, the voice recognition and natural language processing technology can be used to analyze and compare the matching degree of content, pronunciation, intonation, etc. between the response voice data and the standard response voice data to determine whether the two match; for the response delay time, which is time data, the time difference between the response delay time and the standard response delay time can be compared to determine whether the two match; for the mouth movement data, hand movement data and eye movement data, which are image data, the image recognition and image processing technology can be used to analyze and compare the matching degree of movement trajectory, movement frequency, movement direction, contour, etc. between such image data and its corresponding standard image data to determine whether the two match; thus, by accurately matching the standard response data according to the modal type and question type of the response data of each dimension, and matching them in order according to priority, the efficiency and quality of response recognition can be effectively improved.

[0096] In this embodiment, by acquiring the feedback data generated by the target object after learning the current learning content in real time, and extracting the response data of at least one modality from the feedback data, it is ensured that the acquired target response data contains comprehensive and diverse feature data that can characterize the response state, and thus, on the basis of the target response data, by relying on the historical interactive learning data of the target object and / or the question type of the current learning content, the priority of each dimension of the response data is set, so as to identify the response state one by one according to the priority order, effectively realizing the efficient, orderly and targeted identification and processing of the response data, so as to obtain the response state identification results more quickly and accurately, better reflect the accuracy of the user's response to the learning content that he does not know, and thus better guide the user's learning and improve the quality of interaction. The response state identification device provided by the present invention is described below, and the response state identification device described below and the response state identification method described above can be referenced to each other.

[0097] Figure 6 Schematic diagram of the structure of a response status identification device provided by the present invention; Figure 6 As shown, the device includes: a data acquisition unit 610, a processing unit 620, and an identification unit 630. The data acquisition unit 610 is used to obtain feedback data, which is generated by the target object after learning the current learning content; the processing unit 620 is used to extract response data based on the feedback data to obtain target response data; and the identification unit 630 is used to identify the response state based on the target response data to obtain the response state identification result of the target object.

[0098] The device provided in this embodiment obtains the feedback data generated by the target object after learning the current learning content in real time, and extracts the response data based on the feedback data to obtain target response data that can fully reflect the response status characteristics. On the basis of the target response data, the device simulates the teacher to perform response status recognition, so as to perform real-time, effective and accurate user response status detection, so as to better reflect the accuracy of the user's response to the learning content that he does not know, thereby better guiding the user's learning and improving the quality of interaction.

[0099] In some embodiments, the feedback data includes image data and audio data; The processing unit is specifically used to: extract the response delay time based on the image data and the audio data; extract the response action data based on the image data, the response action data including at least one of mouth action data, hand action data and eye action data; extract the response voice data based on the audio data; and obtain the target response data based on the response delay time, the response action data and the response voice data.

[0100] In some embodiments, the processing unit is further used to: perform response start time detection based on the image data to obtain a first response start time; perform response start time detection based on the audio data to obtain a second response start time; determine the response delay time based on the first response start time and the second response start time, as well as the response request time corresponding to the current learning content.

[0101] In some embodiments, the identification unit is specifically used to: obtain the priority level of the response data of each dimension in the target response data based on the historical interactive learning data of the target object and / or the question type of the current learning content; according to the priority level, perform response status identification on the response data of each dimension in turn to obtain the response status identification result.

[0102] In some embodiments, the identification unit is further used to: obtain the response preference characteristics of the target object based on the historical interactive learning data; update the initial level configuration mode corresponding to the target response data based on the response preference characteristics and / or the question type to obtain an updated level configuration mode; and configure the priority level of the response data of each dimension according to the updated level configuration mode to obtain the priority level of the response data of each dimension.

[0103] In some embodiments, the identification unit is further used to: determine the standard response data corresponding to the response data in each dimension in the standard response set corresponding to the current learning content according to the modal type of the response data in each dimension and the question type; match the response data in each dimension with the standard response data corresponding to the response data in each dimension in turn according to the priority level until all response data are matched, or the response data in any dimension matches the standard response data corresponding to the response data in any dimension; perform response status identification based on the matching result to obtain the response status identification result.

[0104] In some embodiments, the identification unit is further used to: match the standard response subsets of the response data in each dimension in the standard response set according to the modality type of the response data in each dimension; and match the standard response data corresponding to the response data in each dimension in the standard response subsets of the response data in each dimension according to the question type.

[0105] The device provided by the present invention is used to execute the above-mentioned method embodiments. Please refer to the above-mentioned embodiments for the specific processes and detailed contents, which will not be repeated here.

[0106] Figure 7 An example of a physical structure diagram of an electronic device is shown below. Figure 7As shown, the electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communication bus 740. The processor 710, the communications interface 720, and the memory 730 communicate with each other via the communication bus 740. The processor 710 may invoke logic instructions in the memory 730 to execute a response state identification method, which includes: obtaining feedback data, the feedback data being generated by a target subject after learning current learning content; extracting response data based on the feedback data to obtain target response data; and performing response state identification based on the target response data to obtain a response state identification result for the target subject.

[0107] Furthermore, the logic instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0108] On the other hand, the present invention also provides a computer program product, which includes a computer program, and the computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the response state identification method provided by the above methods, and the method includes: obtaining feedback data, which is generated by the target object after learning the current learning content; extracting response data based on the feedback data to obtain target response data; and performing response state identification based on the target response data to obtain the response state identification result of the target object.

[0109] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it is implemented to execute the response state identification method provided by the above-mentioned methods. The method includes: obtaining feedback data, which is generated by the target object after learning the current learning content; extracting response data based on the feedback data to obtain target response data; and performing response state identification based on the target response data to obtain a response state identification result of the target object.

[0110] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0111] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for identifying a response state, characterized in that: include: Acquiring feedback data, wherein the feedback data is generated by the target subject after learning the current learning content; Extracting response data based on the feedback data to obtain target response data; Response status recognition is performed according to the target response data to obtain a response status recognition result of the target object.

2. The response status identification method according to claim 1, characterized in that: The feedback data includes image data and audio data; Extracting response data based on the feedback data to obtain target response data includes: extracting a response delay time based on the image data and the audio data; Extracting response action data based on the image data, wherein the response action data includes at least one of mouth action data, hand action data, and eye action data; Extracting answer voice data based on the audio data; The target response data is acquired according to the response delay time, the response action data and the response voice data.

3. The response status identification method according to claim 2, characterized in that: The extracting of the response delay time according to the image data and the audio data includes: Perform response start time detection according to the image data to obtain a first response start time; Performing response start time detection according to the audio data to obtain a second response start time; The response delay time is determined according to the first response start time, the second response start time, and the response request time corresponding to the current learning content.

4. The response status identification method according to any one of claims 1 to 3, characterized in that: The step of performing response status identification according to the target response data to obtain a response status identification result of the target object includes: Obtaining the priority of each dimension of the target response data according to the historical interactive learning data of the target object and / or the topic type of the current learning content; According to the priority level, the response status of the response data of each dimension is sequentially identified to obtain the response status identification result.

5. The response status identification method according to claim 4, characterized in that: The obtaining of the priority of each dimension of the target response data according to the historical interactive learning data of the target object and / or the topic type of the current learning content includes: Acquiring the response preference characteristics of the target object based on the historical interactive learning data; updating the initial level configuration mode corresponding to the target response data according to the response preference feature and / or the question type to obtain an updated level configuration mode; According to the updated level configuration mode, priority level configuration is performed on the response data of each dimension respectively to obtain the priority level of the response data of each dimension.

6. The response status identification method according to claim 4, characterized in that: According to the priority level, the response status of the response data of each dimension is sequentially identified to obtain the response status identification result, including: Determining, according to the modality type of the response data in each dimension and the question type, standard response data corresponding to the response data in each dimension in the standard response set corresponding to the current learning content; According to the priority level, the response data of each dimension is matched with the standard response data corresponding to the response data of each dimension in turn until all response data are matched, or the response data of any dimension is matched with the standard response data corresponding to the response data of any dimension; Response status recognition is performed according to the matching result to obtain the response status recognition result.

7. The response status identification method according to claim 6, characterized in that: The determining, based on the modality type of the response data in each dimension and the question type, standard response data corresponding to the response data in each dimension from the standard response set corresponding to the current learning content includes: According to the modality type of the response data in each dimension, respectively matching the standard response set to obtain a standard response subset of the response data in each dimension; According to the question type, standard response data corresponding to the response data in each dimension are obtained by matching the standard response subsets of the response data in each dimension.

8. A response status recognition device, characterized in that: include: A data collection unit is used to obtain feedback data, wherein the feedback data is generated by the target subject after learning the current learning content; a processing unit, configured to extract response data based on the feedback data to obtain target response data; The identification unit is used to perform response status identification according to the target response data to obtain a response status identification result of the target object.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the response status identification method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the response status identification method according to any one of claims 1 to 7 is implemented.

11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the response status identification method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Online learning system based on emotional state

    CN110334626A

  • Video interaction method

    CN112866744A

  • Training method and device based on virtual reality, virtual reality equipment and storage medium

    CN116741010A

  • Online teaching artificial intelligence tutoring method, medium and system

    CN118053331A

  • Campus safety management method based on smart campus and related equipment

    CN118917564A