A multi-modal emotion recognition and intelligent guidance device for English learning

By using a multimodal emotion recognition device, combined with somatosensory data and learning interaction logs, the learned helplessness of users can be identified and intervened, which solves the problem of insufficient identification of deep psychological obstacles in existing English learning systems, and improves learning effectiveness and user motivation.

CN122116473APending Publication Date: 2026-05-29SUZHOU VOCATIONAL UNIVERSITY (SUZHOU OPEN UNIVERSITY)

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU VOCATIONAL UNIVERSITY (SUZHOU OPEN UNIVERSITY)
Filing Date
2026-02-25
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing English learning systems struggle to identify and intervene in users' deep-seated psychological barriers, particularly learned helplessness, and lack the ability to effectively identify and intervene in complex psychological states.

Method used

Employing a multimodal emotion recognition and intelligent intervention device, the system collects users' video stream data and stress distribution data through a somatosensory data collector. Combined with learning interaction logs, it extracts negative interaction behaviors, posture withdrawal, and stress shift features to generate a comprehensive behavioral pattern vector. A psychological risk assessment model is then used to identify learned helplessness risk, and intervention content is automatically generated for intervention.

Benefits of technology

By accurately identifying users' learned helplessness, and automatically generating guidance content for cognitive restructuring and psychological intervention, the effectiveness of English learning and users' learning motivation have been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122116473A_ABST
    Figure CN122116473A_ABST
Patent Text Reader

Abstract

The application discloses a multi-modal emotion recognition and intelligent guidance device for English learning, comprising a learning terminal, a somatosensory data collector and a central processor. The learning terminal is used for presenting learning content and recording user learning interaction logs. The somatosensory data collector is used for collecting user video stream data and stress distribution data. The central processor comprises a multi-modal feature extraction module, a risk assessment module, a guidance decision module and a content generation module. The device extracts negative interaction behavior features from the logs, extracts posture withdrawal features from the video stream and extracts stress deviation features from the stress data. The user's learned helplessness risk level is obtained through a psychological risk assessment model, a target guidance strategy is matched according to the risk level, and guidance content for cognitive reconstruction is generated and embedded in the learning content. The device can more accurately identify deep psychological risks and provide targeted intervention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a multimodal emotion recognition and intelligent guidance device for English learning. Background Technology

[0002] With the deep integration of information technology and education, intelligent English learning systems have been widely used. Existing adaptive learning systems mainly adjust the path and difficulty of knowledge presentation based on explicit learning data such as the user's correctness of answers and response time, aiming to improve the efficiency of knowledge transfer.

[0003] However, as a second language, English differs greatly from Chinese in terms of grammar, pronunciation, and logical thinking. Chinese students are easily corrected for inaccurate pronunciation and grammatical errors. In the advanced stage, they are limited by vocabulary, reading speed, and fluency of spoken English, making it difficult to see rapid progress.

[0004] English learning, especially long-term or high-intensity learning, is not only a cognitive process, but also accompanied by complex emotional and psychological experiences. When users encounter bottlenecks or slow progress, they are prone to feelings of frustration, anxiety, and even fall into a state of learned helplessness. That is, after repeated failures, they form a negative attribution pattern that no matter how hard they try, they will fail, thus losing their motivation to learn. Existing learning systems lack the ability to identify and intervene in such deep and persistent psychological obstacles.

[0005] In addition, advancements in affective computing have made it possible to identify users' real-time emotions through multimodal data (such as facial expressions and voice), and there have been attempts to apply it to educational scenarios.

[0006] However, existing technical solutions mostly focus on short-term, overt basic emotions (such as happiness and sadness), making it difficult to effectively identify complex psychological states such as learned helplessness, which are closely related to long-term behavioral patterns and cognitive attribution. Summary of the Invention

[0007] This application aims to at least partially solve one of the technical problems in the aforementioned technologies.

[0008] To achieve the above objectives, the first aspect of this application proposes a multimodal emotion recognition and intelligent guidance device for English learning, comprising: a learning terminal unit, a somatosensory data acquisition unit, and a central processing unit; the learning terminal unit is used to present learning content and record the user's learning interaction log; the somatosensory data acquisition unit is used to collect video stream data of the user's upper body and pressure distribution data between the user's buttocks and the seat; the central processing unit includes a multimodal feature extraction module, a risk assessment module, a guidance decision module, and a content generation module; the multimodal feature extraction module is used to receive the learning interaction log and extract negative interaction behavior features representing answering behavior based on the same time window, receive the video stream data and extract posture withdrawal features based on the user's head and shoulders, and receive... The system describes the pressure distribution data and extracts pressure shift features that characterize sitting posture stability; the risk assessment module is used to fuse the negative interaction behavior features, the posture withdrawal features, and the pressure shift features to generate a comprehensive behavior pattern vector, and input the comprehensive behavior pattern vector into a preset psychological risk assessment model to obtain the user's learned helplessness risk level; the guidance decision module is used to select a target guidance strategy from a preset strategy library based on the current risk level and the posture withdrawal features and / or the pressure shift features when the learned helplessness risk level exceeds a preset level threshold; the content generation module is used to generate guidance content for cognitive reconstruction of the user based on the target guidance strategy, and control the learning terminal to embed the guidance content into the learning content for presentation.

[0009] In addition, the multimodal emotion recognition and intelligent guidance device for English learning proposed in this application may also have the following additional technical features:

[0010] As a further description of the above technical solution: the multimodal feature extraction module includes a visual feature extraction unit; the visual feature extraction unit is used to perform human head and shoulder region localization on video stream data, and obtain a two-dimensional image coordinate sequence of the head center point, left shoulder point and right shoulder point; calculate the shoulder line angle formed by the left shoulder point and the right shoulder point, and the vertical offset of the head center point relative to the midpoint of the shoulder line; within the time window, calculate the variance of the shoulder line angle and the mean of the vertical offset, and use the weighted sum of the two as the posture retreat feature.

[0011] As a further description of the above technical solution: the multimodal feature extraction module also includes a pressure feature extraction unit, which is used to calculate the average pressure difference between the left and right sides of the user's seat based on the pressure distribution data within the time window, as a first pressure feature characterizing the shift of the body's center of gravity and the asymmetry of the sitting posture; calculate the standard deviation of the overall pressure value fluctuating over time within the time window, as a second pressure feature characterizing the user's body tension; and combine the first pressure feature and the second pressure feature to constitute the pressure offset feature.

[0012] As a further description of the above technical solution: the multimodal feature extraction module also includes a log feature extraction unit, which is used to extract at least two types of negative interaction behavior features based on the learning interaction logs within the time window. The negative interaction behavior features are used to quantitatively characterize the avoidance tendency and hesitation tendency shown by the user during the question-answering process.

[0013] As a further description of the above technical solution: the risk assessment module includes a feature fusion unit and a model calculation unit, wherein the feature fusion unit is used to normalize the posture withdrawal feature, the stress offset feature, and the negative interaction behavior feature respectively, and concatenate them into the comprehensive behavior pattern vector; the model calculation unit is used to input the comprehensive behavior pattern vector into the preset psychological risk assessment model, the psychological risk assessment model is used to output the predicted probability of the user falling into a learned helplessness psychological state; the learned helplessness risk level is divided according to the interval in which the predicted probability falls.

[0014] As a further description of the above technical solution: the step of selecting a target guidance strategy from a preset strategy library based on the current risk level and the posture retreat feature and / or the pressure shift feature specifically includes: determining whether the posture retreat feature exceeds a first somatosensory threshold; determining whether the pressure shift feature exceeds a second somatosensory threshold; constructing a combination item based on the learned helplessness risk level and the comparison result of the posture retreat feature and the pressure shift feature; and matching a unique target guidance strategy from the preset strategy library based on the combination item.

[0015] As a further description of the above technical solution: the generation of the guidance content includes: based on the target guidance strategy, retrieving historical answer records that meet preset success conditions from the user's historical learning records, and generating review guidance text to guide the user to review the historical answer records; based on the learning interaction log, generating visual information reflecting the user's long-term ability improvement trend; based on the knowledge points and difficulty of the user's current learning task, generating task decomposition guidance text that breaks down the learning task into multiple sub-steps; wherein, at least one of the review guidance text, the visual information, and the task decomposition guidance text is integrated into the guidance content.

[0016] As a further description of the above technical solution: the somatosensory data acquisition device includes a pressure sensing module installed on the user's seat and a visual acquisition module oriented towards the user.

[0017] As a further description of the above technical solution: the pressure sensing module includes a flexible pad, pressure sensors, elastic straps, a mounting compartment, and anti-slip strips. Multiple pressure sensors are arranged in two groups and embedded on the left and right sides of the top surface of the flexible pad. The elastic straps are located on the sides of the flexible pad to bypass the seat and be locked in place by buckles. The mounting compartment is detachably located on the rear side of the flexible pad, and the signal processing chip and data interface of the pressure sensors are integrated within the mounting compartment. The anti-slip strips are located at the bottom of the flexible pad.

[0018] As a further description of the above technical solution: the visual acquisition module includes a mounting base, a telescopic rod, a light-shielding and image-stabilizing compartment, and a camera, wherein the mounting base is detachably mounted on the learning terminal; the telescopic rod is rotatably mounted on the mounting base and locked by a damping knob; the light-shielding and image-stabilizing compartment is rotatably mounted on the telescopic rod and locked by another damping knob; the camera is mounted inside the light-shielding and image-stabilizing compartment.

[0019] The multimodal emotion recognition and intelligent guidance device for English learning proposed in this application constructs a multi-dimensional comprehensive behavioral pattern vector by integrating negative interaction behavior features, posture withdrawal features, and stress shift features. This overcomes the one-sidedness of judging solely based on correct or incorrect answers or single facial expressions, and can more accurately and reliably identify deep psychological risks related to long-term behavioral patterns and physiological states, such as "learned helplessness." Unlike simple emotional comfort, this device can automatically generate guidance content for cognitive restructuring after identifying psychological risks, guiding users to recall small successes, etc. By displaying objective progress curves and providing task decomposition guidance, interventions are directly targeted at the core cognitive biases of "learned helplessness." Visual analysis of the user's head and shoulder posture withdrawal characteristics captures the user's unconscious body language under frustration. Pressure sensors detect pressure shift characteristics in the hip pressure distribution, quantifying the instability of sitting posture and frequent shifts in the center of gravity caused by anxiety and tension. These somatosensory characteristics, as physiological behavioral indicators, are highly correlated with psychological state and are not easily concealed by the user's subjective will, thus greatly enriching the dimensions of emotion recognition and improving the robustness and accuracy of recognition.

[0020] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0021] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0022] Figure 1 This is a system block diagram of a multimodal emotion recognition and intelligent guidance device for English learning according to an embodiment of this application;

[0023] Figure 2 This is a flowchart of the visual feature extraction unit in a multimodal feature extraction module according to an embodiment of this application;

[0024] Figure 3 This is a flowchart of the pressure feature extraction unit in a multimodal feature extraction module according to an embodiment of this application;

[0025] Figure 4 This is a flowchart of the log feature extraction unit in a multimodal feature extraction module according to an embodiment of this application.

[0026] Figure 5 This is a schematic diagram of a multimodal emotion recognition and intelligent guidance device for English learning according to an embodiment of this application;

[0027] Figure 6This is a schematic diagram of the structure of a pressure sensor module according to an embodiment of this application;

[0028] Figure 7 This is a schematic diagram of the structure of a pressure sensor module according to another embodiment of this application;

[0029] Figure 8 This is a schematic diagram of the structure of a visual acquisition module according to an embodiment of this application;

[0030] As shown in the figure: 100, pressure sensor module; 110, flexible pad; 120, pressure sensor; 130, elastic strap; 131, buckle; 140, mounting compartment; 150, anti-slip strip; 200, vision acquisition module; 210, mounting base; 211, hook; 212, support pad; 213, screw; 214, throttle; 220, telescopic rod; 230, light-shielding and image-stabilizing compartment; 240, camera; 300, main unit box. Detailed Implementation

[0031] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0032] The following description, in conjunction with the accompanying drawings, describes an embodiment of the multimodal emotion recognition and intelligent guidance device for English learning.

[0033] like Figure 1 As shown in the figure, the multimodal emotion recognition and intelligent guidance device for English learning according to the embodiments of this application includes a learning terminal unit, a somatosensory data collector, and a central processing unit.

[0034] The learning terminal includes a display screen and input devices, used to present learning content and record the user's learning interaction logs.

[0035] The motion data acquisition device includes pressure sensing modules located on the left and right sides of the user's seat and a vision acquisition module facing the user, used to collect pressure distribution data between the user's buttocks and the seat and video stream data of the user's upper body, respectively.

[0036] The central processing unit includes a multimodal feature extraction module, a risk assessment module, a guidance decision module, and a content generation module.

[0037] The multimodal feature extraction module is used to receive learning interaction logs and extract negative interaction behavior features representing answering behavior based on the same time window, receive video stream data and extract posture withdrawal features based on the user's head and shoulders, and receive pressure distribution data and extract pressure offset features representing sitting posture stability.

[0038] It should be noted that by analyzing users' performance in answering questions, body posture, and sitting posture changes during the same period, this multi-dimensional and synchronous data enables the system to accurately determine whether users are experiencing complex psychological states such as learned helplessness, rather than simply judging emotional fluctuations.

[0039] For clarity, in the embodiments of this application, such as Figure 2 As shown, the multimodal feature extraction module includes a visual feature extraction unit.

[0040] The visual feature extraction unit is used to perform human head and shoulder region localization on video stream data, and obtain two-dimensional image coordinate sequences of the head center point, left shoulder point and right shoulder point, without performing facial recognition, in order to enhance privacy protection.

[0041] Calculate the angle of the shoulder line formed by the left and right shoulder points, and the vertical offset of the head center point relative to the midpoint of the shoulder line.

[0042] Within the time window, the variance of the shoulder line angle change and the mean of the vertical offset are calculated, and the weighted sum of the two is used as the posture retreat feature.

[0043] Specifically, within the time window, the visual feature extraction unit uses a face detection model (e.g., MTCNN) and a human pose estimation algorithm (e.g., MediaPipe Pose) to locate and obtain the two-dimensional pixel coordinates of the user's head center point, left shoulder point, and right shoulder point in the image coordinate system for each sampling moment. Based on the obtained key point coordinates, it calculates the angle between the line segment formed by the left and right shoulder points and the horizontal direction of the image. Based on the relative position of the head center point coordinates and the midpoint of the shoulder line, it calculates the vertical offset of the head.

[0044] For example, setting a time window During this period, the visual acquisition module operates at a fixed frame rate. Continuously capture the user's video stream, therefore, in this window Within, a total of Each frame of an image corresponds to a precise timestamp. .

[0045] Using existing face detection and human pose estimation algorithms, output the coordinates of the user's left shoulder in this frame. Coordinates of the right shoulder point and the coordinates of the head center point .

[0046] Calculate the first shoulder coordinate based on the left and right shoulder coordinates. Instantaneous shoulder line angle of the frame Its formula is:

[0047]

[0048] Calculate the first Instantaneous vertical offset of the frame header Its formula is:

[0049]

[0050] Complete the entire time window All After processing the frames, two frames of length were obtained. The data sequence, i.e., the shoulder line angle sequence: Head vertical offset sequence: .

[0051] Then calculate the variance of the shoulder line angle. Its formula is:

[0052]

[0053] in, The shoulder line angle in the window Mean and variance within Quantified in window The value indicates the degree of fluctuation in the angle of the user's shoulders. The larger the value, the more unstable the shoulder posture and the more frequent the swaying or drooping changes.

[0054] Next, calculate the mean of the vertical head offset. Its formula is:

[0055]

[0056] Among them, the mean Quantified in window The average vertical position of the user's head relative to the midline of the shoulders reflects the degree of vertical offset of the user's head relative to the torso.

[0057] Finally, the variance of the calculated shoulder line angle was analyzed. Mean of vertical offset from head The features are fused to obtain a comprehensive feature value, which is used as the posture retreat feature.

[0058] The formula for calculating the comprehensive eigenvalue is as follows:

[0059]

[0060] in, and These are preset weighting coefficients used to balance the influence of the two indicators; The higher the value, the more obvious the user's unconscious physical avoidance or lethargic posture.

[0061] For clarity, in the embodiments of this application, such as Figure 3 As shown, the multimodal feature extraction module also includes a pressure feature extraction unit.

[0062] The pressure feature extraction unit is used to calculate the average pressure difference between the left and right sides of the user's seat based on the pressure distribution data within the time window, which serves as the first pressure feature characterizing the shift of the body's center of gravity and the asymmetry of the sitting posture.

[0063] The standard deviation of the overall pressure value fluctuating over time within the calculated time window is used as a second pressure feature characterizing the user's physical tension.

[0064] The first pressure feature and the second pressure feature together constitute the pressure offset feature.

[0065] For example, within the same time window as the visual feature extraction unit Inside, calculate the average instantaneous pressure on both the left and right sides respectively. and in This represents the sampling time.

[0066] Let the left region contain One sensing unit.

[0067] Its set of readings is .

[0068] Then calculate the average pressure on the left side of the seat at that moment. The formula is

[0069]

[0070] in Indicates at the sampling time The first in the left area The raw pressure readings collected by each sensing unit.

[0071] Suppose the right region contains One sensing unit.

[0072] Its set of readings is .

[0073] Calculate the instantaneous average pressure in the right region at that moment. The formula is:

[0074]

[0075] in, Indicates at the sampling time The right-hand area The raw pressure readings collected by each sensing unit.

[0076] Then, sum all the pressure readings in the left and right regions to obtain the total pressure applied to the seat surface at that moment. Its formula is:

[0077]

[0078] Next, within the time window The internal sequence is constructed and its features are statistically analyzed. After processing the window... All After sampling points, we obtained three lengths of The data sequence.

[0079] Left-side average pressure sequence: ;

[0080] Right-hand mean pressure sequence: ;

[0081] Overall pressure value sequence: .

[0082] Calculate the time window based on the average pressure sequences on both sides. The average pressure on the left and right sides of the body is calculated, and the absolute value of the difference between the average pressure on the left and right sides is used as the first pressure characteristic to characterize the lateral shift of the body's center of gravity and the asymmetry of the sitting posture.

[0083] Among them, the average pressure on the left side The formula is:

[0084]

[0085] Right-side pressure mean The formula is:

[0086]

[0087] Then the absolute value of the difference between the two (i.e., the first pressure characteristic) Its formula is:

[0088]

[0089] The larger the feature value, the more significant the asymmetry of the user's posture, indicating that the user's body is continuously tilted to one side during the window period.

[0090] After obtaining the first pressure feature, the second pressure feature is calculated based on the overall pressure value sequence.

[0091] First, calculate the standard deviation of the overall pressure value sequence, which serves as the second pressure characteristic value. This is an indicator used to characterize the overall tension of the body.

[0092]

[0093] in, It is the overall pressure value within the time window The mean within, this characteristic value The larger the value, the more drastic the fluctuation in the user's total stress, reflecting higher muscle tension and postural instability.

[0094] Finally, the first and second pressure features mentioned above are combined to form a pressure offset feature vector, which is used in the subsequent risk assessment module. .

[0095] For clarity, in the embodiments of this application, such as Figure 4 As shown, the multimodal feature extraction module also includes a log feature extraction unit, which is used to extract at least two types of negative interaction behavior features based on the learning interaction logs within the time window. The negative interaction behavior features are used to quantitatively characterize the avoidance tendency and hesitation tendency shown by users during the answering process.

[0096] It should be noted that the learning interaction log records all interaction events between the user and the learning terminal. Each event includes an event type (such as starting to answer, modifying the answer, abandoning submission, etc.) and a precise timestamp. The processing range of the log feature extraction unit is a time window. All events recorded internally.

[0097] Specifically, the log feature extraction unit first reads the window sequentially. All log entries within the database are analyzed, and two types of event sequences related to negative psychological states are identified based on preset rules.

[0098] For example, Category events (avoidance-oriented events) are log entries identified as either abandoned submissions or timed-out unanswered events. Each time such an entry is detected, it is recorded as an occurrence. Class events, set time windows The total number of such events identified internally is .

[0099] The category of events (hesitation tendency events) is to identify a combination of log entries that meet the following composite conditions: (1) the event involves a question that has been marked by the system as being mastered by the user or whose difficulty is lower than the preset difficulty threshold; (2) the final submitted answer is recorded in the answer sequence of the question; (3) the modification number attribute value of the answer is greater than or equal to the preset number threshold (for example, the preset number threshold is set to 3).

[0100] Each time a question is detected that satisfies all of the above conditions, it is counted as an occurrence. Class events, set time windows The total number of such events identified internally is At the same time, for each For similar events, calculate the hesitation time from the first presentation of the problem to the final submission (e.g., calculate the time difference between the final submission and the first appearance of the problem as the hesitation time). ,in, Representing the indivual Class of events.

[0101] Then in the completion time window After identifying and classifying all events, calculate Frequency of avoidance of similar events (As the first negative interaction behavior characteristic), its formula is:

[0102]

[0103] This feature quantifies the frequency with which users abandon answering questions per unit of time. The higher the value, the stronger the user's avoidance tendency during the window period.

[0104] By calculating the average hesitation intensity As a second negative interaction behavior feature, it is used to characterize the degree of hesitation of users within the time window.

[0105] First calculate all Total hesitation time for this type of event: Then, the average hesitation intensity is calculated based on the total hesitation duration. Its formula is:

[0106]

[0107] This feature quantifies the average time users spend excessively modifying simple problems they already understand. The higher the value, the stronger the user's tendency to hesitate and agonize during the window period.

[0108] Finally, the first and second negative interaction behavior features are combined to form a negative interaction behavior feature vector. For subsequent risk assessment, i.e. .

[0109] The risk assessment module is used to integrate negative interaction behavior features, posture withdrawal features, and stress shift features to generate a comprehensive behavioral pattern vector. This comprehensive behavioral pattern vector is then input into a pre-set psychological risk assessment model to obtain the user's learned helplessness risk level.

[0110] For clarity, in the embodiments of this application, the risk assessment module includes a feature fusion unit and a model calculation unit.

[0111] The feature fusion unit normalizes the posture retreat features, pressure offset features, and negative interaction behavior features to eliminate differences in the dimensions and numerical ranges of different features, bringing them to the same scale. Normalization can employ the min-max scaling method, mapping each feature value to the [0, 1] interval. These normalized features are then concatenated in a predetermined order to form a comprehensive behavior pattern vector. .

[0112] The model computation unit is used to input the comprehensive behavioral pattern vector into a pre-set psychological risk assessment model. For example, the psychological risk assessment model can be a machine learning model trained on a large amount of historical data based on the gradient boosting tree (i.e., GBDT model) algorithm.

[0113] Understandably, GBDT is a well-known ensemble learning algorithm in this technical field for classification tasks.

[0114] In application, the psychological risk assessment model receives a comprehensive behavioral pattern vector. It outputs a predicted probability value. This probability represents the likelihood that the psychological risk assessment model determines that the user is currently in a state of learned helplessness. The learned helplessness risk level is divided according to a preset probability interval, for example:

[0115]

[0116] The guidance decision module is used to select a target guidance strategy from a preset strategy library when the learned helplessness risk level exceeds a preset level threshold, based on the current risk level and posture withdrawal characteristics and / or pressure offset characteristics.

[0117] The specific steps for obtaining the target guidance strategy are: determining whether the posture retreat feature exceeds the first somatosensory threshold; and determining whether the pressure offset feature exceeds the second somatosensory threshold.

[0118] It should be noted that the evacuation decision module reads the original posture retreat characteristics and pressure offset characteristics (rather than the normalized values) that led to the current risk assessment, compares them with a preset threshold, and generates a Boolean judgment result:

[0119] Determine the characteristics of posture withdrawal Does it exceed the first sensory threshold? ,Right now .

[0120] The first pressure feature for determining pressure offset characteristics Second pressure characteristics Whether they exceed the corresponding second somatosensory threshold , ,Right now .

[0121] The risk level (e.g., high risk is mapped to the numeric code 2, medium risk to 1, and low risk to 0) is combined with the two Boolean conditions mentioned above to form a unique combination.

[0122] For example: Combination term (2, , The value 1 indicates a high risk and significant posture retreat, but no significant pressure shift.

[0123] Combination term (2, , This indicates a high risk and significant pressure shift, but no significant posture retreat.

[0124] Combination term (2, , This indicates a high risk, with significant posture retreat and pressure shift.

[0125] Use the above combination items to query the preset strategy library. The strategy library pre-stores the target guidance strategies corresponding to different combination items. Each strategy defines the core intervention orientation of the guidance content (e.g., mainly cognitive restructuring to resist behavioral withdrawal, mainly relaxation exercises to relieve physiological tension, or comprehensive intervention), as well as the template identifier and intensity parameters required for content generation.

[0126] The content generation module is used to generate guidance content that reconstructs users' cognition based on the target guidance strategy, and to control the learning terminal to embed the guidance content into the learning content for presentation.

[0127] It should be noted that the content generation module receives the target guidance strategy from the guidance decision module. Based on the intervention orientation identifier contained in the target guidance strategy, it calls the corresponding content template and material library to generate guidance content with a specific focus.

[0128] For example, if the target guidance strategy is mainly based on cognitive restructuring and resistance to withdrawal, then the generated guidance content will focus on guiding users to review their success records, showcasing their long-term progress curves, and adjusting their attribution cognition.

[0129] If the target guidance strategy focuses on relieving physiological tension through relaxation exercises, the generated guidance content will emphasize guiding users to perform physical relaxation training such as deep breathing and posture adjustment.

[0130] If the target guidance strategy is comprehensive intervention, the generated guidance content integrates the two types of guidance information mentioned above.

[0131] Then, the content generation module controls the learning terminal unit to dynamically embed the generated guidance content into the currently presented regular English learning content in the form of non-intrusive floating information boxes, edge animations, or audio prompts, and present it in a coordinated manner.

[0132] In another embodiment of this application, the generation of the diversion content includes:

[0133] For example, if the goal guidance strategy is mainly based on cognitive restructuring, confrontational behavior, and withdrawal, then based on the goal guidance strategy, the system retrieves historical answer records that meet the preset success conditions from the user's historical learning records and generates review guidance text to guide the user to review the historical answer records.

[0134] It should be noted that the preset success conditions can be: (1) Answer result: correct; (2) Answer time is less than the average time of the question; (3) The question knowledge point tag is the user's recent weak knowledge points integration.

[0135] The content generation module then accesses the user's personal learning database and retrieves and filters all historical answer records that meet the above success criteria from the learning records of the most recent historical period (such as the past week or month).

[0136] Select one or several of the most representative records from the search results (e.g., the earliest correct record or the record most relevant to the current learning topic). Based on the specific information of these records (such as question content, time taken, and knowledge points), generate a natural language-formatted review guide text. For example: "Look, just two days ago, you completed this exercise on the 'present perfect tense' without any errors, and here's a thumbnail of the question. You took 20% faster than average, proving that you are fully capable of mastering this knowledge point." This guides users to focus on the positive evidence they have overlooked, thus combating the cognitive bias typical of learned helplessness—only remembering failures and ignoring successes.

[0137] For example, based on learning interaction logs, the content generation module generates visual information that reflects the user's long-term ability improvement trend.

[0138] The content generation module extracts evaluation data sequences from users' learning interaction logs over multiple consecutive time periods (such as the past 12 weeks) on relevant skill dimensions (such as vocabulary, reading comprehension accuracy, and grammar score).

[0139] The content generation module uses a chart generation engine to process the above data sequence into intuitive visualizations, such as smooth trend line charts. These charts can use time as the horizontal axis and ability value as the vertical axis, marking the starting point, peak point, and current point. Concise explanatory text is added next to the charts, highlighting observable positive trends, such as: "This is the trend of your reading comprehension accuracy over the past three months. Although there were fluctuations this week, the overall curve is steadily rising. Your long-term efforts are bringing real change," to combat cognitive distortions caused by short-term setbacks.

[0140] For example, based on the knowledge points and difficulty of the user's current learning task, the content generation module generates task breakdown guidance text that breaks down the learning task into multiple sub-steps.

[0141] The content generation module analyzes the attributes of the user's current learning task, including the knowledge points it belongs to, the difficulty level set by the system, and the structure of the task itself.

[0142] Based on the task attributes, pre-defined decomposition rules or algorithms are invoked to break down the current task into 3-5 ordered sub-steps. For example, for an English cloze test, the decomposition result is:

[0143] Step 1: Quickly read through the entire text once, ignoring the blank spaces and only grasping the main idea.

[0144] Step 2: Read it a second time, and only fill in the 3 blanks you are most confident about.

[0145] Step 3: On the third pass, focus on filling in the remaining blanks, paying particular attention to the sentence's grammatical structure.

[0146] The decomposed sequence of steps is transformed into clear and encouraging task breakdown guidance text, which breaks down macro tasks into a series of controllable micro steps, thereby lowering the starting threshold and enhancing users' sense of self-efficacy.

[0147] Among them, at least one of the following is integrated into the guidance content: review guidance text, visual information, and task breakdown guidance text.

[0148] It should be noted that, based on the intervention focus defined in the goal guidance strategy, at least one of the following should be integrated: review guidance text, visual information, and task breakdown guidance text.

[0149] For example, for the combination term (2, , The value of '(')' indicates high risk and significant postural withdrawal, but no significant stress shift. Therefore, the strategy focuses on cognitive restructuring to counteract behavioral withdrawal. This involves integrating guided review texts and visual information to provide evidence of success and objective progress to combat cognitive withdrawal.

[0150] For the combination term (2, , The value of '(')' indicates high risk and significant stress shift, but no significant postural withdrawal. Therefore, the strategy focuses on alleviating physiological tension, which may involve integrating task breakdown guidance text and indirectly alleviating anxiety by reducing the perceived threat of the task, or incorporating simple relaxation techniques.

[0151] In summary, the multimodal emotion recognition and intelligent guidance device for English learning according to the embodiments of this application constructs a multi-dimensional comprehensive behavioral pattern vector by integrating negative interaction behavior features, posture withdrawal features, and stress shift features. This overcomes the one-sidedness of judging solely based on correct or incorrect answers or single facial expressions, and can more accurately and reliably identify deep psychological risks related to long-term behavioral patterns and physiological states, such as "learned helplessness." Unlike simple emotional comfort, this device can automatically generate guidance content for cognitive reconstruction after identifying psychological risks, guiding users to review minor... By demonstrating success, showcasing objective progress curves, and providing task breakdown guidance, interventions are directly targeted at the core cognitive biases of "learned helplessness." Visual analysis of the user's head and shoulder posture withdrawal characteristics captures the user's unconscious body language under frustration. Pressure sensors detect pressure shift characteristics in the hip pressure distribution, quantifying the instability in sitting posture and frequent shifts in the center of gravity caused by anxiety and tension. These somatosensory characteristics, as physiological behavioral indicators, are highly correlated with psychological state and are not easily concealed by the user's subjective feelings, thus greatly enriching the dimensions of emotion recognition and improving the robustness and accuracy of recognition.

[0152] In another embodiment of this application, such as Figure 5 As shown, the hardware computing units of the multimodal feature extraction module, risk assessment module, guidance decision module, and content generation module are integrated into the host box 300. The host box 300 is a miniaturized aluminum alloy box with a hollow heat dissipation design and a built-in micro cooling fan to ensure heat dissipation during long-term computing and avoid processing lag caused by overheating.

[0153] The main unit box 300 has a centralized interface compartment on its side, where the connection interfaces of the pressure sensing module 100, the vision acquisition module 200, and the learning terminal are centrally arranged. The interface compartment is equipped with a dust cover, which should be closed when not in use to prevent dust from entering. The compartment is equipped with a cable fixing clip 131, which can lock and fix the connection cables to prevent them from falling off.

[0154] The bottom of the host box 300 is equipped with anti-slip silicone feet and an adhesive fixing sticker, which can fix the host box 300 in any position on the desktop to prevent it from being moved by collision. The top of the host box 300 has a magnetic cable storage slot, which can store excess connecting cables, solve the problem of messy desktop cables, and reduce signal interference caused by tangled cables.

[0155] In another embodiment of this application, such as Figure 5 As shown, the motion data acquisition device includes a pressure sensing module 100 mounted on the user's seat and a visual acquisition module 200 positioned facing the user.

[0156] For clarity, in the embodiments of this application, such as Figure 6 and Figure 7 As shown, the pressure sensing module 100 includes a flexible pad 110, a pressure sensor 120, an elastic strap 130, an installation chamber 140, and an anti-slip strip 150.

[0157] Multiple pressure sensors 120 are embedded in two groups on the left and right sides of the top surface of the flexible pad 110. Elastic straps 130 are set on the side of the flexible pad 110 to bypass the seat and be locked by buckles 131. The mounting compartment 140 is detachably set on the rear side of the flexible pad 110. The signal processing chip and data interface of the pressure sensor 120 are integrated in the mounting compartment 140. Anti-slip strips 150 are set on the bottom of the flexible pad 110.

[0158] As one possible scenario, the central area of ​​the flexible pad 110 is a flexible silicone pad, and the pressure sensor 120 is evenly embedded inside the silicone pad (distributed according to the pressure area of ​​the buttocks, with the left and right area sensing units independently partitioned, and the left and right pressure calculation logic matched with pressure feature extraction). The surface of the silicone pad is treated with an anti-slip texture to prevent sensor displacement caused by the user's sitting posture movement.

[0159] It should be noted that the elastic strap 130 is stretchable and can be adapted to the width of the seat (such as study chair, desk chair, office chair) and length. The inner side of the strap is equipped with an anti-slip rubber pad to prevent the elastic strap 130 from slipping at the contact point with the seat. At the same time, the elastic strap 130 can be quickly removed and is adapted to different types of seats, such as those without armrests and those with armrests.

[0160] Understandably, the housing of the installation compartment 140 is a magnetically detachable design. When the flexible pad 110 is not in use, the data transmission compartment can be removed and stored separately to prevent the wiring from aging.

[0161] In addition, the anti-slip strip 150 is an adhesive silicone strip that can be used to help fix it to the smooth seat surface, forming a double fixation with the strap, thus solving the problem of data distortion caused by the displacement of the pressure sensing module 100.

[0162] For clarity, in the embodiments of this application, such as Figure 8 As shown, the visual acquisition module 200 is a dedicated adjustable acquisition structure adapted to the multimodal emotion recognition device for English learning. The whole is a detachable and multi-dimensionally adjustable integrated design, focusing on the core functions of acquiring the user's upper body video stream and accurately locating the two-dimensional coordinates of the head and shoulders in this application.

[0163] The visual acquisition module 200 includes a mounting base 210, a telescopic rod 220, a light-shielding and image-stabilizing compartment 230, and a camera 240.

[0164] The mounting base 210 is detachably mounted on the learning terminal, the telescopic rod 220 is rotatably mounted on the mounting base 210 and locked by a damping knob, the light-shielding and image-stabilizing compartment 230 is rotatably mounted on the telescopic rod 220 and locked by another damping knob, and the camera 240 is mounted inside the light-shielding and image-stabilizing compartment 230.

[0165] Understandably, the components are connected by a rotatable hinge and an independent damping knob for locking. The mounting base 210 is the core fixing carrier for the module and the learning terminal. The telescopic rod 220 realizes the adjustment of the acquisition distance and coarse angle. The light-shielding and anti-shake compartment 230 realizes the fine angle adjustment of acquisition and the protection of the camera 240. The camera 240 is the core component for video stream acquisition. The cooperation of each component realizes the precise control of the acquisition position and angle.

[0166] It should be noted that the light-shielding and image-stabilizing chamber 230 reduces the impact of ambient light reflection on video stream acquisition. A micro-shock-absorbing pad is set between the light-shielding and image-stabilizing chamber 230 and the camera 240 to prevent image blurring caused by slight shaking and ensure the accuracy of head and shoulder coordinate acquisition.

[0167] In addition, the telescopic pole 220 adopts a multi-segment axial telescopic pole structure, and the lifting height can be adjusted between 30cm and 80cm. The telescopic pole 220 can rotate at multiple angles in the horizontal or pitch direction along the hinge, realizing coarse angle adjustment of the acquisition direction of the camera 240. After adjusting to the target angle, tightening the damping knob at the hinge can realize the relative locking and fixation of the telescopic pole 220 and the mounting base 210, preventing angle deviation caused by touching or shaking during use. It is suitable for the upper body acquisition height of users of different heights such as primary school students, middle school students, and adults, and matches the needs of upper body video stream acquisition.

[0168] The telescopic pole 220's telescopic function can adjust the acquisition distance between the camera 240 and the user's upper body according to the distance between the user and the learning terminal, ensuring that the acquisition field of view can completely cover the user's head and shoulder area, and meet the acquisition field of view requirements for locating the center point of the head, the left shoulder point, and the right shoulder point.

[0169] Understandably, the light-shielding and image-stabilizing compartment 230 can be tilted or rotated at a small angle along the hinge to achieve fine adjustment of the camera 240's acquisition angle, accurately aligning with the user's upper body center line and ensuring the accuracy of head and shoulder coordinate acquisition; after adjustment, tightening the damping knob at this position can lock and fix the light-shielding and image-stabilizing compartment 230 to the telescopic rod 220, preventing the compartment from shaking.

[0170] The light-shielding and image-stabilizing compartment 230 features a light-shielding, enclosed design with a reserved mounting cavity for the camera 240, providing independent installation and protection space for the camera 240. The acquisition window is only opened on the side facing the user to ensure the field of view of the camera 240 lens.

[0171] As one possible configuration, the mounting base 210 includes a hook 211, a support pad 212, a screw 213, and a throttle 214.

[0172] Among them, the hook 211 is an asymmetrical buckle plate of unequal length, and the whole is U-shaped with the opening facing downward. Its short end is the front fitting end and the long end is the back abutting end. When the hook 211 is fastened on the top of the learning terminal, the short end fits tightly against the front screenless area at the top of the display screen, and the long end naturally extends to the back of the top of the display screen. The U-shaped structure achieves the initial buckle 131 positioning with the learning terminal, and the design of the short end fitting against the front screenless area can completely avoid obstructing the screen of the learning terminal.

[0173] The screw 213 and the long end of the hook 211 are connected by a threaded connection. The axis of the screw 213 is perpendicular to the back of the display screen. The end of the screw 213 facing the back of the display screen is rotatably connected to the support pad 212. The other end of the screw 213 is fixedly connected to the handle 214, which is a manual operation component used to drive the screw 213 to rotate.

[0174] When the handle 214 is turned manually, the screw 213 pushes the support pad 212 to the back of the top of the display screen to tighten until the support pad 212 is in close contact with the back of the display screen. Through the positioning of the hook 211 and the tightening of the support pad 212 by the screw 213, the mounting base 210 and the learning terminal are firmly fixed.

[0175] In addition to being fixed to the learning terminal, the hook 211 can also be fixed to other suitable locations such as the desktop, depending on the specific scenario.

[0176] It should be noted that the visual acquisition function integrated into the learning terminal (such as the camera unit built into the display screen) has a fixed and non-adjustable acquisition angle and position. Due to the limitations of the terminal's placement and size, it is prone to problems such as field of view deviation and missed acquisition of key areas. For example, the camera unit built into the display screen is mostly located at the top center of the screen. If the user's sitting posture is slightly off, or if the user is too tall or too short, or if the terminal is placed at an angle, the camera unit may not be able to fully cover the head and shoulder areas, or may only be able to acquire the face or part of the upper body. This cannot meet the core acquisition requirements of this application for locating the center point of the head, the left shoulder point, and the right shoulder point.

[0177] The visual acquisition module 200 of this application is an independently adjustable structure. It is fixed by the mounting base 210, the telescopic rod 220 extends and retracts, and the light-shielding and anti-shake chamber 230 rotates in multiple dimensions. It can achieve coarse and fine adjustment of the acquisition angle and acquisition distance. It can also lock and fix the adjusted position by the damping knob. Regardless of the user's height, sitting posture, or the placement of the learning terminal, the camera 240 can be accurately aimed at the midline of the user's upper body to ensure that the acquisition field of view fully covers the key areas of the head and shoulders. This provides complete and accurate raw visual data for subsequent extraction of posture retreat features, avoiding the field of view defects of integrated acquisition from the source.

[0178] In addition, the visual acquisition function integrated into the learning terminal is deeply bound to the terminal hardware and can only be used on that terminal. Due to the limitations of the terminal hardware design, it cannot be adapted to other learning terminals. For example, the built-in camera unit of a tablet can only be used on that tablet. If it is replaced with a desktop monitor or a laptop, there is no corresponding acquisition function. If the original terminal is damaged or replaced, the entire visual acquisition process of emotion recognition will fail.

[0179] The visual acquisition module 200 of this application achieves detachable and quick-installation fixation through the mounting base 210 tightened by the unequal length hooks 211 and the screw 213, without the need for hardware binding with the learning terminal. It can be adapted to various learning terminals with screens such as tablets, desktop displays, and laptops. Moreover, the screw 213 can be freely extended and retracted to adapt to the top structure of terminals with different thicknesses and sizes. At the same time, installation and disassembly do not require tools, and can be quickly switched between different terminals, completely getting rid of the terminal binding limitations of integrated acquisition and adapting to the actual scenarios of users changing or using different learning terminals.

[0180] In the description of this specification, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0181] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0182] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A multimodal emotion recognition and intelligent guidance device for English learning, characterized in that, include: Learning terminal, motion-sensing data acquisition device, and central processing unit; The learning terminal is used to present learning content and record the user's learning interaction log; The somatosensory data acquisition device is used to collect video stream data of the user's upper body and pressure distribution data between the user's buttocks and the seat; The central processing unit includes a multimodal feature extraction module, a risk assessment module, a guidance decision module, and a content generation module; The multimodal feature extraction module is used to receive the learning interaction log and extract negative interaction behavior features representing answering behavior based on the same time window, receive the video stream data and extract posture withdrawal features based on the user's head and shoulders, and receive the pressure distribution data and extract pressure offset features representing sitting posture stability. The risk assessment module is used to integrate the negative interaction behavior features, the posture withdrawal features, and the stress shift features to generate a comprehensive behavior pattern vector, and input the comprehensive behavior pattern vector into a preset psychological risk assessment model to obtain the user's learned helplessness risk level. The guidance decision module is used to select a target guidance strategy from a preset strategy library when the learned helplessness risk level exceeds a preset level threshold, based on the current risk level and the posture retreat feature and / or the pressure offset feature. The content generation module is used to generate guidance content for cognitive reconstruction of users based on the target guidance strategy, and to control the learning terminal to embed the guidance content into the learning content for presentation.

2. The multimodal emotion recognition and intelligent guidance device for English learning according to claim 1, characterized in that, The multimodal feature extraction module includes a visual feature extraction unit; The visual feature extraction unit is used to perform human head and shoulder region localization on video stream data and obtain a two-dimensional image coordinate sequence of the head center point, left shoulder point and right shoulder point; It also calculates the angle of the shoulder line formed by the left and right shoulder points, and the vertical offset of the head center point relative to the midpoint of the shoulder line. Within the time window, the variance of the shoulder line angle and the mean of the vertical offset are calculated, and the weighted sum of the two is used as the posture retreat feature.

3. The multimodal emotion recognition and intelligent guidance device for English learning according to claim 1, characterized in that, The multimodal feature extraction module also includes a pressure feature extraction unit, which is used to calculate the average pressure difference between the left and right sides of the user's seat based on the pressure distribution data within the time window, as a first pressure feature characterizing the shift of the body's center of gravity and the asymmetry of the sitting posture. Calculate the standard deviation of the overall pressure value fluctuating over time within the time window, and use it as a second pressure feature characterizing the user's physical tension. The first pressure feature and the second pressure feature together constitute the pressure offset feature.

4. The multimodal emotion recognition and intelligent guidance device for English learning according to claim 1, characterized in that, The multimodal feature extraction module further includes a log feature extraction unit, which is used to extract at least two types of negative interaction behavior features based on the learning interaction logs within the time window. The negative interaction behavior features are used to quantitatively characterize the avoidance tendency and hesitation tendency shown by users during the question-answering process.

5. The multimodal emotion recognition and intelligent guidance device for English learning according to claim 1, characterized in that, The risk assessment module includes a feature fusion unit and a model calculation unit, wherein... The feature fusion unit is used to normalize the posture retreat feature, the pressure offset feature, and the negative interaction behavior feature respectively, and then concatenate them into the comprehensive behavior pattern vector. The model calculation unit is used to input the comprehensive behavioral pattern vector into the preset psychological risk assessment model, which is used to output the predicted probability of the user falling into a learned helplessness psychological state; the learned helplessness risk level is divided according to the interval in which the predicted probability falls.

6. The multimodal emotion recognition and intelligent guidance device for English learning according to claim 1, characterized in that, The step of selecting a target evacuation strategy from a pre-set strategy library based on the current risk level and the posture retreat characteristics and / or the pressure offset characteristics specifically includes: Determine whether the posture retreat feature exceeds the first somatosensory threshold; Determine whether the pressure offset feature exceeds the second somatosensory threshold; A combination term is formed based on the comparison results of the learned helplessness risk level, the posture withdrawal feature, and the pressure shift feature; Based on the combination item, a unique target diversion strategy is matched from the preset strategy library.

7. The multimodal emotion recognition and intelligent guidance device for English learning according to claim 6, characterized in that, The generation of the diversion content includes: Based on the target guidance strategy, retrieve historical answer records that meet the preset success conditions from the user's historical learning records, and generate review guidance text to guide the user to review the historical answer records; Based on the learning interaction log, visual information reflecting the user's long-term ability improvement trend is generated; Based on the knowledge points and difficulty of the user's current learning task, generate task decomposition guidance text that breaks down the learning task into multiple sub-steps; Among them, at least one of the review guidance text, the visualization information, and the task decomposition guidance text is integrated into the guidance content.

8. The multimodal emotion recognition and intelligent guidance device for English learning according to claim 1, characterized in that, The somatosensory data acquisition device includes a pressure sensing module (100) mounted on the user's seat and a visual acquisition module (200) positioned facing the user.

9. The multimodal emotion recognition and intelligent guidance device for English learning according to claim 8, characterized in that, The pressure sensing module (100) includes a flexible pad (110), a pressure sensing module (120), an elastic strap (130), an installation chamber (140), and an anti-slip strip (150), wherein, Multiple pressure sensors (120) are embedded in two groups on the left and right sides of the top surface of the flexible pad (110); The elastic strap (130) is provided on the side of the flexible pad (110) to go around the seat and be locked by a buckle (131); The mounting chamber (140) is detachably disposed on the rear side of the flexible pad (110), wherein the signal processing chip and data interface of the pressure sensor (120) are integrated in the mounting chamber (140); The anti-slip strip (150) is disposed at the bottom of the flexible pad (110).

10. The multimodal emotion recognition and intelligent guidance device for English learning according to claim 8, characterized in that, The visual acquisition module (200) includes a mounting base (210), a telescopic rod (220), a light-shielding and image-stabilizing compartment (230), and a camera (240), wherein, The mounting base (210) is detachably mounted on the learning terminal; The telescopic rod (220) is rotatably mounted on the mounting base (210) and locked by a damping knob; The light-shielding and anti-shake compartment (230) is rotatably mounted on the telescopic rod (220) and locked by another damping knob; The camera (240) is located inside the light-shielding and image-stabilizing compartment (230).