Psychological state self-adaptive grading method and system based on multiple modes

By employing multimodal perception and adaptive grading methods, this study addresses the accuracy degradation issue in existing multimodal emotion recognition systems when environmental changes and modalities are missing. It enables comprehensive assessment and timely grading and early warning of children's psychological states, improving the system's adaptability and robustness, and making it suitable for resource-constrained environments.

CN121817888APending Publication Date: 2026-04-10NANJING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING UNIV OF POSTS & TELECOMM
Filing Date
2025-12-12
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing multimodal emotion recognition systems suffer from decreased recognition accuracy when faced with environmental changes and modality loss. They also lack tiered early warning and auditable mechanisms, making it impossible to effectively assess the psychological state of ordinary children and provide timely intervention.

Method used

A multimodal perception method is adopted, which collects and recognizes visual, speech, behavioral and text data, and combines sliding window adaptive weight allocation, exponential normalization and two-stage time smoothing mechanism to dynamically adjust weights. The baseline mean and variance are introduced to determine the time-varying global threshold for hierarchical early warning. An autoregressive correction factor is introduced at the decision level to achieve adaptive hierarchical classification of psychological state.

Benefits of technology

It enables a comprehensive and objective assessment of children's psychological state, provides timely and tiered early warnings and notifies guardians, improves the accuracy and stability of early warnings, is suitable for resource-constrained environments, reduces costs, and supports both personalized and universal use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121817888A_ABST
    Figure CN121817888A_ABST
Patent Text Reader

Abstract

The invention discloses a psychological state self-adaptive grading method and system based on multiple modes, and belongs to the technical field of artificial intelligence. Collecting four modes of vision, voice, behavior and text of an identified person, and calculating through each mode identification sub-model to obtain a single-mode confidence coefficient; a self-adaptive weight distribution mechanism based on a sliding window is introduced to determine a dynamic stability score of each mode; obtaining a time-varying weight by adopting index normalization to obtain a multi-modal fusion score; a two-stage time smoothing mechanism is adopted to obtain a time sequence score finally used for early warning judgment; according to the time-varying global threshold, the sensitivity bandwidth, the proportion of the upper threshold and the continuous threshold exceeding duration, determining a grading early warning trigger criterion, and then performing grading early warning on the time sequence score to obtain an early warning grade; an autoregressive correction factor is introduced into a decision-making layer, so that the early warning level depends on the state of the previous moment, and the early warning level depending on the previous moment is obtained. The method not only has multi-mode perception, but also can perform graded early warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a multimodal adaptive hierarchical method and system for psychological states, belonging to the field of artificial intelligence technology. Background Technology

[0002] Currently, mental health issues such as depression and anxiety are prevalent among children and adolescents, while professional psychological intervention resources are relatively insufficient, with approximately 90% of children failing to receive effective treatment. With the deep integration of artificial intelligence and healthcare, digital intervention technologies are becoming an important means to improve the efficiency of mental health management. Intervention studies targeting children with autism have shown that by constructing multimodal interactive environments and integrating technologies such as computer vision and speech recognition into virtual training scenarios, it is possible to estimate children's attention and emotional states and dynamically adjust training activities based on observed psychological states.

[0003] Existing patent CN102354349A discloses a multimodal early intervention system for improving the social skills of children with autism. The system consists of a multi-touch screen, three cameras, and a computer. The cameras, equipped with microphones, collect video and audio signals from the child, while the touchscreen captures the child's movements. The system comprises six modules: visual signal processing, speech signal processing, a physical interactive interface, multimodal fusion, an intelligent control console, and real-scene simulation. Specifically, the visual processing module analyzes facial expressions and head posture, the speech processing module analyzes tone and voice features, the touch module recognizes touchscreen actions, and the multimodal fusion module integrates information such as head position, eye contact, facial expressions, voice, and gestures to generate the child's current learning and psychological state. The intelligent control console generates game interactions between the child and virtual animated characters based on designed learning tasks, while the real-scene simulation module displays different virtual scenes and sound effects on the screen according to console commands. This technical solution utilizes a real-time human-computer game interaction environment, assesses the child's psychological state through multimodal signal fusion, and dynamically adjusts training activities accordingly, forming a virtuous cycle.

[0004] The aforementioned solutions primarily focus on social skills training for children with autism, and lack sufficient support for the detection and timely warning functions of psychological states in typical children. Furthermore, existing multimodal emotion recognition and psychological intervention terminals typically include a camera, microphone, and several interactive sensors (such as touch sensors and buttons). The workflow of such systems is generally as follows: 1. Data Acquisition: Acquire facial images, voice signals, and interactive operation data; 2. Feature extraction: Facial expression features are extracted using image processing algorithms, and speech features such as pitch, volume, and formants are extracted using acoustic analysis methods; 3. Fusion and Determination: Input the features of each modality into a preset model or algorithm for fusion, and output the emotion category and confidence level; 4. Intervention and Feedback: When negative emotions are detected, immediate feedback is provided through methods such as voice playback, lighting changes, and screen prompts.

[0005] The advantage of this type of technology lies in its ability to comprehensively utilize visual and auditory information, improving the coverage and accuracy of emotion recognition, and achieving a certain degree of emotion intervention through hardware interaction. However, this multimodal emotion recognition fusion strategy is simplistic: it often employs fixed weighting or simple voting methods for multimodal fusion, lacking the ability to dynamically adjust weights based on environmental changes (such as noise, lighting, and occlusion). It is not robust to modality loss: when data for a particular modality is missing or of poor quality (such as facial occlusion or audio signal distortion), the overall accuracy of the judgment decreases significantly. Furthermore, it lacks tiered early warning and auditable mechanisms: existing systems mostly provide instant feedback, lacking event tiering, long-term emotion profile recording, and traceable evidence-based data management mechanisms. Summary of the Invention

[0006] Purpose of the invention: This invention provides a multimodal adaptive classification method for psychological states that has multimodal perception and can provide hierarchical early warning.

[0007] Technical solution: To achieve the above objectives, the technical solution adopted by this invention is as follows: A multimodal adaptive hierarchical method for mental states includes the following steps: Step 1: Collect four modalities of the subject: visual, speech, behavior, and text.

[0008] Step 2: The collected visual, speech, behavioral, and text modalities are used to calculate the single-modal confidence scores using the respective modal recognition sub-models.

[0009] Step 3: The obtained single-mode confidence scores are used to determine the dynamic stability score of each mode by introducing an adaptive weight allocation mechanism based on a sliding window.

[0010] Step 4: Apply exponential normalization to the dynamic stability score of each modality to obtain time-varying weights and thus obtain the multimodal fusion score.

[0011] Step 5: Apply a two-stage time smoothing mechanism to the obtained multimodal fusion score to obtain the final time-series score used for early warning determination.

[0012] Step 6, Introduce the baseline mean With baseline variance Determine the time-varying global threshold. Analyze the sensitivity bandwidth based on a combination of modal variance and global fluctuation. Use a frequency-duration-intensity hybrid decision rule to determine the proportion exceeding the upper threshold and the duration of continuous over-threshold occurrences.

[0013] Step 7: Determine the graded early warning triggering criteria based on the time-varying global threshold, sensitivity bandwidth, the proportion of the upper threshold, and the duration of continuous over-threshold. Obtain the early warning level by graded early warning based on the time series score through the graded early warning triggering criteria.

[0014] Step 8: Introduce an autoregressive correction factor at the decision-making level to make the warning level dependent on the state at the previous moment, thus obtaining the warning level dependent on the state at the previous moment.

[0015] The preferred formula for calculating the confidence level of a single mode is:

[0016] in, Indicates the confidence level of a single mode. This represents the Sigmoid activation function. This represents the recognition sub-model of the m-th modality. This represents the recognition sub-model of the m-th modality. This indicates a modal identifier used to distinguish different data modalities. Representing visual modality, Represents speech modality, Represents behavioral modality, Represents the text modality.

[0017] Preferred formula: The dynamic stability score for each mode is calculated as follows:

[0018]

[0019] in, This represents the dynamic stability score for each mode. Indicates the length of the sliding window. This represents the single-mode confidence of the ti-th mode at time t. This represents the sliding window variance of the confidence level of the m-th modality at time t. This represents the average value of the sliding window for the m-th mode at time t. This represents the dynamic stability score of the m-th mode at time t. This is the adjustable weighting coefficient.

[0020] The preferred formula for calculating the multimodal fusion score is:

[0021] in, Indicates the multimodal fusion score. Represents the weight sensitivity coefficient. This represents the dynamic stability score of the k-th mode at time t. Indicates the modal traversal index. This represents the dynamic stability score of the k-th mode at time t.

[0022] The preferred formula for calculating the time series score is:

[0023]

[0024] in, This represents the score after short-term exponential smoothing in the first stage at time t. Indicates the short-term smoothing coefficient. Indicates the time series score. This represents the score after long-term adaptive trend smoothing in the second stage at time t. This represents the long-term smoothing coefficient.

[0025] The preferred formula for calculating the time-varying global threshold is:

[0026] in, Indicates the time-varying global threshold. This represents the normal baseline mean of the psychological state of the person being identified. This represents the baseline variance adjustment factor. The normal baseline variance representing the mental state of the identified individual. This represents the rate of change adjustment coefficient. This represents the rate of change of the time series score at time t.

[0027] The formula for calculating sensitivity bandwidth is:

[0028] in, This represents the sensitivity bandwidth at time t. This represents the modal variance adjustment coefficient. This represents the time-varying weight of the m-th mode at time t. This represents the sliding window variance of the confidence level of the m-th modality at time t. This represents the single-mode confidence time series of the m-th mode. This represents the baseline variance adjustment factor. The normal baseline variance representing the mental state of the identified individual. This represents the minimum constant.

[0029] The formulas for calculating the proportion exceeding the upper threshold and the duration of continuous over-threshold are as follows:

[0030]

[0031] in, This represents the proportion above the threshold. This indicates the number of samples used for threshold determination. This indicates that the indicator function takes a value of 1 when the condition is met, and 0 otherwise, and is used to count the number of times the threshold is exceeded. Let represent the time series score at time ti. Indicates the time-varying global threshold. This represents the sensitivity bandwidth at time t. Indicates the duration of continuous exceedance. Indicates the sampling time interval.

[0032] The preferred formula for calculating the warning level is:

[0033] in, Indicates the time-varying global threshold. Indicates the sensitivity bandwidth. This represents the proportion above the threshold. Indicates the duration of continuous exceedance.

[0034] The preferred formula for calculating the warning level based on the previous moment is:

[0035] in, This represents the final warning level at time t, which depends on the state at the previous time. This indicates the k-th level warning. This represents the autoregressive correction factor for the k-th level warning. Let represent the probability that the condition for a level k warning is met at time t. Indicates an indicator function.

[0036] Another objective of this invention is to provide a multimodal adaptive grading system for psychological states, employing the aforementioned multimodal adaptive grading method for psychological states, comprising a data acquisition unit, a modality recognition sub-model unit, a dynamic stability unit, a multimodal fusion unit, a temporal scoring unit, a warning parameter determination unit, a graded warning triggering criterion unit, a temporal warning level unit, and an output unit, wherein: The acquisition unit is used to acquire four modalities of the subject: visual, speech, behavior, and text.

[0037] The modality recognition sub-model unit is used to collect four modalities: visual, speech, behavior, and text. The single-modality confidence is calculated by each modality recognition sub-model.

[0038] The dynamic stability unit is used to determine the dynamic stability score of each mode by introducing an adaptive weight allocation mechanism based on a sliding window, based on the obtained single-mode confidence.

[0039] The multimodal fusion unit is used to obtain time-varying weights by exponentially normalizing the dynamic stability scores of each obtained modality, thereby obtaining the multimodal fusion score.

[0040] The time-series scoring unit is used to apply a two-stage time smoothing mechanism to the obtained multimodal fusion score to obtain the final time-series score used for early warning determination.

[0041] The early warning parameter determination unit is used to introduce a baseline mean. With baseline variance Determine the time-varying global threshold. Analyze the sensitivity bandwidth based on a combination of modal variance and global fluctuation. Use a frequency-duration-intensity hybrid decision rule to determine the proportion exceeding the upper threshold and the duration of continuous over-threshold occurrences.

[0042] The graded early warning triggering criterion unit is used to determine the graded early warning triggering criterion based on the time-varying global threshold, sensitivity bandwidth, the proportion of the upper threshold, and the duration of continuous over-threshold. The graded early warning triggering criterion is used to classify the time series score to obtain the early warning level.

[0043] The time-series early warning level unit is used to introduce an autoregressive correction factor at the decision-making level so that the early warning level depends on the state at the previous moment, thus obtaining an early warning level that depends on the state at the previous moment.

[0044] The output unit is used to output the warning level based on the previous moment.

[0045] Another object of the present invention is to provide an electronic device comprising: at least one processor, at least one memory, and a communication interface. The processor, memory, and communication interface communicate with each other. The memory stores program instructions executable by the processor, which invokes the program instructions to execute the aforementioned multimodal-based adaptive hierarchical method for psychological states.

[0046] Compared with the prior art, the present invention has the following advantages: 1. Multimodal comprehensive perception: The system uses a combination of multiple sensors to collect multimodal data on children, including but not limited to facial expressions, tone of voice, and body movements, to more comprehensively and objectively assess children’s psychological state and make up for the shortcomings of single-modality perception.

[0047] 2. Early warning function: When the system detects persistent negative emotions or potential risky behaviors in children, it can issue tiered early warnings, promptly notify guardians, and thus achieve early intervention to prevent the problem from worsening.

[0048] 3. Universality and usability: This invention can be deployed without a professional laboratory environment, has low cost, and is convenient for widespread use. Attached Figure Description

[0049] Figure 1 The flowchart shows the main body of the adaptive grading method for mental states based on multimodality. Figure 2 A flowchart; Figure 3 This is a schematic diagram of a three-order interaction system; Figure 4 For the early warning classification process; Figure 5 This is a schematic diagram of the terminal structure. Detailed Implementation

[0050] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that these examples are for illustrative purposes only and are not intended to limit the scope of the invention. After reading this invention, any modifications of the invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.

[0051] Example 1 This embodiment provides a multimodal adaptive grading method for mental states, such as... Figure 1-5 As shown, it includes the following steps: Step 1, at each time step The system collects four modalities from the subject: visual, speech, behavioral, and textual data.

[0052] The terminal incorporates sensors such as cameras and microphones to collect children's visual data (facial expressions), auditory data (audio events and voice text), and tactile / motor data (touch operations). This multimodal input data is entered into the system for subsequent analysis.

[0053] Step 2: The collected visual, speech, behavioral and text modal inputs are processed by each modal recognition sub-module to obtain the single-modal confidence score.

[0054] Single-mode confidence:

[0055] in, Indicates the confidence level of a single mode. This represents the Sigmoid activation function. This represents the recognition sub-model of the m-th modality. This represents the recognition sub-model of the m-th modality. This indicates a modal identifier used to distinguish different data modalities. Representing visual modality, Represents speech modality, Represents behavioral modality, Represents the text modality.

[0056] Step 3: To reflect the differences in dynamic reliability between modes, the obtained single-mode confidence scores are used to determine the dynamic stability score of each mode by introducing an adaptive weight allocation mechanism based on a sliding window.

[0057] Defined in length of The mean and variance within the time window are:

[0058] Based on this, the dynamic stability score for each mode is calculated:

[0059] in, This represents the dynamic stability score for each mode. Indicates the length of the sliding window. This represents the single-mode confidence of the ti-th mode at time t. This represents the sliding window variance of the confidence level of the m-th modality at time t. This represents the average value of the sliding window for the m-th mode at time t. This represents the dynamic stability score of the m-th mode at time t. The third term represents the adjustable weighting coefficients and reflects the inhibition factor of the modal change rate.

[0060] Step 4: Apply exponential normalization to the dynamic stability score of each modality to obtain time-varying weights and thus obtain the multimodal fusion score.

[0061] Time-varying weights are obtained using exponential normalization:

[0062] The multimodal fusion score is defined as:

[0063] in, Indicates the multimodal fusion score. Represents the weight sensitivity coefficient. Indicates the multimodal fusion score. Represents the weight sensitivity coefficient. This represents the dynamic stability score of the k-th mode at time t. Indicates the modal traversal index. This represents the dynamic stability score of the k-th mode at time t.

[0064] Step 5: In order to suppress instantaneous fluctuations and introduce time inertia, a two-stage time smoothing mechanism is used to obtain the final time series score used for early warning determination on the obtained multimodal fusion score.

[0065] A two-stage time smoothing mechanism is adopted to achieve a tiered fusion of short-term and long-term trends. The first stage is short-term exponential smoothing:

[0066] The second is the long-term adaptive trend term:

[0067] in, This represents the score after short-term exponential smoothing in the first stage at time t. Indicates the short-term smoothing coefficient. Indicates the time series score. This represents the score after long-term adaptive trend smoothing in the second stage at time t. Indicates the long-term smoothing coefficient. .

[0068] Step 6, Introduce the baseline mean With baseline variance Determine the time-varying global threshold. Analyze the sensitivity bandwidth based on a combination of modal variance and global fluctuation. Use a frequency-duration-intensity hybrid decision rule to determine the proportion exceeding the upper threshold and the duration of continuous over-threshold occurrences.

[0069] To ensure adaptability to individual baselines, a baseline mean is introduced. With baseline variance Determine the time-varying global threshold.

[0070]

[0071] in, Indicates the time-varying global threshold. Indicates the baseline mean. This represents the baseline variance adjustment factor. Indicates the baseline variance. This represents the rate of change adjustment coefficient. This represents the time series score used for early warning determination at time t. Adjust the weights for stability and rate of change separately. Parameters Through finite difference form Approximation reflects the rate of change in scores.

[0072] To describe the width of the risk range, the sensitivity bandwidth is determined jointly based on modal variance and global volatility.

[0073]

[0074] in, This represents the sensitivity bandwidth at time t. This represents the modal variance adjustment coefficient. This represents the time-varying weight of the m-th mode at time t. This represents the sliding window variance of the confidence level of the m-th modality at time t. This represents the single-mode confidence time series of the m-th mode. This represents the baseline variance adjustment factor. The normal baseline variance representing the mental state of the identified individual. Represents the local minimum constant. This allows the threshold bandwidth to adaptively adjust with modal fluctuations and overall emotional fluctuations.

[0075] Let the most recent In the sampling, the proportion exceeding the upper threshold is:

[0076] Simultaneously define the duration of continuous threshold exceedance:

[0077] in, This represents the proportion above the threshold. This indicates the number of samples used for threshold determination. This indicates that the indicator function takes a value of 1 when the condition is met, and 0 otherwise, and is used to count the number of times the threshold is exceeded. Let represent the time series score at time ti. Indicates the time-varying global threshold. This represents the sensitivity bandwidth at time t. Indicates the duration of continuous exceedance. Indicates the sampling time interval.

[0078] Step 7: Determine the graded early warning triggering criteria based on the time-varying global threshold, sensitivity bandwidth, the proportion of the upper threshold, and the duration of continuous over-threshold. Obtain the early warning level by graded early warning based on the time series score through the graded early warning triggering criteria.

[0079] The criteria for triggering a tiered early warning are as follows:

[0080] in, Indicates the time-varying global threshold. Indicates the sensitivity bandwidth. This represents the proportion above the threshold. Indicates the duration of continuous exceedance. For frequency threshold, This is the time threshold. If continuous detection satisfies the three-level conditions in the above formula, the system triggers a higher-level intervention process.

[0081] Step 8: To further improve system stability, an autoregressive correction factor is introduced at the decision-making level to make the warning level dependent on the state at the previous moment, thus obtaining the warning level dependent on the state at the previous moment.

[0082]

[0083] in, This represents the final warning level at time t, which depends on the state at the previous time. This indicates the k-th level warning. This represents the autoregressive correction factor for the k-th level warning. Indicates an indicator function, Let represent the probability that the condition for a level k warning is met at time t. For the first The autoregressive design ensures the temporal continuity and physical rationality of changes in early warning levels, reducing instantaneous jumps, by meeting the level conditions.

[0084] This invention assesses children's emotional and behavioral states by fusing multi-dimensional data, including visual, speech, audio events, and actions. If at least one modality scores exceed a threshold and this is consistent over time, it is considered a consistent support result; otherwise, it is a low-confidence result and enters the auxiliary decision chain.

[0085] This invention provides two triggering methods for the early warning mechanism.

[0086] ① Method 1 - Comprehensive Emotional State Early Warning: The system continuously collects and analyzes children's emotional state data, generating a visualized emotional trend change chart. By setting an emotional baseline and trend judgment algorithm, when the system detects that a child's emotional indicators show a persistent negative deviation (e.g., emotional scores are significantly lower than the average level or show a monotonous downward trend over multiple consecutive assessment periods), it determines that there is a potential mental health risk. At this time, the system will automatically trigger a non-emergency early warning process, and through the associated mobile application (APP) or SMS platform, it will push information including an overview of the child's recent emotional profile (such as emotional fluctuation patterns and correlations with key events) and a quantitative assessment report of emotional trends to the child's guardian.

[0087] ② Method Two - Real-time Emotional State Warning: The system integrates multi-source behavioral signals to construct a multimodal psychological state recognition model, enabling efficient assessment of a child's immediate emotional state. Dynamic warning thresholds are set for different emotional dimensions. When real-time analysis results indicate that the child's current emotional intensity or a specific negative emotion (such as extreme sadness, anger, or anxiety) momentarily exceeds a preset safety threshold, the system will immediately trigger the highest-level emergency alarm mechanism. This mechanism will prioritize sending an immediate alarm message to the guardian via the fastest and most reliable communication channel (such as a strong push notification from the app, SMS, or a pre-set emergency contact number), containing a description of the emergency emotional state, a timestamp of the event, and suggested initial response measures.

[0088] This invention enables multimodal emotion fusion calculation and hierarchical early warning decision-making in resource-constrained environments. The introduction of dynamic thresholds, variance correction, and two-layer time smoothing makes the model more adaptable and robust, significantly improving the accuracy and stability of early warning judgment.

[0089] In another embodiment of the present invention, an electronic device is provided, comprising: at least one processor, at least one memory, and a communication interface. The processor, memory, and communication interface communicate with each other. The memory stores program instructions executable by the processor, which invokes the program instructions to execute the described multimodal-based adaptive grading method for mental states.

[0090] Another embodiment of the present invention provides a multimodal adaptive grading system for psychological states, employing the aforementioned multimodal adaptive grading method for psychological states, comprising a data acquisition unit, a modality recognition sub-model unit, a dynamic stability unit, a multimodal fusion unit, a temporal scoring unit, a warning parameter determination unit, a graded warning triggering criterion unit, a temporal warning level unit, and an output unit, wherein: The acquisition unit is used to acquire four modalities of the subject: visual, speech, behavior, and text.

[0091] The modality recognition sub-model unit is used to collect four modalities: visual, speech, behavior, and text. The single-modality confidence is calculated by each modality recognition sub-model.

[0092] The dynamic stability unit is used to determine the dynamic stability score of each mode by introducing an adaptive weight allocation mechanism based on a sliding window, based on the obtained single-mode confidence.

[0093] The multimodal fusion unit is used to obtain time-varying weights by exponentially normalizing the dynamic stability scores of each obtained modality, thereby obtaining the multimodal fusion score.

[0094] The time-series scoring unit is used to apply a two-stage time smoothing mechanism to the obtained multimodal fusion score to obtain the final time-series score used for early warning determination.

[0095] The early warning parameter determination unit is used to introduce a baseline mean. With baseline variance Determine the time-varying global threshold. Analyze the sensitivity bandwidth based on a combination of modal variance and global fluctuation. Use a frequency-duration-intensity hybrid decision rule to determine the proportion exceeding the upper threshold and the duration of continuous over-threshold occurrences.

[0096] The graded early warning triggering criterion unit is used to determine the graded early warning triggering criterion based on the time-varying global threshold, sensitivity bandwidth, the proportion of the upper threshold, and the duration of continuous over-threshold. The graded early warning triggering criterion is used to classify the time series score to obtain the early warning level.

[0097] The time-series early warning level unit is used to introduce an autoregressive correction factor at the decision-making level so that the early warning level depends on the state at the previous moment, thus obtaining an early warning level that depends on the state at the previous moment.

[0098] The output unit is used to output the warning level based on the previous moment.

[0099] The system employs algorithms such as image recognition, voice emotion analysis, and behavior recognition to process sensor data. The visual analysis component identifies facial expressions, the voice analysis component extracts audio events and speech-to-text, and the touch / motion component detects children's actions on toys. Through multimodal information fusion technology, the system integrates these sensory results to form an assessment of the child's current attention span, emotional state, and other psychological indicators (referencing existing multimodal fusion methods). The assessment results are mapped to a pre-set mental model to obtain the child's immediate psychological evaluation value.

[0100] This system also features an adaptive interactive control unit: based on the aforementioned assessment results and preset psychological interaction strategies, and through the fusion analysis of real-time psychological assessment and historical emotional state data, it can generate content for interaction with children, such as situational stories and interactive voice messages, to engage in beneficial interactive behaviors. For example, if anxiety is assessed in a child, relaxing content is played and deep breathing is prompted; if a child is assessed as depressed, a mini-game is initiated to shift their mood. The intelligent control module interacts with children through an eye-screen display and speakers, simulating virtual character dialogues or providing task feedback to enhance fun and immersion. The interactive content can also be personalized based on the child's age, interests, preferences, and historical data.

[0101] Real-time Interaction and Tiered Early Warning Unit: This unit detects a child's emotional state and provides feedback on the assessment results. If the assessment results fall within a preset danger threshold, an early warning mechanism is triggered. For general emotional state changes (Level 0 or 1), the system uses personalized / non-emergency proactive interaction methods such as voice to help the child self-regulate. For emergency warnings (Level 2), the system automatically generates a prompt message and sends alerts and suggestions to the guardian via a mobile app or SMS. It also enters emergency proactive interaction mode. For crisis warnings (Level 3), the system automatically generates an alarm message and sends multiple alerts (based on sensitive words) to the guardian and public institutions via a mobile app or SMS, and enters crisis proactive interaction mode.

[0102] The psychological management unit maintains a personalized profile for each child, recording their daily emotional fluctuations. Utilizing storage and data processing units, it regularly updates children's psychological profiles and continuously optimizes interaction strategies. Through long-term tracking, the system can evaluate interaction effectiveness, statistically analyze trends, and provide guardians with periodic psychological reports. Furthermore, the system supports cloud expansion, enabling collaboration with a child psychological management cloud platform, uploading and analyzing local assessment results to further enhance reliability.

[0103] Unique Zodiac Character System: This invention introduces a unique "Zodiac Character System" into the terminal. Each terminal has preset differentiated personality and zodiac attributes, allowing users to choose their playmate character based on personal preferences. By combining preset personality traits with real-time interactive scenarios, the system generates cognitively and emotionally driven dialogue strategies, achieving anthropomorphic and differentiated deep companionship and enhancing long-term interaction and engagement between children and the terminal.

[0104] The accompanying growth memory unit can record emotional data for a long time without deleting related records. This invention further incorporates a personalized companionship mechanism based on interactive memory. The system records and analyzes children's historical interaction data and preference characteristics to form a dynamic psychological profile, and generates growth companionship strategies based on this profile to achieve nurturing companionship.

[0105] Data upload and sentiment profile generation unit: The terminal summarizes the event summary, timeline, key modality data and aggregated sentiment scores of the current session based on local timed caching. After anonymization and encryption, it automatically initiates the upload when the available Wi-Fi network is restored. After receiving the data, the cloud stores the summary, merges it with existing individual profiles and generates or updates periodic sentiment profiles. The cloud analysis results can drive model updates and the distribution of interaction suggestions. Both upload and profile change operations are recorded to support subsequent traceability and compliance management.

[0106] Power-on / off and standby strategies: When the device is powered on or woken up, it completes local file loading, sensor self-testing, and initialization of the sensing subsystem, enters a continuous detection and interaction ready state, and retains the context memory of the current interaction until power-off; when the user actively shuts down or the standby strategy is triggered by prolonged idle time, the system terminates the real-time sensing and interaction process in an orderly manner according to predetermined steps, saves the current cache, completes event summary merging, performs local encrypted writing, and marks the status as offline; in abnormal power outage scenarios, the system relies on periodic checkpoints to ensure the consistency and recoverability of recorded elements, and ensures the integrity of data before and after power-off and subsequent traceability.

[0107] This invention constructs a three-level responsive children's emotional interaction system, which achieves dynamic and personalized psychological support through the following hierarchical mechanism: ① Personalized and proactive interaction: After powering on, the terminal loads the emotional profile (including user preference data) from the cloud and performs dynamic optimization to achieve personalized and proactive care driven by emotional data.

[0108] ② Real-time personalized interaction: Based on multi-sensor fusion technology, in the comprehensive emotional state early warning mechanism, when the early warning level is 0 (no early warning), the user preference data is updated based on the current real-time emotional state feature value and the historical emotional state feature library, and real-time personalized interaction is generated based on the updated user preference data.

[0109] ③ Non-urgent proactive interaction: In the comprehensive emotional state early warning mechanism, when the early warning level is level 1 (non-urgent early warning), the system integrates children's long-term emotional data, generates structured emotional profiles through the psychological state analysis engine, and automatically matches and generates non-urgent proactive interaction content. The aim is to promote children's emotional resilience development and alleviate negative emotions through data-driven approaches.

[0110] ④ Emergency Proactive Interaction: In the comprehensive emotional state early warning mechanism, when the early warning level is level 2 (emergency early warning), the system will call the professional psychological knowledge base, automatically match the preset emergency interaction behavior template, and generate emergency proactive interaction content.

[0111] ⑤ Crisis-based Proactive Interaction: In the comprehensive emotional state early warning mechanism, when the early warning level is 3 (crisis warning), and the multimodal emotion recognition system detects that real-time emotional indicators exceed the clinical early warning threshold (such as acute anxiety or agitation), the crisis interaction protocol is immediately activated. The system combines the current emotional physiological representation (facial movement unit intensity) and situational context, calls upon a professional psychological knowledge base to generate immediate suggested behavioral plans, and implements emergency guidance through a high-priority interaction channel to block the path of emotional deterioration and prevent potential behavioral risks.

[0112] This invention constructs a three-level responsive interactive linkage mechanism. When long-term trend analysis triggers a non-emergency warning (Level 1), the system activates an addressable RGB breathing light module. By dynamically generating a healing spectrum that conforms to color psychology theory, it implements non-invasive phototherapy interaction during periods of persistent negative emotional deviation in children. Its dual-channel value lies in providing children with subconscious emotional regulation stimulation and building a visual emotional state indicator for guardians, effectively compensating for the blind spots of mobile terminal information push. When multimodal real-time recognition triggers an emergency warning (Level 2), the system simultaneously executes a personalized emotional interaction protocol and a structured warning push, realizing a coordinated response of real-time guidance on the user end and precise alarm on the management end. When real-time emotional indicators exceed the clinical safety threshold and trigger a crisis warning (Level 3), the system automatically connects to the social support network while initiating a crisis stabilization interaction program and pushing emergency alarm information.

[0113] This invention organically combines multimodal recognition, human-computer interaction, and psychological interaction, not only meeting the real-time and convenient needs of children's psychological interaction in a home environment, but also expanding the application scope of existing technologies in the field of children's psychological monitoring and early warning, as detailed below: 1. Improved Perception Accuracy and Targeted Interaction: Employing a multimodal fusion assessment mechanism, this system can capture subtle changes in children's emotions and attention. Compared to single-sensor solutions, this approach can more accurately determine a child's psychological state, thereby making subsequent interactions more targeted. For example, the system not only recognizes the sound of crying but also senses facial expressions of sadness, comprehensively judging whether comforting is needed.

[0114] 2. Real-time interaction and closed-loop feedback: This invention combines gamified interaction with intelligent feedback, providing guidance and encouragement to children during the interaction process, promptly correcting emotional or behavioral deviations, and enhancing children's participation and psychological interaction effects. As mentioned earlier, some existing systems utilize similar human-computer game interaction environments to allow children to explore scenarios and improve their social skills in a free space. This solution further integrates real-time emotional interaction, forming a closed loop of recognition-interaction-evaluation.

[0115] 3. Guardian / Public Institution Involvement and Early Warning: With the added early warning function, guardians can be informed of children's psychological risks immediately without the need for 24 / 7 supervision. Referring to educational practices, when the system detects self-harm or depressive expressions in children, the crisis early warning module can immediately trigger a red alert and notify guardians and community public institutions. This invention applies a similar early warning principle to home terminals, issuing alarms for abnormal emotional fluctuations, achieving timely control and protection.

[0116] 4. Universality and Economy: Compared with existing solutions that require specialized venues and equipment, the interactive terminal of this invention can be manufactured as a home toy device, which is small in size and easy to set up. Software upgrades can support various game content without frequent hardware replacements, significantly reducing long-term operating costs and user burden, and improving the accessibility of psychological interaction.

[0117] 5. Continuous Personalized Management: Leveraging long-term data management capabilities, the system can be tailored to the individual needs of different children, thereby achieving personalized interaction. Existing research also indicates that a child-centered learning and interaction approach can make training more effective. This invention combines the management module with the interaction module, ensuring engagement while addressing individual differences among children, enabling better tracking of progress and improved interaction effectiveness.

[0118] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A multimodal adaptive hierarchical method for psychological states, characterized in that, Includes the following steps: Step 1: Collect four modalities of the subject: visual, speech, behavior, and text. Step 2: The collected visual, speech, behavioral, and text modalities are used to calculate the single-modal confidence scores through the respective modal recognition sub-models; Step 3: The obtained single-modal confidence scores are used to determine the dynamic stability score of each modality by introducing an adaptive weight allocation mechanism based on a sliding window. Step 4: Apply exponential normalization to the dynamic stability score of each modality to obtain time-varying weights and thus the multimodal fusion score; Step 5: Apply a two-stage time smoothing mechanism to the obtained multimodal fusion score to obtain the final time-series score used for early warning determination; Step 6, Introduce the baseline mean With baseline variance Determine the time-varying global threshold; determine the sensitivity bandwidth based on modal variance and global fluctuation; and determine the proportion exceeding the upper threshold and the duration of continuous over-threshold using a frequency-duration-intensity hybrid judgment rule. Step 7: Determine the graded early warning triggering criteria based on the time-varying global threshold, sensitivity bandwidth, the proportion of the upper threshold, and the duration of continuous over-threshold. Then, use the graded early warning triggering criteria to classify the time series score and obtain the early warning level. Step 8: Introduce an autoregressive correction factor at the decision-making level to make the warning level dependent on the state at the previous moment, thus obtaining the warning level dependent on the state at the previous moment.

2. The adaptive hierarchical method for mental states based on multimodality as described in claim 1, characterized in that: The formula for calculating the confidence level of a single mode is: in, Indicates the confidence level of a single mode. This represents the Sigmoid activation function. This represents the recognition sub-model of the m-th modality. This represents the recognition sub-model of the m-th modality. This indicates a modal identifier used to distinguish different data modalities. Representing visual modality, Represents speech modality, Represents behavioral modality, Represents the text modality.

3. The adaptive grading method for mental states based on multimodality according to claim 2, characterized in that: The formula for calculating the dynamic stability score for each mode is as follows: in, This represents the dynamic stability score for each mode. Indicates the length of the sliding window. This represents the single-mode confidence of the ti-th mode at time t. This represents the sliding window variance of the confidence level of the m-th modality at time t. This represents the average value of the sliding window for the m-th mode at time t. This represents the dynamic stability score of the m-th mode at time t. This is the adjustable weighting coefficient.

4. The adaptive hierarchical method for mental states based on multimodality according to claim 3, characterized in that: The formula for calculating the multimodal fusion score is: in, Indicates the multimodal fusion score. Represents the weight sensitivity coefficient. This represents the dynamic stability score of the k-th mode at time t. Indicates the modal traversal index. This represents the dynamic stability score of the k-th mode at time t.

5. The adaptive hierarchical method for mental states based on multimodality according to claim 4, characterized in that: The formula for calculating the time series score is: in, This represents the score after short-term exponential smoothing in the first stage at time t. Indicates the short-term smoothing coefficient. Indicates the time series score. This represents the score after long-term adaptive trend smoothing in the second stage at time t. This represents the long-term smoothing coefficient.

6. The adaptive grading method for mental states based on multimodality according to claim 5, characterized in that: The formula for calculating the time-varying global threshold is: in, Indicates the time-varying global threshold. This represents the normal baseline mean of the psychological state of the person being identified. This represents the baseline variance adjustment factor. The normal baseline variance representing the mental state of the identified individual. This represents the rate of change adjustment coefficient. This represents the rate of change of the time series score at time t; The formula for calculating sensitivity bandwidth is: in, This represents the sensitivity bandwidth at time t. This represents the modal variance adjustment coefficient. This represents the time-varying weight of the m-th mode at time t. This represents the sliding window variance of the confidence level of the m-th modality at time t. This represents the single-mode confidence time series of the m-th mode. This represents the baseline variance adjustment factor. The normal baseline variance representing the mental state of the identified individual. Represents the minimum constant; The formulas for calculating the proportion exceeding the upper threshold and the duration of continuous over-threshold are as follows: in, This represents the proportion above the threshold. This indicates the number of samples used for threshold determination. This indicates that the indicator function takes a value of 1 when the condition is met, and 0 otherwise, and is used to count the number of times the threshold is exceeded. Let represent the time series score at time ti. Indicates the time-varying global threshold. This represents the sensitivity bandwidth at time t. Indicates the duration of continuous exceedance. Indicates the sampling time interval.

7. The adaptive grading method for mental states based on multimodality according to claim 6, characterized in that: The formula for calculating the warning level is: in, Indicates the time-varying global threshold. Indicates the sensitivity bandwidth. This represents the proportion above the threshold. Indicates the duration of continuous exceedance. This indicates the proportion exceeding the upper threshold (F). t The criterion for determining the threshold is T, where T represents the duration of continuous over-threshold (D). t The criteria for determining ( ).

8. The adaptive hierarchical method for mental states based on multimodality according to claim 7, characterized in that: The formula for calculating the warning level based on the previous moment is: in, This represents the final warning level at time t, which depends on the state at the previous time. This indicates the k-th level warning. This represents the autoregressive correction factor for the k-th level warning. Let represent the probability that the condition for a level k warning is met at time t. Indicates an indicator function.

9. A multimodal adaptive classification system for mental states, employing the multimodal adaptive classification method for mental states as described in claim 1, characterized in that: It includes a data acquisition unit, a modality recognition sub-model unit, a dynamic stability unit, a multimodal fusion unit, a time-series scoring unit, a warning parameter determination unit, a graded warning triggering criterion unit, a time-series warning level unit, and an output unit, wherein: The acquisition unit is used to acquire four modalities of the subject: visual, speech, behavior, and text. The modality recognition sub-model unit collects four modalities: visual, speech, behavior, and text. Each modality recognition sub-model then calculates the single-modality confidence score. The dynamic stability unit is used to determine the dynamic stability score of each mode by introducing an adaptive weight allocation mechanism based on a sliding window, based on the obtained single-mode confidence. The multimodal fusion unit is used to obtain time-varying weights by exponentially normalizing the dynamic stability scores of each obtained modality, thereby obtaining the multimodal fusion score. The time-series scoring unit is used to apply a two-stage time smoothing mechanism to the obtained multimodal fusion score to obtain the final time-series score used for early warning determination; The early warning parameter determination unit is used to introduce a baseline mean. With baseline variance Determine the time-varying global threshold; determine the sensitivity bandwidth based on modal variance and global fluctuation; and determine the proportion exceeding the upper threshold and the duration of continuous over-threshold using a frequency-duration-intensity hybrid judgment rule. The graded early warning triggering criterion unit is used to determine the graded early warning triggering criterion based on the time-varying global threshold, sensitivity bandwidth, the proportion of the upper threshold and the duration of continuous over-threshold, and to obtain the early warning level by graded early warning based on the time series score through the graded early warning triggering criterion. The time-series early warning level unit is used to introduce an autoregressive correction factor at the decision-making level so that the early warning level depends on the state at the previous moment, thus obtaining an early warning level that depends on the previous moment. The output unit is used to output the warning level based on the previous moment.

10. An electronic device, characterized in that, include: At least one processor, at least one memory, and a communication interface; The processor, memory, and communication interface communicate with each other; The memory stores program instructions that can be executed by the processor, which invokes the program instructions to execute the multimodal adaptive grading method for mental states as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Human-machine interaction multi-mode early intervention system for improving social interaction capacity of autistic children

    CN102354349A