Education robot-based intelligent system and method for fusing mental health evaluation in interest scene
Through anonymous design, quantum hybrid encryption, physical isolation and other technical means, combined with multimodal biometric collection and intelligent dialogue guidance, the problems of privacy protection and assessment efficiency in the mental health assessment of middle school students in mountainous schools have been solved, and efficient and secure remote mental health assessment has been achieved.
Patent Information
- Application Number
- CN202510772304.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies in student mental health assessments have problems such as insufficient privacy protection, inability to adapt to individual differences, and inability to achieve real-time dynamic feedback. Especially in mountainous schools where real experts are difficult to reach, there is a lack of effective remote mental health assessment methods.
It adopts a five-level security system that includes anonymous design, quantum hybrid encryption, blockchain evidence storage, and physical isolation. Combined with multimodal biometric collection and embedded security modules, it uses educational robots to conduct mental health assessments. It utilizes federated learning and privacy protection mechanisms to achieve data anonymization and encrypted transmission, ensuring student privacy. It also conducts real-time assessments through intelligent dialogue guidance and dynamic emotional regulation.
It has achieved privacy protection and efficient assessment of students' mental health in mountain schools, reduced the incidence of psychological crisis events, improved the accuracy and acceptability of assessments, ensured data security and privacy, and complied with medical ethics requirements.
Smart Images

Figure CN120636769A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the application technology of artificial intelligence in the field of psychological education, and involves the cross-technical field of psychological health analysis and data encryption. Background Art
[0002] 1. The severity of student mental health issues: A 2023 survey by the China Youth Research Center revealed that 32% of middle school students and 40% of college students experience psychological distress. In extreme cases, self-harm and suicide are frequent, posing a serious threat to their lives. According to a study published by the China Social Sciences Network in November 2024, the detection rates of self-harm among junior high school, high school, and college students in China were 22%, 22.8%, and 16.2%, respectively. These data indicate the urgency of establishing a scientific system for early identification and intervention of mental health risks.
[0003] 2. According to the "2021-2022 Mental Health Literacy Survey Report" by the Institute of Psychology, Chinese Academy of Sciences, 68% of parents are resistant to psychological assessments, mainly worried that "schools will label their children" or "privacy information will be improperly used."
[0004] 3. The critical role of early warning: Research shows that "early detection and early intervention" can effectively reduce the risk of psychological crises. The Institute of Psychology, Chinese Academy of Sciences (2024) published the "Blue Book on Mental Health," emphasizing that timely identification of psychological problems can significantly improve students' social adaptability and reduce long-term negative impacts. Tracking data from Beijing Normal University (2023) shows that a systematic early warning mechanism can reduce the incidence of psychological crises by 40%.
[0005] 4. In mountainous schools, mental health education is relatively scarce, and it is not convenient for real experts to go to mountainous schools. Therefore, it is a good way for real experts to use physical educational robots to carry out remote psychological education, and to use robots to detect and intervene in mentally unhealthy students early.
[0006] 5. Traditional mental health assessments have three major technical bottlenecks: They are hosted by schools, are real-name tests, and explicit questionnaires trigger psychological defense mechanisms in participants; static question banks are difficult to adapt to individual differences; and manual analysis cannot provide real-time dynamic feedback. 6. Disclaimer: This invention does not involve medical diagnosis. It merely aims to identify students suspected of mental health problems from a third-party perspective and provide early warning suggestions to parents, who voluntarily adopt them and seek diagnosis and treatment at a formal medical institution. Summary of the Invention
[0007] The present invention aims to at least solve the technical problems existing in the existing technology or related technology, and to conduct psychological assessments based on students' personalized interests and hobbies by intelligently integrating psychological health test questions into educational robots. Summary of the Invention
[0008] 1. Strict privacy and security protection measures 1. Overview and Business Process 1.1.1 In the early stages, it is necessary to ensure that parents and students are fully aware of the system's functions and operation methods, and only after signing a voluntary agreement can the mental health re-testing system be activated.
[0009] 1.1.2 The entire system is anonymous and does not contain any personal information such as names or phone numbers. All data is linked and transmitted through the student's unique ID, protecting student privacy from the source.
[0010] 1.1.3 Live online experts review an anonymous list of students identified as exhibiting psychological abnormalities. Once confirmed, a physical educational robot on campus is triggered to conduct a mental health retest on the selected students. This method can also be used to conduct mental health testing for students who are resistant to traditional mental health scales. Software-based educational robots can also be used for mental health testing, but without biometric data, biometric data can be obtained through integration with smart wristbands.
[0011] 1.1.4 The campus local server stores the facial feature codes of the school's students and converts them into facial feature codes in an irregular manner. The facial feature code is associated with the student's unique code ID. All associated data is desensitized and does not contain any personalized identifiers such as name, mobile phone number, and parent information. The above-mentioned associated information is encrypted using the National Security Office algorithm combined with quantum key encryption, stored on the campus local server, and blockchain technology is used to prevent encrypted data and checksums from being tampered with. Once the system detects a data change, it will immediately notify the network gate to disconnect access to the stored content, and only manual recovery is supported. In addition, after the facial feature code is associated with the code ID, the system will delete the original facial image information within 10 seconds.
[0012] (Note: In addition to facial feature codes, one or more biometric feature codes such as fingerprint codes, voiceprint codes, iris codes, and brainwave codes may also be used. Facial feature codes, which are more mature in technology, are preferred.) 1.1.5 The physical educational robot only receives the list containing the code ID and then initiates the retest task. The robot uses facial recognition technology to generate a facial feature code, converts it into a random facial feature code, and then associates it with the student's unique code ID.
[0013] 1.1.6. Educational robots can be deployed at the entrances of campus gates, dormitories, and teaching buildings. A federated learning-driven privacy protection mechanism can be used to build a horizontal federated learning framework. Each educational robot on campus can participate in model training as a client. Differential privacy noise injection (ε=0.5, δ=1e-5) is used to protect local gradient update parameters.
[0014] 1.1.7. After the physical educational robot completes the identification of the retest subject using facial recognition technology, it shall immediately delete the original facial image, recognition data, and other information.
[0015] 1.1.8 Data management adheres to the principle of minimization: Normal mental health data will not be stored and will only be marked as "normal." Data that requires storage will be encrypted using the National Security Administration's algorithm combined with quantum key encryption and stored on a local campus server. Physical isolation is implemented via a network gateway on a daily basis, with network access only when data is in use. Blockchain technology is utilized throughout the entire process to prevent data tampering and monitor intrusions in real time. Any anomalies detected will be immediately notified to the network gateway, physically disconnecting information transmission.
[0016] 1.1.9. Authority Control: The school cannot obtain test results without parental authorization, in compliance with the requirements of the Personal Information Protection Law.
[0017] 1.2. Technical Solution This invention builds a five-level security system consisting of "front-end anonymous collection - dynamic feature conversion - quantum hybrid encryption - blockchain evidence storage - physical isolation control". The specific technical solution includes the following innovative details: 1.2.1. Multimodal Biometric Anonymization System (1) The biometric feature acquisition module is equipped with a hardware-level encryption chip and uses the ISO / IEC 19794 standard to extract facial features. The original feature vector is nonlinearly transformed by a chaotic random number generator to generate a 128-bit irreversible facial feature code. (2) The dynamic association engine uses double hash chain technology to associate the facial feature code with the student’s unique ID (USIC): The first-layer hash: H1=SM3(facial feature code||timestamp); the second-layer hash: H2=SM3(USIC||institution code||random salt value); the association relationship is stored as a mapping table: H1→H2; (3) The multimodal compatible interface integrates the voiceprint feature extraction algorithm, uses MFCC (Mel-Frequency Cepstral Coefficients) coefficients combined with the GMM-UBM (Gaussian Mixture Model-Universal Background Model) model to generate voiceprint codes (Voiceprint), and supports fusion verification with iris features (Iris Code).
[0018] 1.2.2. Quantum-enhanced hybrid encryption system (1) Data encryption uses a layered encryption architecture: The bottom layer uses the SM4 algorithm (a block cipher algorithm developed by the National Security Bureau) for data block encryption. The middle layer uses the NTRU (Number Theory Research Unit) post-quantum cryptography system, which is resistant to the Shor algorithm, to encapsulate the key. The top layer implements quantum key distribution (QKD) through the BB84 protocol, with a key update period of ≤30 seconds. (2) The blockchain evidence storage module deploys a consortium chain architecture, sets up three types of verification nodes: education management agencies, school nodes, and health departments, adopts an improved PBFT (Practical Byzantine Fault Tolerance) consensus algorithm, and sets the Byzantine fault tolerance threshold to ≥33%; (3) Encrypted data is stored in fragments using Reed-Solomon coding, which divides the data into n fragments (n≥5). The original data can be restored by satisfying any k fragments (k=3).
[0019] 1.2.3. Intelligent physical isolation system In addition to network security equipment such as firewalls and intrusion detection, add network gateways for physical isolation.
[0020] (1) The Data Diode uses an FPGA-based hardware design to implement unidirectional transmission at the physical layer: Transmitter: Integrated 850nm VCSEL laser emission array, transmission rate ≥10Gbps; Receiver: Equipped with InGaAs photodetector, bit error rate ≤1×10⁻¹²; (2) The anomaly monitoring system deploys a Deep Packet Inspection (DPI) engine with a built-in threat signature library containing the latest vulnerability signatures such as CVE-2023. When an anomaly is detected, the smart contract automatically triggers the following actions: Cut off the optical path of the network gate within 0.5 seconds; activate the piezoelectric ceramic crusher to destroy the key storage unit; and send a physical alarm signal through the LoRa wireless module.
[0021] 1.2.4. Educational Robot Safety Execution Unit (1) The embedded security module integrates a national secret level 2 security chip to achieve: The facial recognition algorithm runs within the TrustZone security domain; the signature code conversion process is fully hardware-encrypted; and the storage medium uses FRAM (Ferroelectric RAM) to prevent physical detection. (2) The data self-destruction mechanism uses dual triggering: Timing trigger: Automatically erase temporary data 30 seconds after the task is completed; Abnormal trigger: Activate the data degaussing coil when the shell is detected to be opened.
[0022] (3) The biometric collector is equipped with near-infrared liveness detection and uses a CNN convolutional neural network to identify photo / video attacks. The liveness detection accuracy is ≥99.97%.
[0023] 1.2.5. Dynamic Risk Assessment Subsystem (1) Deploy a Hidden Markov Model (HMM) to analyze the psychological data stream in real time, and trigger the review process when the abnormal probability exceeds a threshold θ (θ=0.85); (2) The risk level assessment matrix uses the fuzzy logic algorithm to comprehensively consider the three-dimensional parameters of the SCL-90 scale data, behavioral analysis data, and physiological indicators; (3) The data desensitization engine supports dynamic desensitization strategies and automatically selects the desensitization intensity based on the visitor role: Teacher side: retains the psychological level code (such as A1 / B2); administrator side: displays the encrypted hash value; expert side: requires quantum key decryption to obtain complete data.
[0024] 1.2.6. Technical Features: (1). The chaotic random number generator is implemented using the Lorenz attractor equation: dx / dt = σ(yx) dy / dt = x(ρ z) y dz / dt = xy βz Where σ=10, ρ=28, β=8 / 3, the initial value is generated by the physical entropy source (2) The consortium chain smart contract contains the following key functions: solidity function dataIntegrityCheck(bytes32 dataHash) public returns(bool) {require(nodeType[msg.sender] == NodeType.Validator); uint confirmations = 0;for (uint i=0; i <validators.length; i++) { if (checkSignature(dataHash, validators[i])) { confirmations++;}} return confirmations >= (validators.length * 2) / 3;} (3). The liveness detection CNN network structure includes an example: Input layer: 256×256 RGB image; Convolutional layer: 5×5 kernel, stride 2, 32 channels; Max pooling layer: 3×3 window; Fully connected layer: 1024 neurons; Output layer: Sigmoid activation function.
[0025] 2. Overview of main business processes 2.1. Voluntary Authorization In the initial phase, parents or students of legal age voluntarily sign an agreement, fully understanding the system's full functionality. The educational robot then conducts an anonymous, comprehensive initial analysis of the students, identifying those suspected of mental health problems. Remote, live experts can use system settings to automatically trigger the educational robot to initiate proactive retesting, or manually trigger the physical robot to initiate proactive retesting.
[0026] Traditional mental health assessments currently available include online mental health assessments, smartphone apps or mini-programs, and paper questionnaires. For students who are reluctant to take traditional mental health assessments, this method can also be used to conduct psychological assessments on topics of interest. Students can also choose to take the integrated mental health assessment anonymously.
[0027] 2.2. Identifying the Suspect When a suspected student appears near an educational robot (e.g., at the entrance of a teaching building or dormitory), the robot uses facial recognition technology to identify the student. Refer to the security section in Section 1. After the information is associated, the system deletes the original facial image and uses only the student's ID for subsequent communication.
[0028] 2.3. Acquisition of biological data The physical educational robot initiates the interaction, greeting the suspected student and inquiring about their willingness to communicate. After the student understands the system's functions and obtains their consent, the educational robot shakes hands with the student. Various sensors and instruments on its hands, body, and head collect real-time physiological data.
[0029] 2.4. Traditional Mental Health Assessment The physical educational robot, combined with the student's anonymous psychological profile, provides a traditional mental health risk analysis scale in the form of multiple-choice questions for the suspected student to answer. Students can complete the questions using a touchscreen or voice dialogue, with the voice answers converted to text in real time. The system then assigns a quantitative score based on the scale's scoring rules, ultimately providing a preliminary assessment of whether a student has a mental health risk.
[0030] 2.5. Implanting Mental Health Test Questions Based on Interests At the same time, the physical educational robot can identify the emotions of suspected students based on anonymous psychological profiles. It can also proactively explore students' interests and hobbies during conversations, raising targeted topics of interest. It can also predict students' emotional trends in real time during conversations and dynamically adjust the conversation content based on the emotion recognition results, naturally incorporating psychological test questions. Because the topics are of interest to the students, the psychological test questions incorporated into the conversation are more easily accepted by the students.
[0031] 2.6. Comprehensive Psychological Assessment The system converts the collected close-range physiological data and conversation content (including psychological test questions) into text information, and the psychological test questions are scored based on a standardized psychological test scale.
[0032] 2.7. Graded warning The results of the mental health retest are output as a three-level risk index (low, medium, and high). If the risk index is high, the system immediately triggers an emergency response process with a live online expert. After the results are confirmed by the live online expert, parents will be advised to accompany the student to a professional institution for diagnosis and medical treatment.
[0033] 3. Overall technical solution 3.1. Student Identification and Privacy Protection Mechanism For the suspected students screened out in the first round of detection, the system automatically deletes the original facial image information and only transmits subsequent information using the code ID (2), combining k-anonymity (k ≥ 3) and differential privacy (ε = 0.5) technology to ensure the security of student identity information.
[0034] Edge computing localization processing uses an embedded artificial intelligence computing platform to deploy a local model to achieve real-time data processing on the device side. The sensitive data retention period is ≤72 hours, ensuring privacy and security from the hardware level.
[0035] 3.2. Multimodal Data Acquisition and Fusion Technology Comprehensive physiological data acquisition educational robot integrated with all-in-one sensor array: Hand sensors: laser Doppler blood flowmeter (measuring microcirculation), flexible strain gauge pressure sensor (grip force distribution), miniature body temperature probe (accuracy ±0.1°C), bioelectric acquisition module (16-channel differential amplification), accelerometer (detecting tremor frequency); Expandable devices: Hot-swappable interfaces support access to external devices such as eye trackers, near-infrared eye trackers (sampling rate ≥ 250Hz, accuracy 0.5°), and electronic noses, enabling multi-dimensional collection of physiological data (blood flow, body temperature, bioelectricity), eye movement data (gaze point heat map, pupil diameter changes), body movements (handshake strength, tremor frequency), and odor information (primarily for detecting alcohol or contraband).
[0036] A multimodal data fusion assessment model builds a mental health assessment system consisting of nine primary indicators and 32 secondary indicators. It uses an improved LightGBM algorithm for outlier detection, reducing the false positive rate to 3.2%. A hybrid LSTM+CNN model fuses time-series physiological data with image and speech features to output emotion classification (anxiety / depression / normal) and confidence scores, enabling in-depth fusion analysis of physiological and behavioral data.
[0037] 3.3. Intelligent Conversation Guidance and Dynamic Emotional Regulation BERT-based emotion prediction and topic guidance utilizes the BERT model to identify emotions in student speech-to-text data, outputting positive, negative, and neutral emotional states in real time. By combining student profiles and interests captured during conversations, conversations are dynamically adjusted: In-depth discussion is facilitated when students are positive, soothing language is used to guide conversations when negative, and relaxed communication is maintained when neutral, achieving intelligent matching of interests, emotions, and topics.
[0038] Reinforcement learning-based dialogue strategy optimization defines a state space encompassing emotional state, response content, and number of conversation turns. It also designs an action space encompassing continuing the conversation, switching topics, asking a test question, and ending the conversation. The dialogue strategy is optimized using a reward function R = α • conversation duration + β • test question completion rate γ • emotional deterioration. A deep learning architecture (Transformer model) with an attention mechanism and a Proximal Policy Optimization (PPO) algorithm is employed to dynamically adjust the dialogue strategy based on student feedback, improving interaction effectiveness and the efficiency of acquiring psychological health information.
[0039] 3.4. Embedded Psychological Assessment and Risk Grading Contextualized test questions are naturally embedded. 12 contextualized test question templates are designed. Using a natural insertion algorithm, psychological test questions covering sleep quality, social interaction, and other dimensions are integrated into the conversation. A multimodal verification mechanism is used to cross-validate test answers with physiological data to improve assessment accuracy.
[0040] Three-Level Risk Index Output and Emergency Response: The system integrates physiological data, conversation content, and test scores to output a three-level risk index of low, medium, and high. A high risk index triggers an emergency response process, where a remote, live expert confirms the results and recommends seeking medical attention. The system does not make any medical diagnoses, in compliance with medical ethics requirements (medical statement).
[0041] 3.5. High-precision attention distraction testing and dynamic adjustment Designing a distraction test based on eye tracking technology: The multi-dimensional evaluation system collects eye movement features such as gaze point heat maps, scanning speed, and pupil diameter changes, and combines them with environmental parameters such as interaction distance and ambient lighting. It extracts attention features through a spatiotemporal joint analysis model (short-time Fourier transform (STFT), hidden Markov model (HMM), dynamic time warping (DTW)), calculates core indicators such as focus maintenance rate (FMR), attention shift cost (ASC), and interference inhibition index (DII) to evaluate students' attention status.
[0042] The hardware layer of the layered processing architecture integrates a near-infrared eye tracker, a wide-angle camera, and a multi-axis force feedback manipulator; the data acquisition layer realizes the synchronous collection of high-precision eye movement data and environmental data; the algorithm layer uses GPU-accelerated deep spatiotemporal three-dimensional residual network (ResNet3D) for real-time processing, ensuring response delay ≤300ms (95% of scenarios).
[0043] A three-level confidence assessment system is established for misjudgment correction and quality control, and the review process is initiated when the confidence level is low (10.3). Multi-period baseline comparison (at least three independent measurements) and a manual review interface (two-factor confirmation mechanism) (10.3) are used, combined with an improved exponentially weighted moving average (EWMA) control chart algorithm (α=0.2, UCL=2.3) to dynamically trigger retesting (12) to ensure the reliability of the assessment results.
[0044] 4. Main innovation: Mental health assessment system based on interest topics 4.1. Business Process and Overview Embedding psychological testing questions within hobby scenarios can significantly improve student acceptance. The mental health assessment scales used in this system can be manually selected by live experts or automatically matched by the system's algorithms. The physical educational robot has been trained on over 100,000 consultation cases. It anonymously collects information on students' interests and hobbies through anonymous mental health records or daily conversations, with the knowledge and permission of the testee.
[0045] 4.1.1. The system of the present invention includes over a hundred professional assessment scales, including the "Mental Health Scale for Chinese Middle School Students (MHS-CISS)" (Professor Zheng Richang, 2002). The specific assessment dimensions are shown in the following table: For example, the "Mental Health Scale for Chinese Middle School Students (MHS-CISS)" (Professor Zheng Richang, 2002) is inserted into the topic of interest.
[0046] |Subscale Name|Core Assessment Content|Example of Typical Items| | Anxiety around people | Nervousness, low self-esteem, and fear in interpersonal relationships | I feel uncomfortable around strangers | |Tendency to be lonely|Loneliness, isolation, and negativity towards relationships|I feel like I have no real friends| | Self-blame tendency | Excessive self-denial and guilt for mistakes or shortcomings | I often blame myself for things I do wrong | |Study stress|Anxiety and stress about academic performance and test rankings|I get nervous when I think about exams| | Hypersensitivity | Oversensitivity to evaluations and circumstances, and mood swings | When someone laughs, I suspect they are laughing at me | |Physical symptoms|Somatization caused by psychological stress|I often feel dizzy or have headaches| |Paranoid tendencies|Suspicion, stubbornness, hostility|I feel like people are talking bad about me behind my back| |Maladjustment|Difficulty adjusting to school life and changes in the environment|I have difficulty adjusting to the new learning environment| |Emotional instability|Frequent mood swings, irritability, or depression|My moods often fluctuate greatly| |Impulsive tendencies|I have trouble controlling impulsive behavior|I sometimes can't help but throw things or hit people| 4.1.1. Example In order to better understand the invention, the following examples are given: (1) For example, a student named A in Beijing has interests in geology (especially caves), complex machinery, and rare creatures, and prefers light rock music with a cheerful rhythm. Based on student A's interests, the system intelligently integrates the above psychological test questions into chat topics related to his preferences. The physical educational robot automatically selects assessment scenarios based on the priority of interests and hobbies, and the scenarios can be sorted based on parameters such as landscape level or number of visitors.
[0047] (2) The system first plays a video of Shihua Cave in Beijing. The video can be viewed through a variety of devices such as display screens, projectors, and VR glasses, and is accompanied by background music that Student A likes. The robot will perform emotional analysis and adjust the content of the conversation: Robot: "Student A, does this cave look familiar to you?" Student A: "Yes, it's beautiful. I've been there." Robot: "Shihua Cave is very famous. Was it crowded when you went there?" Student A: "Yes, a lot of people!" Robot: "Do you feel uncomfortable meeting strangers when you visit Shihua Cave?" (corresponding to the psychological scale item "I feel uncomfortable around strangers") Student A: "No discomfort at all." Robot: "Do you feel uncomfortable when you meet unfamiliar teachers and classmates now?" Student A: "No discomfort at all." Robot: "Would you go visit caves with your good friends? For example, who are your good friends? Are they your true friends?" (corresponding to the psychological scale item "I feel I don't have any real friends") Student A: "I will go with my good friends. My good friends are Xiao Huang, Xiao Zhu, Xiao Tian, Xiao Song, and Xiao Qiu. They are all my real good friends." The robot simultaneously integrates cave knowledge, exploration equipment information, and historical cave stories, naturally integrating psychological test questions into the conversation scene: Robot: "You like to go to unexplored, pristine caves. If you don't bring a flashlight, do you blame yourself for doing something wrong? Or if you accidentally break your favorite machine model, do you often blame yourself?" Student A (positive response): "No, I can just go back and get my flashlight if I forget it. If I break the model, I can just ask Dad to fix it." The following is an example of a negative response: Student A (negative response): "Yes!" Robot: "Occasionally, sometimes, or often?" (Rating scale: 1 for occasionally, 2 for sometimes, 3 for often) Student A: "Sometimes I blame myself." (3) Leveraging the curiosity engine In the above case, through sentiment analysis, it was found that student A was not interested in Shihua Cave because he had been there many times. At this time, he could look for a cave with a better level and a higher beauty rating than Shihua Cave, and refer to the top ten caves in the world selected by some professional tourism websites, such as Huanglong Cave in Hunan.
[0048] The robot began to play videos and music of Huanglong Cave on the screen. Student A had indeed never been to Huanglong Cave, and he would be attracted by the new landscape. By analyzing student A's emotions through eye movements, etc., it was determined that A was attracted by the new cave. Then the robot conducted embedded questions and answers with reference to the scale.
[0049] Robot: "Student A, you have never been to Huanglong Cave in Hunan, right?", and the video and music (or commentary) of Huanglong Cave begin to play on the screen.
[0050] Student A: "Yes, it's beautiful." Robot: "Look, there are a lot of tourists!" Student A: "Yes, a lot of people!" Robot: "If you go to this cave, would you feel uncomfortable meeting strangers?" (corresponding to the psychological scale item "I feel uncomfortable around strangers") Student A: "No discomfort at all." Robot: "Do you feel uncomfortable when you meet unfamiliar teachers and classmates now?" Student A: "No discomfort at all." Note: The above conversations are only examples. Due to space constraints, we will not list them one by one. In actual application, more scenarios and questions can be expanded according to assessment needs. During the process, we will identify students' emotions and adjust the content of the conversation.
[0051] The technical solution includes: 4.2. Intelligent scene construction system 4.2.1 Multi-source data acquisition module Using Natural Language Processing (NLP) to analyze anonymous mental health records and daily conversations Interest weight calculation model: W_i = α•log(N_i) + β•TF-IDF(s_i) Among them, α and β are adjustment coefficients obtained through training of 100,000 consulting cases, N_i is the frequency of occurrence of interest points, and s_i is the semantic association strength. 4.2.2 Interest Graph Generation Technology Use Bidirectional Gated Graph Convolutional Network (Bi-GGCN) and Graph Attention Network (GAT) to process multi-source heterogeneous data: Input feature vector X=[x_1,...,x_n]^T∈R^(n×d), d=256-dimensional feature Graph convolution operation formula: H^(l+1)=σ(∑_(k=1)^K D_k^(-1 / 2) A_k D_k^(-1 / 2) H^(l)W_k^(l)) Dynamic interest modeling technology: Bimodal Data Fusion Mechanism: Improved Siamese Network Architecture Sequential dialogue features: Gated recurrent unit (GRU) processing h_t = GRU(x_t, h_{t-1}) Structural features of archives: Graph attention network constructs interest association graph Cross-modal contrast loss function: L_cmc = -log[exp(s_ij) / Σ_k exp(s_ik)] Time-aware weight decay factor: W_i(t) = W_i^0 * e^(-λ(t-t_0)) 4.2.3 Scene Matching Engine Generative Adversarial Scene Synthesizer (1) Multimodal Conditional GAN (MC-GAN) Generator improved design: Using the U-Net++ architecture (improved U-Net), the input layer contains the concatenation of interest tag vectors (Interest Tags), scale dimension encoding (Scale Dimensions) and scene parameters (Scene Parameters): I = Concat(Tag_emb, Dim_emb, Param_emb) ∈ R^(128×128×256) The generator outputs a virtual scene image `S∈R^(128×128×3)` and focuses on the area of interest through the spatial attention mechanism. The attention weight is calculated as: `A = Softmax(Conv3x3(ReLU(Conv1x1(I))))` Discriminator innovations: A multi-scale feature pyramid structure is introduced to integrate local details with global semantics, and the loss function adds semantic consistency constraints: `L_total = L_GAN + λ⋅L_consistency` Where `L_consistency = ||E(S_fake) - E(S_real)||_2`, `E(⋅)` is the pre-trained scene semantic encoder.
[0052] (2) Dynamic adversarial training strategy: Adopting a curriculum learning mechanism to gradually increase the complexity of the scenario: Stage 1: Generate a single point of interest scene (such as a cave structure) Phase 2: Generate multi-interest cross-scenes (such as cave + mechanical model + music elements) The difficulty of the discriminator increases exponentially with the number of training rounds: `D_k = D_base + ⌊log2(k)⌋⋅ΔD`, `k` is the number of training rounds.
[0053] Psychological scale adaptation model (1) Cross-modal feature alignment technology: Use Deep Contrastive Learning to construct a mapping relationship between scale items and scene elements: The scale text is encoded into `h_text∈R^768` by the psychology pre-trained model PsyBERT, and the scene features are extracted into `h_scene∈R^512` by ResNet-18. The matching degree is calculated by bilinear fusion: `Score = σ(h_text^TW h_scene + b)` Where `W∈R^(768×512)` is the learnable parameter matrix and `σ` is the Sigmoid function.
[0054] Innovatively introduce adversarial perturbation robustness training: add limited amplitude noise `δ` (`||δ||_2≤ 0.1`) to the input scene to force the model to learn invariant features.
[0055] (2) Real-time optimization mechanism: A lightweight convolutional network (MobileNetV3-Small) is deployed for online feature extraction, with an inference time of <15ms (measured on NVIDIA Jetson TX2) and support for dynamic updates of scale adaptation rules.
[0056] 4.2.4 Dynamic Scene Generator (1) Dialogue path planning based on reinforcement learning: the optimal path P = argmax(Σγ^t R(s_t,a_t)).
[0057] (2) The reward function R comprehensively considers: question concealment, answer credibility, and dimension coverage.
[0058] (3) Gradual training of course learning strategies: primary stage: single point of interest scene generation (such as cave structure); advanced stage: multi-point of interest cross-scene generation (such as cave + mechanical model).
[0059] 4.3. Multimodal Sentiment Analysis System 4.3.1 Visual Emotion Perception Module Cascaded expression recognition model: Improved MobileNetV3 detects 68 facial key points in real time; Temporal Convolutional Network (TCN) processes sequential data: y_t=ReLU(W*X_(tk)^(t+k)+b); combined with the OpenFace feature library, it outputs a 7-dimensional emotion probability distribution; 3D Separable Convolutional Network (CNN) for facial micro-expression processing: FLOPs = 0.38 × baseline model (tested on the FER-2013 dataset) 4.3.2 Language Sentiment Analysis Module PsyBERT, a psychological domain adaptive pre-training model: Joint training of MLM (masked language modeling) and TSD (topic-specific denoising) based on RoBERTa Emotion intensity calculation: E_s=Σ_(i=1)^n α_i•h_i^T•w_e Deep Residual Shrinkage Network (DRSN) processes speech features: MFCC' = DRSN(MFCC_raw) ⊕ (MFCC_raw ⊙ S(mask)) 4.3.3 Multimodal Fusion Decision Gated fusion mechanism: F = σ(W_f•[V;T])⊙V + (1-σ(W_f•[V;T]))⊙T Cross-modal consistency constraint: L_align = ||E_v(v) - E_a(a)||_2 + ||E_v(v) - E_t(t)||_2 4.4. Adaptive Dialogue System 4.4.1 Curiosity-driven engine Knowledge-enhanced Q-Learning framework: The state space S includes: interest matching, emotion fluctuation value, and historical question-answering path Reward function: R=0.4•(1-|E_t-E_(t-1)|)+0.3•I_s+0.3•K_r Monte Carlo Tree Search (MCTS) optimizes the dialogue path: UCB formula: UCB=Q(s,a)+c√(ln N(s) / n(s,a)) Cognitive load balancing mechanism: C_t=β_1•N_ent + β_2•T_avg + β_3•D_KL(p||q) 4.4.2 Hidden Interaction Module Hybrid Generative Adversarial Network (Hybrid GAN) generation problem: The generator is based on the Transformer-XL architecture; the discriminator uses the RoBERTa-wwm model to detect natural fluency (≥0.82) and hiddenness (≥0.75); Two-channel attention gating: h_fusion = g ⊙ h_topic + (1-g) ⊙ h_test.
[0060] 4.5. Real-time Optimization Engine 4.5.1 Online Incremental Learning System Improved Elastic Weight Consolidation (EWC) algorithm: Dynamic parameter importance matrix: F_ij = (1-η)F_ij + η(∂L / ∂θ_i)(∂L / ∂θ_j), single sample update time < 8ms (measured on NVIDIA Jetson TX2) 4.5.2 Adversarial Training Enhancement Constrained Adversarial Attack (CAA) generator: limited to synonym replacement with PMI>6.5; semantic perturbation amplitude δ<0.15 (BERT-base calculation).
[0061] 4.6. Key technical parameters 4.6.1 Model Training Details Dataset: 100,000 consulting cases, 8:1:1 split; Optimizer: AdamW (learning rate 3e-5, weight decay 0.01); batch size 32, early stopping strategy (patience=5) 4.6.2 Real-time performance indicators Response latency <800ms (NVIDIA Jetson AGX); memory usage <1.2GB (after quantization); conversation coherence BERTScore >0.87.
[0062] 4.6.3 Comparison of technical effects |Innovation Dimensions| Traditional Methods| Innovations of This Invention| Technological Breakthroughs| |Interest Modeling|Manual Annotation + Keyword Matching|Bimodal Dynamic Graph + Time Decay Model|F1 score increased by 32.7% | |Problem concealment|Fixed template insertion|Hybrid GAN generation + attention gating|Concealment 82.5% | |Multimodal fusion|Weighted summation|Asymmetric features + functions| AUC 0.927 (95%CI 0.914-0.939)| | Real-time optimization | Offline batch update | Online EWC + adversarial training | Performance degradation <2.3% | 4.6.4 Pilot Test Results | Indicator | Traditional Method | This Invention | Improvement | Assessment Acceptance Rate | 63.2% | 91.7% | +45.1% | Question Validity | 72.4 points | 88.6 points | +22.4% | | Anomaly Detection Timeliness | 14.5 days | 3.2 days | +353% | 4.7. Special Notes All parameters have undergone ethical review by psychology experts and comply with the confidentiality requirements of Article 23 of the Mental Health Law.
[0063] Industry firsts: Reinforcement learning for dynamic scenario generation; Multimodal hidden evaluation; Cross-modal alignment of conversation data and scales.
[0064] 5. Dynamic Psychological Assessment Scale 5.1. Standardized scale Construct a hierarchical scale system: the total scale consists of 10 subscales.
[0065] 5.1.1 Intelligent Scale Generation System Workflow of the literature mining module: search for keywords and create feature scales by citing research reports.
[0066] 5.1.2 Sample Scale For example, insert the Mental Health Scale for Chinese Middle School Students (MHS-CISS) (Professor Zheng Richang, 2002) into the interest scene. |Subscale Name|Core Assessment Content|Weight| |Anxiety towards people|Tension, inferiority, and fear in interpersonal communication|5% | |Loneliness Tendency|Loneliness, isolation, and negativity toward relationships|5% Self-blame tendency | Excessive self-denial and guilt over mistakes or shortcomings | 10% | |Study pressure|Anxiety and pressure about academic tasks|Exam rankings|10% Hypersensitivity | Excessive sensitivity to evaluation and environment, and mood swings | 10% | |Physical symptoms|Somatization caused by psychological stress|10% | Paranoid tendencies | Suspicion, stubbornness, hostility | 10% | |Maladjustment|Difficulty adapting to school life and environmental changes|5% |Emotional instability|Frequent mood swings, irritability, or depression|15% | |Impulsive tendencies|Difficulty controlling impulsive behavior|20% (Note: Weight parameters are continuously updated and optimized.) The scores of each subscale are added up to obtain the total score, and the students' mental health is evaluated based on the sub-scores and the total score.
[0067] 5.2 Intelligent Scale Matching System 5.2.1 Multimodal Scale Parsing Engine Using Deep Semantic Parsing Network (DSPN) to automatically parse standardized scales: (1) Construct a Chinese psychological scale parser based on RoBERTa-wwm and extract item features through a bidirectional attention mechanism; (2) Design scale dimension association map: Use graph neural network (GNN) to establish the topological relationship between each component scale. The formula is expressed as: G = (V,E), V∈R^d, E_ij = σ(W_g•[h_i;h_j]) Where d = 768-dimensional BERT embedding vector, σ is the Sigmoid function 5.2.2 Dynamic Dimension Weight Allocation Develop a weight optimization model based on Deep Reinforcement Learning (DRL): State space: S = [emotional fluctuation value, response credibility, scene matching degree] Action space: A = {weight adjustment coefficient Δw_i | i∈component scale dimension} Reward function: R = 0.6•Corr(Δw, E_m) + 0.4•(1 - KL(p||q)) Among them, E_m is the result of multimodal sentiment analysis, and KL divergence ensures the rationality of weight distribution 5.3 Hidden Question Generation Based on Scale 5.3.1 Hybrid Generative Adversarial Network (Hybrid GAN) (1) Generator architecture: GPT-2 → Transformer-XL → Semantic Constraint Layer ↘ Interest Feature Embedding → Attention Fusion (2) The discriminator adopts a dual-channel detection mechanism: Naturalness judgment: RoBERTa-wwm calculation fluency score S_f ≥ 0.82 Concealment judgment: CNN detection of psychological scale feature similarity S_c ≤0.15 5.3.2 Context-aware embedding algorithm Design a dynamic attention gating mechanism: h_fusion = σ(W_g•[h_topic; h_test]) ⊙ h_topic + (1-σ(W_g•[h_topic;h_test])) ⊙ h_test Where h_topic∈R^512 is the topic feature vector, h_test∈R^512 is the test question feature vector 5.4 Real-time Adaptive Optimization 5.4.1 Federated Incremental Learning Framework Establish a three-level privacy protection mechanism: (1) Local Differential Privacy (LDP): Add Laplace noise, ε=0.3; (2) Secure Multi-Party Computation (SMC): Paillier homomorphic encryption; (3) Model parameter aggregation: Improved FedAvg algorithm, dynamically adjusting the learning rate η_t=η_0•e^(-λt).
[0068] 5.4.2 Bayesian Dynamic Parameter Adjustment Construct a Gaussian Process Regression (GPR) model: f(w) ~ GP(m(w), k(w,w')), the kernel function uses Matern 5 / 2: k(w,w') = σ_f^2(1 + √5r + 5r^2 / 3)exp(-√5r), r=||w-w'|| / l Dynamic update cycle Δt=30 minutes 5.5 Multi-dimensional Verification System 5.5.1 Adversarial Verification Module (1) Design of psychological scale feature confusion detector: Use t-SNE dimensionality reduction to analyze the question distribution and ensure that the KL divergence D_KL between hidden questions and normal topics is less than 0.15 (2) Detection formula: D_KL(P||Q) = ΣP(x)log(P(x) / Q(x)).
[0069] 5.5.2 Ethical compliance monitoring Developing an ethical review system based on a rule engine: Keyword filtering library: contains 2000+ sensitive words (such as self-harm, suicide, etc.); Real-time emotion warning: triggers manual intervention when E_m > threshold θ_e is detected; Threshold dynamic calculation: θ_e = μ_e + 3σ_e (based on historical sentiment data statistics); 5.6 Comparison of technical effects | Innovation Dimensions | Traditional Methods | Inventive Solution | AI Technology Advantages | | Scale Matching | Manual Experience Matching | DSPN+GNN Dynamic Parsing | Matching Speed Increased by 18 Times | | Question Generation | Template Replacement | Hybrid GAN Generation | 41.7% Concealment Improvement | | Weight optimization | Fixed weight | DRL dynamic adjustment | Anomaly detection rate increased by 29.3% | |Privacy Protection|Data Desensitization|LDP+SMC Federated Learning|Privacy Leakage Risk Reduced by 76.2%| The invention of this module achieves the following through the above artificial intelligence technology innovation: (1) The first dynamic scale system that combines deep semantic parsing and reinforcement learning; (2) The first to optimize psychological scale parameters under the federated learning framework; (3) A breakthrough in using hybrid generative adversarial networks to achieve covert implantation of scale questions; (4) Establish a multi-level ethical protection mechanism in accordance with the requirements of the “Artificial Intelligence Ethics Review Standards”.
[0070] 6. Other innovations 6.1. Intelligent Evaluation Architecture for Multimodal Fusion It is the first to deeply integrate students' anonymous physiological data (blood flow, body temperature, bioelectricity, odor), behavioral data (handshake strength, eye movement trajectory), and language data (voice emotion, conversation content). Through the LSTM+CNN hybrid model and the improved LightGBM algorithm, a high-precision mental health assessment system is constructed with a false alarm rate as low as 3.2%, breaking through the limitations of traditional single scale assessment.
[0071] A two-stream network architecture based on the attention mechanism is adopted. The physiological signal stream processes the five-in-one sensor time series data (sampling frequency 100Hz, input dimension 5×300 time steps) through the gated recurrent unit (GRU), and the behavioral feature stream uses 3DResNet to process eye tracking video (resolution 640×480@30fps). Features are fused through the cross-attention module, and the weight distribution formula is α = σ(W_a[h_phy||h_beh]) (σ is the Sigmoid function, W_a is the learnable parameter matrix).
[0072] 6.2. Dynamic Adaptive Human-Computer Interaction Strategy Based on the BERT model's real-time emotion prediction and reinforcement learning (PPO algorithm), the conversation strategy is dynamically optimized to achieve a closed-loop interaction of "emotion perception-topic guidance-test embedding". The conversation fluency response delay is ≤300ms, and the topic switching threshold (interest matching degree >70%, emotion fluctuation >20%) is intelligently adjusted, significantly improving student participation and data acquisition effectiveness.
[0073] A hierarchical reinforcement learning framework was designed. The high-level strategy selected the conversation mode (topic extension / test insertion / emotional support) based on the PPO algorithm. The state space S = (emotion intensity ∈ [0, 1], topic matching ∈ [0, 1], number of conversation turns ∈ N), and the action space A = {continue the current topic, insert PHQ-2 test questions, switch to soothing topics}. The low-level strategy dynamically embedded test questions through the Pointer-Generator Network. For example, when a student answered "I often suffer from insomnia recently," the strategy generated "That sounds a bit troublesome (emotional feedback). Can you tell me what time you usually go to bed? (test question embedding)"?
[0074] 6.6. End-to-end privacy and security protection solution It integrates edge computing local processing, federated learning model training, differential privacy data desensitization and other technologies to achieve full-link encryption of "data collection-transmission-storage-processing", with a sensitive data retention period of ≤72 hours. It strictly complies with the "Personal Information Protection Law" and ISO / IEC 27701 privacy protection certification requirements, and solves the privacy leakage risks in traditional psychological assessments.
[0075] An improved locality-sensitive hashing (LSH) algorithm is used to de-identify facial features, and the 128-dimensional feature vector is encoded into a 256-bit binary hash code. The Hamming distance threshold is set ≥15 to ensure k-anonymity (k≥3). Differential privacy noise injection (ε=0.5, δ=1e-5) and secure multi-party computation (SMPC) are used in the federated learning framework. Feature-level desensitization only retains the relative change rate (±%) of biometric features.
[0076] 6.7. Accurate Active Retest Trigger Mechanism An improved EWMA control chart algorithm is used to dynamically monitor the risk index. Combined with three-level confidence assessment and multi-period baseline comparison, accurate capture of abnormal psychological states can be achieved. High-risk situations automatically trigger the intervention of real experts, building a three-level protection system of "machine initial screening-intelligent retesting-manual intervention".
[0077] Multimodal anomaly detection is verified through semantic-physiological conflict. For example, if the voice answer "emotionally calm" but the skin conductance level (SCL) rise rate is >0.05μS / s (baseline 0.02μS / s) or the pupil diameter fluctuation rate is >15%, it will be marked as abnormal; eye movement analysis uses STFT to analyze the micro-saccade spectrum (3-5Hz component indicates dispersion), the HMM model is used to construct the gaze pattern transition probability (maintain state probability 0.8, transition probability 0.2), and the DTW algorithm is used to match abnormal trajectories (similarity with the negative depressive psychological manifestation pattern library >65%).
[0078] 6.8. Sensor-Algorithm Joint Optimization and Real-Time Guarantee A five-in-one sensor array was developed on a 15×15mm flexible circuit board. The pressure sensor uses a graphene nanostructure (sensitivity 0.1N), the bioelectric module has a common-mode rejection ratio of >100dB, and the clock deviation between hardware sampling and algorithm processing is <1ms. A double-buffered pipeline architecture was constructed, sensor data was preprocessed on the FPGA (FIR filtering + downsampling), and a GPU-accelerated hybrid inference engine (TensorRT deployment model) was used. When the delay was >300ms, the eye movement resolution was automatically reduced (640×480→320×240).
[0079] 6.6. Ethical safety mechanism Three-level data firewall: including hardware-layer trusted execution environment (TEE) to isolate sensitive data, algorithm-layer federated learning + homomorphic encryption transmission gradient, and application-layer blockchain access log audit (implemented by Hyperledger Fabric).
[0080] 7. System performance and technical advantages 7.1. Technical parameters and performance indicators | Technical Module | Core Parameters | Performance Index | |Emotion Recognition|LSTM+CNN Hybrid Model|Emotion classification accuracy ≥ 85%, confidence score error ≤ 5%| |Conversation Response Speed | Edge Computing Platform | Response Latency ≤ 300ms (95% of the time) | |Outlier Detection| Improved LightGBM Algorithm|False Positive Rate ≤ 3.2% | |Privacy protection level|Differential privacy + federated learning|ε=0.5, k-anonymity≥3| |Attention Test|Near-Infrared Eye Tracker| Fixation accuracy 0.5°, pupil diameter measurement accuracy ±0.1mm| | Assessment Reliability and Validity | Cronbach's α | Consistency α ≥ 0.8, test-retest reliability ICC ≥ 0.75 (2-week interval) | 7.2. Comparison of the Advantages of Quantitative Technology | Technical Indicators | Traditional Solutions | This Invention | Improvement | Emotion Recognition Accuracy | 68.2% (Unimodal) | 92.3% (Multimodal) | +35.3% | Test Acceptance Rate | 47.5% (Direct Question) | 89.7% (Natural Embedding) | +88.8% | False positive rate: 12.8% (screening) | 3.2% (multimodal validation) | -75% Data processing latency | 850ms (cloud processing) | 210ms (edge computing) | -75.3% | Privacy Protection Strength | AES-128 Encryption | Federated Learning + Homomorphic Encryption | Security Improved by 2 Levels | BRIEF DESCRIPTION OF THE DRAWINGS
[0081] Figure 1 System overall architecture diagram The system includes three modules: multimodal acquisition, intelligent processing, and security assurance: 1.1. Front-end acquisition: The educational robot integrates physiological (blood flow, body temperature), behavioral (eye movement, handshake strength), and language (speech, text) sensors, and transmits the data to the AI model after edge computing pre-processing.
[0082] 1.2. Intelligent Processing: The dynamic scene generation unit embeds psychological test questions into conversations based on the interest graph (Bi-GGCN) and generative adversarial network (MC-GAN). The AI assessment model fuses multimodal data through LSTM+CNN to output a three-level risk index.
[0083] 1.3. Security: Biometrics are anonymized through chaotic transformation, data transmission uses quantum encryption (SM4+NTRU+BB84), and storage combines blockchain evidence with physical isolation gateways to ensure privacy and security.
[0084] Figure 2 Five-level privacy protection technology roadmap 2.1. Biometric anonymization: The facial feature code is transformed into an irreversible code using chaotic random numbers, and a double hash chain (SM3 algorithm) is used to associate the code ID. The original image is deleted within 10 seconds.
[0085] 2.2. Quantum encryption: SM4 block encryption at the bottom layer, NTRU quantum-resistant key encapsulation at the middle layer, and BB84 protocol key distribution at the top layer (update ≤ 30 seconds).
[0086] 2.3. Blockchain Evidence Storage: Consortium chain architecture (education / school / health nodes), improved PBFT consensus (fault tolerance ≥ 33%), and data sharding storage (Reed-Solomon encoding).
[0087] 2.4. Physical Isolation: The gateway implements one-way optical transmission (850nm laser, rate ≥10Gbps). In the event of an abnormality, the optical path is cut off within 0.5 seconds and the key is destroyed.
[0088] 2.5. Permission Control: Dynamic desensitization (teachers see the level code, experts need quantum decryption). The school can only obtain the results after parental authorization.
[0089] Figure 3 Dynamic scene generation technology flow chart 3.1. Interest Modeling: Bi-GGCN processes archives and conversations to generate interest tags; the twin network integrates temporal and structured features to dynamically adjust interest weights.
[0090] 3.2.MC-GAN Architecture: The generator (U-Net++) integrates interest tags and scale dimensions to generate scene images, while the discriminator (feature pyramid) verifies authenticity and concealment. The course generates single / multiple interest scenes in stages, gradually increasing in difficulty.
[0091] 3.3. Scale Adaptation: PsyBERT encodes the scale text, ResNet extracts scene features, bilinear fusion calculates the matching degree, and anti-disturbance training enhances robustness.
[0092] Figure 4 Schematic diagram of multimodal data fusion model 4.1. Physiological Data: LSTM processes time series signals such as blood flow and bioelectricity to capture physiological fluctuations of emotions such as anxiety.
[0093] 4.2. Behavioral Data: 3D ResNet analyzes eye movement videos (gaze point, pupil changes), and CNN identifies handshake strength and tremor frequency to determine the degree of limb tension.
[0094] 4.3. Language Data: DRSN processes speech MFCC features, and PsyBERT analyzes text sentiment. The gating mechanism integrates multimodal features and outputs a risk index (a threshold of 0.85 triggers retesting).
[0095] Figure 5 Dialogue strategy optimization framework diagram 5.1. Hierarchical Reinforcement Learning: The high-level strategy (PPO algorithm) selects the conversation mode (continue / insert test / switch topic) based on sentiment intensity and topic matching, while the low-level strategy generates natural language questions.
[0096] 5.2. Curiosity-driven: Q-Learning combines knowledge gain to optimize exploration strategies, while MCTS balances exploration and exploitation. The real-time adjustment mechanism simplifies problems based on cognitive load and dynamically switches the language when emotions fluctuate. DETAILED DESCRIPTION
[0097] Example 1: Psychological Assessment Scenario in a Mountain Middle School 1.1: Hardware deployment and initialization 1.1.1. Educational robot configuration: Hardware platform: NVIDIA Jetson AGX Xavier edge computing module (32GB RAM, 512 CUDA cores) Sensor Array: Head: FLIR Boson 640 thermal imaging camera (resolution 640×480@30Hz) Hand: FlexiForce A201 pressure sensor (range 0-100N, accuracy ±0.1N) Bioelectric module: ADS1299 24-bit ADC chip (16-channel differential input, CMRR>110dB) Security chip: National secret level 2 encryption module (SM4 / SM3 hardware acceleration) 1.1.2. Network architecture deployment: Python # Federated Learning Node Configuration Example nodes = {"school_01": {"type": "client", "ip": "192.168.1.10", "role": "data_processor"}, "edu_center": {"type": "server", "ip": "10.10.1.1", "role": "model_aggregator"}} 1.2: Biometric anonymization (technical details) 1.2.1. Facial feature code generation process: Input: RGB image (256×256 pixels) Feature extraction: Use the ArcFace model (LResNet100 backbone) to generate a 512-dimensional feature vector Chaotic transformation: Generate irreversible code by iterating the Lorenz system 20 times Lorenz parameter settings: dx / dt = 10(y - x), dy / dt = x(28 - z) - y, dz / dt = xy - (8 / 3)z, Initial values: x0=0.01, y0=0, z0=0 (random perturbation ±0.0001 injected by hardware entropy source) Output: 128-bit binary signature (example: 0x3A5F...E9C2) 1.2.2. Double hash chain association: First-layer hash calculation (SM3 algorithm): H1 = SM3_Hash(Concat(face feature code, Unix timestamp)); Second-layer hash calculation: Salt value = randomly generated 32 bytes (generated by the TPM chip); H2 = SM3_Hash(Concat(student ID, school organization code, salt value)); Mapping table storage structure: Use B+ tree index to speed up queries.
[0098] 1.3: Dynamic scene generation and test question embedding 1.3.1. Interest Graph Construction: Data source: Historical conversation text (entity extraction by BERT) + sensor behavior data Graph update formula: W_i(t) = 0.7*log(number of occurrences of interest points) + 0.3*TF-IDF(semantic association strength) - 0.05*(current time - last occurrence time) / 3600 Output: Dynamic weight sorting (Example: ["Geological Exploration: 0.82", "Mechanical Model: 0.75", "Electronic Music: 0.68"]) 1.3.2. MC-GAN scene generation: Python # Generator Architecture (U-Net++ Improved) class Generator(nn.Module): def __init__(self): super().__init__() self.down1 = ConvBlock(256, 128) # Input channel: interest label (128) + scale dimension (64) + parameter (64) self.up1 = UpConvBlock(256, 64) self.attn = SpatialAttention(scale=0.5) # Spatial attention weight def forward(self, x): x = self.down1(x) x = self.attn(x) * x # Focus on strengthening the area of interest return self.up1(x) # Discriminator loss function loss = adversarial_loss + 0.3*|E(fake_scene) - E(real_scene)|_2 1.3.3. Test question implantation trigger conditions: Real-time monitoring indicators: IF emotional fluctuation value > 0.2 (baseline 0.1) AND interest matching > 70% AND conversation round number > 3 THEN trigger test question implantation Examples of implant strategies: Original dialogue: "It took tens of thousands of years for the stalactites in this cave to form." After implantation: "If you get lost here, will you blame yourself for not being prepared enough?" (corresponding to the Self-blame Tendency Scale) 1.4: Multimodal Data Fusion Analysis 1.4.1. Physiological-Behavioral Data Synchronization: Time alignment mechanism: C / / Hardware-level timestamp synchronization (accuracy ±1ms) void sync_timestamps() { IMU_data.timestamp = get_network_time(NTP_SERVER); bio_sensors.timestamp = IMU_data.timestamp - clock_skew;} 1.4.2. Hybrid Model Inference: Python # LSTM+CNN hybrid architecture class FusionModel(nn.Module): def __init__(self): self.lstm = nn.LSTM(input_size=5, hidden_size=128) # Input: 5-dimensional physiological signal self.cnn = ResNet3D(in_channels=3) # Input: eye movement video stream def forward(self, x_phy, x_video): h_phy, _ = self.lstm(x_phy) # Output shape: (batch, 128) h_video = self.cnn(x_video) # Output shape: (batch, 256) fused = torch.cat([h_phy, h_video], dim=1) return self.classifier(fused) # Output: Anxiety / Depression / Normal 1.5: Safety Control and Emergency Response 1.5.1. Quantum Key Distribution Process: Equipment: QKD transmitter (1550nm laser, 10MHz repetition rate) Protocol execution steps: (1) Base station A sends a sequence of photons with polarization states |H>, |V>, |+>, |->; (2) Base station B randomly selects a measurement basis (Rectilinear / Diagonal); (3) Both parties select the basis through classical channel comparison and retain the matching results; (4). Finally, a 256-bit shared key is generated (valid when the bit error rate is <1%).
[0099] 1.5.2. Data self-destruction mechanism: Trigger conditions: WHEN Shell open signal == True OR Temperature > 85℃ OR Acceleration > 10g THEN Activate degaussing coil (magnetic field strength ≥ 5000 Oe) Hardware response time: from detecting the trigger signal to completing data erasure < 200ms Example 2: University Campus Psychological Intervention Scenario (Differentiated Implementation) 2.1. Hardware configuration adjustment: Additional modules: electronic nose sensor (MQ-3 alcohol sensor + GP2Y1010AU0F dust sensor) Modify encryption policy: use SM9 identification password instead of SM4, support group management 2.2. Dialogue strategy optimization: Reinforcement learning parameter adjustment: Reward function R = 0.5*(1 - |emotion change|) + 0.3*test completion rate + 0.2*conversation depth; PPO algorithm update frequency: The policy network is updated every 50 rounds of dialogue.
[0100] 2.3. Special handling procedures: When the alcohol concentration is detected to be >0.2mg / L: (1) Enable privacy protection mode: disable the camera and only retain voice interaction; (2) The conversation shifts to stress relief: "It sounds like you've been under a lot of pressure lately. Would you like to talk about what's been going on?" (3) Trigger emergency contact notification for one consecutive week (students must actively confirm).
[0101] Technical effect verification data | Test Item | Index Value | Measurement Conditions | |Irreversibility of face anonymization|Hamming distance ≥ 115 / 128 bits|1000 groups of the same person with different expressions| | Emotion Recognition Response Latency | 217±15ms | Jetson AGX Measured | Naturalness of test placement score | 4.8 / 5.0 | Blind test with 50 psychologists | Encrypted data resistance to quantum attacks | Requires >10^3 quantum gate operations | NIST PQC testing framework | | Mean Time Between Failures (MTBF) | 2,150 hours | 72-hour stress test | The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. An intelligent system for mental health assessment based on educational robots, characterized by include: (1) Multimodal biometric anonymization module uses a chaotic random number generator to perform nonlinear transformation on the original biometrics to generate a 128-bit irreversible feature code; (2) Quantum hybrid encryption unit, including the SM4 algorithm block encryption layer, the NTRU post-quantum cryptography middle layer and the BB84 protocol quantum key distribution top layer; (3) Dynamic scenario generator, using multimodal conditional generative adversarial network (MC-GAN) to naturally integrate mental health test questions into interest scenarios; (3) Edge computing security module, integrating the national secret level 2 chip and deploying the federated learning framework, the client uses differential privacy noise injection (ε=0.5, δ=1e-5) for local model training.
2. The system according to claim 1, characterized in that The biometric anonymization module includes: (1) a chaotic random number generator based on the Lorenz attractor equation, with parameters σ=10, ρ=28, β=8 / 3; (2) a double hash chain association engine, with the first layer hash H1=SM3(feature code||timestamp) and the second layer hash H2=SM3(student ID||random salt value); (3) a multimodal compatible interface that supports the fusion verification of MFCC voiceprint code and iris feature code.
3. The system according to claim 1, characterized in that The dynamic scene generator includes: (1) a U-Net++ architecture generator of a generative adversarial network (GAN), whose input layer concatenates interest label vectors and scale dimension encoding; (2) a multi-scale feature pyramid discriminator, whose loss function adds a semantic consistency constraint L_consistency=||E(S_fake)-E(S_real)||2; and (3) a curriculum learning mechanism, which generates single interest point scenes and multi-interest intersection scenes in stages.
4. The system according to claim 1, characterized in that The emotion analysis module includes: (1) a three-dimensional separable convolutional network (3D Separable CNN) to process facial micro-expressions, reducing the computational complexity to 38% of the baseline model; (2) a psychological domain pre-training model PsyBERT, which is jointly trained through masked language modeling (MLM) and topic-specific denoising (TSD); and (3) a gated multimodal fusion mechanism F=σ(W_f•[V;T])⊙V+(1-σ(W_f•[T;V]))⊙T.
5. The system according to claim 1, characterized in that The privacy protection system includes: (1) a physical isolation gateway that uses an 850nm VCSEL laser array to achieve unidirectional transmission (rate ≥10Gbps, bit error rate ≤1×10⁻¹²); (2) a blockchain evidence storage module that deploys an improved PBFT consensus algorithm with a Byzantine fault tolerance threshold ≥33%; and (3) a data self-destruction mechanism that activates a piezoelectric ceramic crusher to destroy the storage unit when the shell is detected to be open.
6. A method for evaluating mental health, characterized in that The method includes the following steps: (1) Calculating the interest weight W_i=α•log(N_i)+β•TF-IDF(s_i) through interest graph generation technology; (2) Adopting reinforcement learning to optimize the dialogue strategy, with the reward function R=0.4•(1-|E_t-E_(t-1)|)+0.3•I_s+0.3•K_r; (3) Deploying Monte Carlo tree search (MCTS) to plan the dialogue path, with the UCB formula being UCB=Q(s,a)+c√(lnN(s) / n(s,a)); (4) Dynamically adjusting the timing of test question implantation, triggering embedding when the emotion fluctuation value is greater than the threshold θ_e and the interest matching degree is greater than 70%.
7. The method according to claim 6, characterized in that The dynamic risk assessment includes: (1) real-time analysis of psychological data streams using a hidden Markov model (HMM) with an abnormal probability threshold of θ = 0.85; (2) monitoring of risk index using an improved EWMA control chart algorithm (α = 0.2, UCL = 2.3); and (3) three-level confidence verification, with multi-period baseline comparison initiated when confidence is low (at least three independent measurements).
8. The method according to claim 6, characterized in that The physiological data acquisition includes: (1) laser Doppler blood flow meter to measure microcirculation data (sampling rate 100 Hz); (2) flexible strain pressure sensor to detect grip force distribution (accuracy ±0.1N); (3) 16-channel bioelectric acquisition module (common mode rejection ratio >100 dB).
9. The method according to claim 6, characterized in that The test question generation includes: (1) generating natural language questions using a hybrid generative adversarial network (Hybrid GAN), with a concealment discrimination threshold S_c ≤ 0.15; (2) a context-aware implantation algorithm, dynamic attention gating h_fusion = σ(W_g•[h_topic;h_test])⊙h_topic+...; and (3) real-time semantic alignment, calculating the matching degree between scale items and scenes through bilinear fusion Score = σ(h_text^TWh_scene).
10. The method according to claim 6, characterized in that The security control includes: (1) Reed-Solomon encoding (n≥5, k=3) for data shard storage; (2) Laplace noise (ε=0.3) is added to local differential privacy (LDP); and (3) Paillier homomorphic encryption is used to transmit gradient parameters in secure multi-party computation (SMC).
Citation Information
Cited By
All-in-One video restoration system, method and device based on expert system
CN121073836A
Road water damage disaster rapid identification method and system based on remote sensing image analysis
CN121415264A
Highway water damage disaster rapid identification method and system based on remote sensing image analysis
CN121415264B
Edge-cloud collaborative architecture-based emotion recognition privacy protection system and method
CN121435274A
Ear-nose-throat department symptom monitoring method and system based on multi-modal data fusion
CN121512476A