A method and device for monitoring the safety of autism children in multi-modal risk perception based on robot interaction tasks

CN122889367APending Publication Date: 2026-10-09FUJIAN UNIV OF TRADITIONAL CHINESE MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610995526.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-10-09

AI Technical Summary

Technical Problem

旨在解决孤独症儿童在康复机器人互动中虽然能够被机器人形象、声音、动作和游戏化任务吸引,但也可能因为声光刺激、任务失败、偏好物被限制、轮替等待、社交要求增加或空间边界变化而出现紧张、回避、尖叫、抓挠、撞击、冲向门口、过近接触机器人等安全风险,而现有孤独症康复机器人在真实互动过程中无法提前识别上述安全风险、无法将多模态指标与具体任务前因关联、以及无法依据风险预测自动调整机器人行为的问题

Benefits of technology

[0065]通过机器人特定交互任务与多模态时钟对齐采集的技术特征,实现了特征提取与具体康复训练任务前因的深度关联,提高了风险感知的主观可解释性与客观精度。 现有技术多依赖泛化的离线视频或脱离上下文的单一生理信号,难以解释诱发情绪波动的真实原因。本方案通过机器人主动执行呼名共同注意、感官舒适表达、轮替协作与模仿、延迟满足与求助以及安全边界与空间转移等特定结构化任务,以机器人自身的事件时间戳为主轴对视觉行为、语音表达、生理状态、空间运动、触觉或力反馈和任务日志等多模态数据进行滑动窗口切片与时钟对齐。这一手段使得采集到的每一组反应潜伏期、视线回避、心率变异性及皮肤电峰值等特征指标,均具备极其明确的交互上下文前因解释,为后续精准捕捉危险升级的先兆奠定了客观且可解释的数据基础。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122889367A_ABST
    Figure CN122889367A_ABST
Patent Text Reader

Abstract

The application relates to an autism child multi-modal risk perception safety monitoring method and device based on a robot interaction task, the method synchronously collects multi-modal original data such as vision, voice, physiology and spatial motion in response to a specific interaction task performed by a robot; clock alignment slicing is performed with the robot event timestamp as a main shaft, and time sequence characteristic indexes are extracted; the characteristic indexes are input into a task-guided time sequence multi-modal risk perception network, quality perception fusion is performed through a cross-modal attention mechanism to obtain a fusion state vector; based on the fusion state vector and an individualized causal risk chain diagram, multi-type risk probabilities and prediction uncertainties are time sequence predicted, and a comprehensive risk score is calculated; control decisions are made according to the comprehensive risk score and a child individualized dynamic risk threshold, and a robot is driven to adaptively perform multi-level safety intervention actions. The scheme effectively reduces the problem behavior escalation probability, and realizes the organic unification of rehabilitation training and safety monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of autism rehabilitation robots and human-computer interaction control technology, specifically relating to a safety monitoring method and device for multimodal data acquisition and temporal risk prediction based on robot interaction tasks during the rehabilitation care and training of children with special needs. Background Technology

[0002] Currently, autism screening and rehabilitation assessment typically employ the following technical approaches: First, the scale and expert observation approach, where therapists or doctors provide conclusions through scales such as ADOS, CARS, and M-CHAT, as well as on-site observation; this approach is highly subjective. Second, the structured paradigm approach, which elicits children's responses through tasks such as name calling, joint attention, imitation, and emotional fragmentation, followed by the collection and analysis of video, audio, or eye-tracking data. Third, the multimodal acquisition approach, which collects behavioral, physiological, linguistic, and social indicators through RGB-D cameras, microphone arrays, eye trackers, wearable physiological sensors, or VR environments. Fourth, the machine learning fusion approach, which calculates autism risk or ability scores using algorithms such as CNN, LSTM, BP neural networks, soft voting, autoencoders, and attention mechanisms. Fifth, the robot-assisted approach, which enhances children's participation and assists in emotional expression, imitation training, and sensory adaptation through humanoid robots or facial expression robots.

[0003] The aforementioned approaches have demonstrated the value of multimodal data acquisition, paradigmatic tasks, and robot interaction. However, most solutions remain at the level of "screening, assessment, reporting, and course recommendation," making it difficult to determine in real time whether "escalation of danger is imminent" during continuous interaction between rehabilitation robots and children. Especially in scenarios such as natural homes, institutional training rooms, and school-based inclusive activities, children's states fluctuate rapidly due to fatigue, environmental noise, task failure, and changes in preferences, making a single assessment result insufficient to represent their real-time safety status.

[0004] Specifically, existing autism screening and rehabilitation assessment systems primarily rely on scales, short-term video paradigms, VR scenarios, or offline multimodal data analysis, focusing on "whether there is a risk of autism" or "scores for a certain ability." However, they lack systematic technical solutions for addressing issues such as "when risk warning signs appear, whether the risk will escalate within the next 30-120 seconds, and how the robot should safely degrade and alert caregivers during real-world interactions with rehabilitation robots." In institutional training and home care scenarios, caregivers often can only intervene after problem behaviors have already occurred, affecting training continuity and potentially increasing the risk of injury to both children and caregivers. Existing technologies typically have the following drawbacks:

[0005] (1) Insufficient real-time performance: Existing systems mostly generate evaluation reports after the task is completed, and cannot provide early warnings when children show signs of risk such as self-harm, aggression, escape, collision, or collapse.

[0006] (2) Lack of robot closed loop in tasks: Existing paradigms are mostly presented by real human evaluators or video / VR stimulation, with robots only serving as tools to attract attention and not undertaking the main body of task arrangement, risk adjustment and safe execution.

[0007] (3) Incomplete modality: Video or eye movement alone is insufficient to capture physiological risks such as increased stress, heart rate variability, skin conductance changes, and rapid breathing; and physiological alone cannot explain the antecedents and specific behaviors of the task.

[0008] (4) Insufficient data on natural states: The short-term clinic paradigm is difficult to cover real-world risk scenarios such as family care, free play, waiting, refusal, change of location, and noise stimulation, and the model's generalization ability is limited.

[0009] (5) Inadequate handling of individual differences: The risk warning signs of autism vary greatly among children. Some show staring avoidance first, some show repeated clapping first, and some show rapid breathing or rushing to the door first. Using a uniform threshold can easily lead to false alarms or missed alarms.

[0010] (6) Lack of robot safety strategy: Existing solutions rarely convert risk scores directly into executable control actions such as robot deceleration, reversal, stopping, reducing sound and light stimulation, switching tasks, or reminding caregivers.

[0011] In summary, overcoming the shortcomings of existing technologies and providing a multimodal risk perception safety solution for autism rehabilitation robots, enabling those skilled in the art to obtain interpretable and reproducible behavioral indicators through specific robot interaction tasks, continuously collect large amounts of long-term data on children in a natural state, establish individualized baselines and risk warning models, and thus complete early warning, task downgrading, robot backoff, stimulus reduction, help prompts, and caregiver notifications before the risk escalates, has become a pressing technical problem to be solved by those skilled in the art. Summary of the Invention

[0012] The purpose of this invention is to provide a method and device for multimodal risk perception and safety monitoring of children with autism based on robot interaction tasks. This invention is primarily aimed at children with autism spectrum disorder (ASD) in rehabilitation training, institutional care, family intervention, and inclusive education settings, and is particularly suitable for children with risks of social communication difficulties, sensory sensitivity, emotional regulation difficulties, stereotyped repetitive behaviors, and self-harm, aggression, or escape behaviors. It aims to address the safety risks that autistic children may experience during rehabilitation robot interactions. While they may be attracted by the robot's image, sound, movements, and gamified tasks, they may also exhibit tension, avoidance, screaming, scratching, bumping, rushing towards the door, or excessively close contact with the robot due to auditory and visual stimuli, task failure, restrictions on preferred objects, waiting, increased social demands, or changes in spatial boundaries. Existing autism rehabilitation robots cannot identify these safety risks in advance during actual interaction, cannot correlate multimodal indicators with specific task antecedents, and cannot automatically adjust robot behavior based on risk prediction.

[0013] To achieve the above objectives, this invention uses a robot to perform specific interactive tasks. During the task, multimodal data such as children's visual behavior, speech expression, physiological state, spatial movement, tactile or force feedback, environmental stimuli, and robot logs are collected simultaneously. This data is then used to construct an individualized temporal risk model, which predicts the probability of sensory overload, emotional breakdown, self-harm or aggression, escape, collision, or training interruption occurring within the next 30, 60, and 120 seconds. Finally, based on the predicted risk level, the robot is adaptively triggered to perform actions such as destimulation, backing away, pausing, demonstrating help-seeking expressions, and alerting the caregiver, thereby achieving the unification of rehabilitation training and safety monitoring for children with special needs.

[0014] The present invention provides a first aspect of a method for multimodal risk perception and safety monitoring of autistic children based on robot interaction tasks, comprising the following steps:

[0015] In response to a specific interactive task performed by the robot, multimodal raw data of the child is collected synchronously. The multimodal raw data includes at least two or more of the following: visual behavior data, speech expression data, physiological state data, spatial movement data, tactile or force feedback data, environmental stimulus data, and the robot's own task log data. The robot includes at least a mobile chassis and mechanical components.

[0016] Using the robot's own event timestamps as the main axis, the multimodal raw data is clock-aligned and sliced, and the temporal feature indicators corresponding to each modality are extracted respectively.

[0017] The extracted temporal feature indicators are input into a task-guided temporal multimodal risk perception network, and quality perception fusion is performed through a cross-modal attention mechanism to obtain a fused state vector.

[0018] Based on the fused state vector, time series prediction is performed, and the risk probability and time series prediction uncertainty of multiple types of risks within multiple preset time windows in the future are output. A comprehensive risk score is calculated based on the risk probability and the time series prediction uncertainty.

[0019] Control decisions are made based on the comprehensive risk score and the child's individualized dynamic risk threshold, driving the robot to adaptively execute corresponding multi-level safety intervention actions.

[0020] Furthermore, the specific interactive tasks include one or more of the following: name-calling-shared attention tasks, sensory comfort expression tasks, alternating cooperation and imitation tasks, delayed gratification and help-seeking tasks, safety boundary and spatial transfer tasks, and natural companionship free play tasks.

[0021] Furthermore, the temporal characteristic indicators corresponding to each mode include:

[0022] Visual behavioral indicators include face detection bounding boxes, facial key points, facial motion unit intensity, blink frequency, gaze direction, co-attention switching, head posture, body skeleton points, hand key points, root mean square velocity of upper limbs, frequency of repeated patting or shaking, and one or more of the following: sitting posture and falling posture.

[0023] Speech and language metrics include one or more of the following: sound pressure level, fundamental frequency, energy, speech rate, pauses, probability of screaming or crying, emotional acoustic features, speech-to-text transcription, specific keywords, semantic completeness, and response delay after robot prompts.

[0024] Physiological indicators include one or more of the following: heart rate, heart rate variability, skin conductance level, peak skin conductance response, respiratory rate, body movement intensity, and skin temperature changes.

[0025] Space and safety indicators include one or more of the following: relative distance between the child and the robot, relative distance between the child and the exit or hazard, relative speed, collision time, number of boundary crossings, speed of the moving chassis, distance between the arm ends of the mechanical components, and changes in the center of pressure of the floor mat.

[0026] Haptic or force feedback metrics: including haptic or force feedback data collected by the robot's collision or touch sensors;

[0027] Interaction and task metrics include one or more of the following: task stage, number of prompts, stimulus intensity, waiting time, number of failures, presentation status of rewards or preferences, number of pauses, number of times the child initiates a task, task completion rate, and recovery time.

[0028] Environmental indicators include one or more of the following: noise decibels, light intensity, number of people in the room, stranger entry incidents, temperature and humidity, and background music or sudden noise incidents.

[0029] Furthermore, clock-aligned slicing is performed on the original multimodal data, including:

[0030] The multimodal raw data is divided into sliding windows of preset length and scrolling according to preset step size;

[0031] The quality of the missing modal data is scored, and a masking mechanism is used to preserve the missing information in the sliding window.

[0032] Furthermore, before inputting the task-guided temporal multimodal risk perception network, the method further includes single-modal encoding of the temporal feature indicators corresponding to each modality:

[0033] The skeletal and motion features in the visual behavior indicators are encoded using a facial expression recognition network and a spatiotemporal graph convolutional network or a video Transformer.

[0034] A speech feature extraction model based on a self-attention mechanism (such as the Conformer or wav2vec2 model) is used to extract the emotional acoustic representation in the speech and language indicators, and keywords are extracted in combination with speech recognition.

[0035] The physiological indicators were encoded using a one-dimensional convolutional neural network and a bidirectional long short-term memory network.

[0036] The task context corresponding to the robot's own task log data is represented by an event embedding vector.

[0037] Furthermore, quality-perceived fusion is performed through a cross-modal attention mechanism to obtain a fused state vector, the calculation formula of which is:

[0038]

[0039] in, For the first Quality-perceived attention weights for each modality For query vector, For the first The key vectors corresponding to each mode and It is a linear transformation matrix. For vector dimensions, For the first Quality score of each modality of data. This is the adjustment factor for the quality score;

[0040] The fused state vector is obtained by fusing the modal codes according to the quality-aware attention weights. The fused state vector is also obtained by weighted summation (or linear transformation after concatenation) of the modal codes according to the quality-aware attention weights.

[0041] Furthermore, before outputting the risk probability, it also includes:

[0042] Construct an individualized risk map with a structure of antecedent event A, child state S, aura behavior P, risky behavior B, and intervention outcome C. The nodes of the individualized risk map include increased noise, task failure, covering ears, backing away, increased heart rate, screaming, and rushing to the door.

[0043] A graph attention network is used to calculate the contribution of each node to future risk.

[0044] Furthermore, the multiple types of risks include sensory overload, self-harm or aggression, escape, collision or fall, training interruption, and social withdrawal; the multiple preset time windows in the future include at least 30 seconds, 60 seconds, and 120 seconds in the future;

[0045] The formula for calculating the comprehensive risk score is as follows:

[0046]

[0047] in, For the current moment The overall risk score, This represents the total number of risk types. For the first Types of risks in the future time window The predicted probability of occurrence within the timeframe, For the first Weighting coefficients for risk classes For the first Uncertainty in the time-series prediction of risk-like conditions The penalty weighting coefficient is for uncertainty.

[0048] Furthermore, the individualized dynamic risk threshold for children is updated online using the following formula:

[0049]

[0050] in, For the updated individualized risk thresholds for children, The preset initial global threshold, This represents the average fluctuation range of the child's stable period across a preset number of tasks. The success rate of the most recent robotic intervention for the children. and This is the dynamic adjustment coefficient.

[0051] Furthermore, the step of driving the robot to adaptively execute corresponding multi-level safety intervention actions includes:

[0052] When the overall risk score is less than a first preset threshold, the robot is controlled to continue the current task;

[0053] When the overall risk score is greater than or equal to a first preset threshold and less than a second preset threshold, the robot is controlled to reduce the difficulty and intensity of the current task.

[0054] When the comprehensive risk score is greater than or equal to the second preset threshold and less than the third preset threshold, the robot is controlled to pause the current task, the mobile chassis is controlled to retreat to a preset safe distance, and the robot is controlled to display prompt information to guide the child to express their needs.

[0055] When the comprehensive risk score is greater than or equal to the third preset threshold, or when self-harm or escape actions are detected, the mobile chassis and mechanical components are controlled to immediately stop moving and an alarm notification is sent to an external terminal.

[0056] Furthermore, when controlling the robot to reduce the difficulty of the current task or controlling the back distance of the mobile chassis, the comprehensive risk score and spatial motion data are embedded into a path cost function for path planning. The path cost function is:

[0057]

[0058] in, Candidate nodes in the path The overall cost of the comprehensive assessment The actual cost of the movement path. For heuristic estimation of path cost, The emotional and escape risk costs for the child when the robot reaches the candidate node. To predict the cost of a collision, The penalty for being near an exit or a pre-designated danger zone, These are the weighting coefficients for each item.

[0059] A second aspect of the present invention provides a multimodal risk perception and safety monitoring device for autistic children based on robot interaction tasks. This device is used to implement the above-described method and includes:

[0060] The data acquisition module is used to synchronously acquire multimodal raw data of children in response to specific interactive tasks performed by the robot. The multimodal raw data includes at least two or more of the following: visual behavior data, speech expression data, physiological state data, spatial movement data, tactile or force feedback data, environmental stimulus data, and the robot's own task log data. The robot includes at least a mobile chassis and mechanical components.

[0061] The feature extraction module is used to perform clock-aligned slicing of the multimodal raw data with the robot's own event timestamp as the main axis, and extract the time-series feature indicators corresponding to each modality respectively.

[0062] The risk perception and prediction module is used to input the extracted temporal feature indicators into a task-guided temporal multimodal risk perception network, perform quality perception fusion through a cross-modal attention mechanism to obtain a fusion state vector, perform temporal prediction based on the fusion state vector, output the risk probability and temporal prediction uncertainty of multiple types of risks in multiple preset time windows in the future, and calculate a comprehensive risk score based on the risk probability and the temporal prediction uncertainty.

[0063] The safety decision control module is used to make control decisions based on the comprehensive risk score and the child's individualized dynamic risk threshold, and drive the robot to adaptively execute corresponding multi-level safety intervention actions.

[0064] Compared with existing technologies, this invention has at least the following significant advantages due to its specific robot interaction task design, temporal multimodal feature fusion, individualized risk chain modeling, and adaptive closed-loop safety intervention control:

[0065] By aligning the collected data with specific robot interaction tasks and multimodal clocks, this approach achieves a deep correlation between feature extraction and the antecedents of specific rehabilitation training tasks, improving the subjective interpretability and objective accuracy of risk perception. Existing technologies often rely on generalized offline videos or single physiological signals detached from context, making it difficult to explain the true causes of emotional fluctuations. This solution utilizes the robot's proactive execution of specific structured tasks such as name-calling for joint attention, sensory comfort expression, alternating cooperation and imitation, delayed gratification and seeking help, and safety boundaries and spatial transfer. Using the robot's own event timestamps as the main axis, it performs sliding window slicing and clock alignment of multimodal data, including visual behavior, verbal expression, physiological state, spatial movement, tactile or force feedback, and task logs. This method ensures that each set of collected feature indicators, such as reaction latency, eye avoidance, heart rate variability, and peak skin conductance, possesses a highly clear antecedent explanation within the interaction context, laying an objective and interpretable data foundation for the subsequent accurate capture of warning signs of escalating danger.

[0066] By employing a task-guided temporal multimodal risk perception network and individualized risk chain graph modeling, this patent achieves temporal multi-label prediction of short-term safety risks, elevating the autism rehabilitation robot from "executing fixed training scripts" to a closed-loop system "capable of identifying children's risk states and safely adjusting tasks." In calculating multimodal attention weights, this patent introduces quality scoring, implementing dynamic masking and quality compensation for missing or low-quality modalities in real-world scenarios, ensuring robustness of cross-modal fusion. Based on this, the system constructs an individualized risk graph with a five-stage temporal topology structure including antecedent event A, child state S, aura behavior P, risky behavior B, and intervention result C, and applies a graph attention network to calculate the contribution of each aura node to future risks. This architecture can simultaneously and proactively output the risk probability and model temporal prediction uncertainty within multiple short time windows (30 seconds, 60 seconds, and 120 seconds) for various types of risks such as sensory overload, self-harm or aggression, escape, collision or fall, training interruption, and social withdrawal, achieving early warning of aura behavior.

[0067] By introducing an online mechanism for updating individualized dynamic risk thresholds for children, this solution effectively addresses the issues of highly sensitive false alarms and low-response missed detections caused by significant individual differences among children with autism. Addressing the industry pain point of vastly different behavioral precursors during emotional escalation among individuals with autism, this solution completely abandons the traditional fixed and uniform threshold criteria. The system can dynamically and online adjust its current individual risk threshold limits based on the average fluctuation range of a specific child's stable periods over a preset number of tasks, as well as the child's recent success rate with robot intervention. This feedback adjustment mechanism allows the system to adaptively adjust its warning sensitivity based on the child's natural baseline and real-time intervention response. This ensures high sensitivity for 24 / 7 monitoring while significantly reducing frequent false alarms and serious missed detections due to individual differences, improving the system's robustness across different children and scenarios.

[0068] By driving the robot to execute multi-level safety intervention actions and path decision-making mechanisms, a fully closed-loop control system is achieved, moving from "passive identification" to "active intervention." This reduces the probability of escalating problematic behaviors and improves the continuity of rehabilitation training and the safety of care. This invention endows the robot with the function of a safety control execution entity. The system directly maps the calculated comprehensive risk score to hardware-level control decisions for the mobile chassis and mechanical components: when a risk initially appears, it automatically reduces the difficulty and intensity of the current task; when the risk escalates further, it actively controls the mobile chassis to retreat to a preset safe distance and presents prompts to guide the child to correctly express their need for help; and it immediately implements an emergency stop when self-harm, escape, or other emergency actions are detected. When controlling the mobile chassis retreat or reducing difficulty, path planning is performed by embedding the comprehensive risk score into the path cost function. This allows the robot to proactively choose a detour path when the child's emotions are agitated or they are approaching physical boundaries, effectively avoiding the risk of injury to special needs children and caregivers at the physical control level.

[0069] By constructing a long-term database combining structured tasks and free play in a natural state, the system provides an objective basis for adjusting personalized rehabilitation courses, achieving a unified approach to rehabilitation training and safety monitoring. The system not only induces reproducible social behavior indicators through standardized tasks but also collects baseline data on children's spontaneous social interactions, emotional fluctuations, and repetitive behaviors during natural play. Furthermore, an active learning mechanism, reviewed by experts, generates risk warning labels, continuously supplementing the long-term rehabilitation database. This results in structured reports that not only display task completion rates and quantitative indicators but also visually present each child's unique risk warning chains and key contribution modalities. This helps assessment experts and therapists identify the triggers leading to emotional escalation and effective robotic interventions, providing scientific and objective data support for the precise adjustment of personalized rehabilitation courses. Attached Figure Description

[0070] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0071] Figure 1 This is a schematic diagram of the multimodal risk perception and safety monitoring method for children with autism according to the present invention;

[0072] Figure 2 This is a system architecture diagram of the present invention;

[0073] Figure 3 This is a flowchart of the robot interaction task and data collection indicators of the present invention;

[0074] Figure 4 This is a flowchart of the temporal multimodal risk perception algorithm of the present invention;

[0075] Figure 5 This is a diagram illustrating the risk perception safety closed loop and robot control of the present invention. Detailed Implementation

[0076] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.

[0077] Example 1

[0078] like Figure 1 As shown, this embodiment provides a multimodal risk perception and safety monitoring method for children with autism based on robot interaction tasks, including the following steps:

[0079] In response to a specific interactive task performed by the robot, multimodal raw data of the child is collected synchronously. The multimodal raw data includes at least two or more of the following: visual behavior data, speech expression data, physiological state data, spatial movement data, tactile or force feedback data, environmental stimulus data, and the robot's own task log data. The robot includes at least a mobile chassis and mechanical components.

[0080] To ensure that subsequent data collection has a clear clinical baseline and initial physical safety boundaries, this embodiment includes a step of establishing an individual child baseline and safety parameters before the robot performs specific interactive tasks: Upon initial use, the system records the child's age, diagnosis, past problematic behaviors, sensory sensitivity type, preferences, taboo stimuli, common expressions, and safety boundaries. The system then continuously collects the child's basic facial expressions, spontaneous repetitive movements, heart rate, heart rate variability, skin conductance, respiration, and natural speech during a 3-5 minute resting period, thus forming an individual baseline. Simultaneously, the robot sets an initial physical safety distance based on the child's height and mobility. In desktop interaction scenarios, the initial safety distance is set to be no less than 0.8 meters, and in mobile following scenarios, it is set to be no less than 1.2 meters. The value of this initial safety distance can be adjusted according to actual safety monitoring standards. The technical advantage of this step is that it establishes a personalized safety monitoring starting point for autistic children with highly heterogeneous and unique sensory sensitivity types, avoiding the physical collision risks associated with traditional fixed mechanical settings or defensive anxiety caused by excessive spatial pressure.

[0081] Using the robot's own event timestamps as the main axis, the multimodal raw data is clock-aligned and sliced, and the temporal feature indicators corresponding to each modality are extracted. Specifically, during the execution of structured robot tasks, the system uses the high-precision system clock of the robot interaction events as the main axis to perform multi-channel millisecond-level alignment of the heterogeneous data streams collected concurrently by various sensors. On this basis, the system not only records the start and end times of each task, prompt type, number of prompts, intensity of output audio-visual stimuli, task failure events, and presentation status of rewards or preferences in real time, but also simultaneously extracts the spatial and environmental features of the multimodality. Among them, on the spatial side, the distance sensor outputs the child's real-time relative physical position and relative speed of movement relative to the robot chassis, relative to the room exit, relative to the corner of the training table, and relative to the on-site caregiver; on the environmental side, real-time noise decibels, light intensity, total number of people in the room, stranger entry events, temperature and humidity, and sudden background noise events are collected simultaneously. This step completely eliminates the feature misalignment problem caused by the asynchronous physical sampling rates of different sensors by rigidly aligning the underlying high-dimensional heterogeneous video streams and physiological sequences with the high-level robot interaction context events on the time axis. This ensures that each extracted set of feature indicators has an extremely clear interactive task context and antecedent explanation of the interaction causality.

[0082] The extracted temporal feature indicators are input into a task-guided temporal multimodal risk perception network, and quality perception fusion is performed through a cross-modal attention mechanism to obtain a fused state vector.

[0083] In the specific algorithm processing, the system first divides the aligned multimodal temporal indicators into time window sequences using a sliding window with a length of 5 seconds and a scrolling step of 0.5 seconds. To address the technical challenges of video data loss caused by the temporary loss of physiological characteristics due to frequent and vigorous movements of children or large-angle facial deflections in real rehabilitation institutions or home environments with free interaction, this embodiment performs real-time quality scoring on each missing or low-quality modal data during slicing using real-time quality assessment rules. A masking mechanism is then used to retain missing information in the current sliding window. Subsequently, the features with quality bias are fed into a task-guided temporal multimodal risk perception network for cross-modal attention fusion calculation. The high-frequency scrolling step of 0.5 seconds ensures that the system has extremely high time sensitivity at the data slicing level, while the masking mechanism based on quality scoring enables the feature fusion algorithm to maintain strong mathematical robustness in the face of complex noise and uncertain packet loss, avoiding the defect of abnormal interruption of the entire monitoring process due to incomplete single-modal data.

[0084] Based on the fused state vector, a temporal prediction is performed, outputting the risk probability and temporal prediction uncertainty of multiple types of risks within multiple preset time windows in the future, and calculating a comprehensive risk score based on the risk probability and the temporal prediction uncertainty; specifically, the task-guided temporal multimodal risk perception network takes the fused features extracted by the sliding window as input, combines the current interactive task context of the robot and the child's individualized baseline, and outputs in parallel the probability of occurrence of six specific risks, namely sensory overload, self-harm or aggression, escape, collision or fall, training interruption, and social withdrawal, within multiple short time windows such as 30 seconds, 60 seconds, and 120 seconds in the future, and outputs the temporal prediction uncertainty of the model for the current judgment result. The system integrates and infers the predicted probabilities of different risk types across various future time windows to form specific risk warning signs: If a child exhibits a specific combination of "sudden and high-level changes in skin conductance, visual ear-covering actions, physical retreat, and an increased probability of audio screaming or crying" during a sensory task, the system will automatically increase the probability output of sensory overload risk; if a child exhibits a specific combination of "grabbing or hitting a box, significantly increased speaking volume, rapid increase in heart rate, and increased task abandonment rate" during a delayed gratification task, the system will automatically increase the probability output of emotional breakdown or aggression risk; if a child exhibits a combination of "physical acceleration towards the room door, continuous non-response to the robot's name call, and a rapid increase in the physical distance from the robot's chassis" during a spatial transfer task, the system will automatically increase the predicted probability of escape risk. This step successfully achieves highly sensitive, multi-label, and forward-looking extrapolation of the warning signs of autistic behavioral outbursts, overcoming the limitation of existing systems that can only report after the fact. It allows the control system and caregivers to know the probability of escalation of danger 30-120 seconds in advance, gaining extremely valuable pre-intervention and control time.

[0085] Control decisions are made based on the comprehensive risk score and the child's individualized dynamic risk threshold, driving the robot to adaptively execute corresponding multi-level safety intervention actions.

[0086] In this decision-making and control step, the system, based on the calculated comprehensive risk score and the individualized risk threshold updated online by the child's current physiological characteristics, directly drives the robot's mobile chassis and mechanical components to adaptively execute actions through a control decision algorithm: when the comprehensive risk score is less than the first preset threshold (e.g., comprehensive risk score less than 0.35), it is determined to be low risk, and the robot continues to execute the current task while maintaining the original training script; when the comprehensive risk score is greater than or equal to the first preset threshold and less than the second preset threshold (e.g., 0.35 ≤ comprehensive risk score < 0.60), it is determined to be medium risk, and the robot automatically reduces the speech rate and volume of the speech synthesis, reduces social requirements, shortens the rule waiting time of the current task, and replaces the selection with image elements on the screen to actively reduce the intensity of audio-visual stimulation and the difficulty of the task; when the comprehensive risk score is greater than or equal to the second preset threshold and less than the third preset threshold... When a threshold (e.g., 0.60 ≤ comprehensive risk score < 0.80) is reached, the system is deemed high-risk. The robot immediately pauses the current interactive task, moves the chassis backward to maintain a preset physical safety distance, and displays a prominent graphical interface for "Pause / Help / Discomfort" on the screen. Simultaneously, it plays calm, soothing phrases to guide the child to express their needs through gestures or touch. When the comprehensive risk score is greater than or equal to the third preset threshold (e.g., comprehensive risk score ≥ 0.80), the system is deemed an emergency risk. Alternatively, if obvious self-harm or escape behaviors are detected directly through visual and tactile sensors, the system triggers an emergency stop decision. The chassis and mechanical components immediately cease all physical movement, an audible and visual low-stimulation alarm is activated, and an alarm notification is sent to the mobile terminal, tablet, or institutional management system of the external caregiver or therapist. Simultaneously, multimodal data for 60 seconds before and after the event is automatically recorded for subsequent review.

[0087] To ensure the hierarchical control logic possesses high interactive control flexibility, this embodiment also allows therapists to manually confirm, correct, or add notes to the system's judgment labels after the report is generated, enabling the model to dynamically update the child's individual risk chain. The following detailed description of this multi-level safety intervention and control process uses specific application examples:

[0088] Scenario 1 (Sensory Overload Closed-Loop Control in a Rehabilitation Center Training Room): A child with special needs enters the training room to perform a sensory comfort expression task. The robot presents sound stimuli at a medium-to-high intensity. The system detects a combination of signals in the child, including brief frowning, backing away, and a rapid increase in skin conductance. The network model calculates that the child's sensory overload risk probability within the next 60 seconds reaches 0.72 (triggering the high-risk interval boundary). The control decision algorithm immediately drives the robot to reduce the speech synthesis speed and volume, controls the mobile chassis to actively move backward by 0.5 meters to create space, and displays a graphical "Pause / Quiet / Continue" card selection interface on the screen to guide the child in expressing their needs. The child then selects "Quiet," and the system detects that their comprehensive risk score rapidly drops to 0.38 (below the second preset threshold). The robot switches back to a low-stimulation training state, ensuring the safe and continuous execution of the rehabilitation task.

[0089] Scenario 2 (Emotional Intervention and Behavioral Reinforcement Control in Delayed Gratification Tasks): When performing a delayed gratification task, the robot locks a preferred toy car in a transparent box and guides the child to ask for help. The child fails to utter the correct verbal expression during their first attempt to grab the box, and their heart rate and upper limb activity increase significantly. Based on this, the model immediately determines that their emotional breakdown risk score has reached 0.65 (greater than the second preset threshold). The control decision algorithm then controls the robot to proactively shorten the rule waiting time, actively demonstrate the verbal expression "help me" in a calm voice, and simultaneously display a help card on the screen. When the child touches the help card as prompted, the robot immediately activates mechanical components to automatically open the transparent box and reward the toy car, successfully reinforcing the help-seeking behavior before an emotional breakdown occurs.

[0090] This step is the first to directly and seamlessly transform the algorithm prediction results at the software level into adaptive risk avoidance actions of the chassis and robotic arm of the special children's rehabilitation robot. It builds a physical closed loop of safety protection at the lowest level, which not only prevents the escalation of problem behaviors at the physical level, but also protects children and caregivers from accidental personal injury to the greatest extent, while taking into account both the continuity of special education and training and the safety of supervision.

[0091] As one implementation method, the specific interactive tasks described in this embodiment include one or more of the following: name-calling-shared attention task, sensory comfort expression task, alternating cooperation and imitation task, delayed gratification and help-seeking task, safety boundary and spatial transfer task, and natural companionship free play task.

[0092] In specific system deployment and training implementation, the above six specific interactive tasks are given a clear standardized structure and monitoring behavior induction purpose in the process to avoid indiscriminate data collection: Among them, in the T1 name-calling-co-attention task, the robot calls the child's name through speech synthesis at a distance of 1.2-1.8 meters, then points to the target toy or presents the target toy on the screen image, and issues a prompt. Indicators such as name-calling response latency, head turning angular velocity, first eye contact point, duration of co-attention, and number of times the target object is gazed are collected to capture early signs of task stress such as persistent non-response or increased eye avoidance; in T2 In the sensory comfort expression task, the robot progressively presents low-intensity changes in sound decibels or light intensity, demonstrating standardized expressions. Data is collected on facial motor unit intensity (such as frowning / squinting / covering ears), peak skin conductance, respiratory rate, and changes in the relative distance between the child and the robot. This data is used to assess sensory overload and self-harm risk, and to immediately control the robot to reduce stimulation. In the T3 alternating collaboration and imitation task, the robot and child take turns performing tasks such as building blocks, passing a ball, or imitating actions. Waiting times and rule changes are set, and data are collected on waiting time, number of attempts to grab objects, imitation accuracy, smoothness of movement trajectories, and frequency of repetitive behaviors. This data is used to capture sensory information. To mitigate the risks of escalating frustration and potential aggression while reducing rule complexity, in the T4 delayed gratification and help-seeking task, preferred toys are briefly placed in a transparent box or key components are deliberately omitted from the task to encourage children to seek help. Indicators such as the probability of help-seeking initiation, help-seeking latency, grasping / slapping of the box, leaving the seat, crying, and increased heart rate are collected to assess the risk of emotional breakdown due to the inability to express needs and to provide alternative expression boards. In the T5 safety boundary and spatial transfer task, the robot guides children from the play area to the training mat or the area opposite the doorway, setting safety boundaries and obstacles along the way, and collecting child-machine interaction data. The robot measures distance to the person, speed approaching the door, number of boundary crossings, collision prediction distance, and sudden acceleration or fall / balance posture to capture escape and collision / fall risks and perform risk-aware path planning. In the T6 natural companionship free play task, the robot switches to a low-initiative guardianship state for 10-30 minutes, responding only when the child actively approaches, calls out, or the risk increases. This is used to collect baselines of spontaneous social initiation, emotional fluctuations, and repetitive behaviors in the unstructured natural state, forming a long-term and realistic individual baseline of the child's natural state, which reduces model underreporting and false positives outside of structured tasks. This step, by completely deconstructing and embedding multimodal safety monitoring into these six types of real robot rehabilitation training and daily companionship processes with clear clinical prognostic significance and triggering logic, perfectly overcomes the shortcomings of insufficient ecological validity in traditional single-room short-term scale assessments, achieving long-term, fully closed-loop behavioral quantification and safety intervention in the natural state.

[0093] As one implementation method, the time-series characteristic indicators corresponding to each mode in this embodiment include:

[0094] Visual behavioral indicators include face detection bounding boxes, facial key points, facial motion unit intensity, blink frequency, gaze direction, co-attention switching, head posture, body skeleton points, hand key points, root mean square velocity of upper limbs, frequency of repeated patting or shaking, and one or more of the following: sitting posture and falling posture.

[0095] Speech and language metrics include one or more of the following: sound pressure level, fundamental frequency, energy, speech rate, pauses, probability of screaming or crying, emotional acoustic features, speech-to-text transcription, specific keywords, semantic completeness, and response delay after robot prompts.

[0096] Physiological indicators include one or more of the following: heart rate, heart rate variability, skin conductance level, peak skin conductance response, respiratory rate, body movement intensity, and skin temperature changes.

[0097] Space and safety indicators include one or more of the following: relative distance between the child and the robot, relative distance between the child and the exit or hazard, relative speed, collision time, number of boundary crossings, speed of the moving chassis, distance between the arm ends of the mechanical components, and changes in the center of pressure of the floor mat.

[0098] Haptic or force feedback metrics: including haptic or force feedback data collected by the robot's collision or touch sensors;

[0099] Interaction and task metrics include one or more of the following: task stage, number of prompts, stimulus intensity, waiting time, number of failures, presentation status of rewards or preferences, number of pauses, number of times the child initiates a task, task completion rate, and recovery time.

[0100] Environmental indicators include one or more of the following: noise decibels, light intensity, number of people in the room, stranger entry incidents, temperature and humidity, and background music or sudden noise incidents.

[0101] In this feature extraction stage, visual behavior indicators use facial key points and body skeleton point models to finely calculate the intensity of facial action units, and output the root mean square velocity of the upper limbs and the frequency of repeated patting / shaking through hand key points; speech and language indicators use automatic speech recognition and transcription to extract discrete keyword semantic information such as rejection words, help-seeking words, and preference words expressed by children locally; physiological indicators use wearable or non-contact sensors to acquire heart rate, heart rate variability (including RMSSD and SDNN time-domain indicators), skin conductance level and peak response, respiratory rate, and skin temperature changes to objectively assess the child's stress and emotional arousal at the lowest level. This step utilizes a comprehensive and cross-modal feature indicator matrix, which can not only extremely sensitively capture the high internal tension and stress increase in autistic children who cannot verbally express themselves due to speech communication barriers through deep physiological arousal indicators such as heart rate and skin conductance surges, but also explain the specific interactive context of this fluctuation through visual posture and the robot's own interaction logs, achieving a comprehensive and high-fidelity quantitative perception of the psychological arousal and external behavioral performance of special children.

[0102] As one implementation method, this embodiment performs clock-aligned slicing on the multimodal raw data, including:

[0103] The multimodal raw data is divided into sliding windows of preset length and scrolling according to preset step size;

[0104] The quality of the missing modal data is scored, and a masking mechanism is used to preserve the missing information in the sliding window.

[0105] In the specific system implementation, the preset length of the sliding window is set to 5 seconds, and the preset scrolling step size is set to 0.5 seconds. The video frame sequence, audio frame, physiological sequence, distance sequence, and robot log are sliced ​​using a high-frequency scrolling window. This 0.5-second refresh rate ensures that the algorithm has extremely high real-time response capabilities at the data flow control level, guaranteeing second-level timing accuracy when risk warnings are detected. This provides a sufficiently generous golden safety avoidance time window for the robot to execute adaptive intervention actions (such as moving the chassis backward, immediately stopping the chassis and mechanical parts) and for the caregiver to issue alarms.

[0106] As one implementation method, this embodiment further includes single-modal encoding of the temporal feature indicators corresponding to each modality before inputting the task-guided temporal multimodal risk perception network:

[0107] The skeletal and motion features in the visual behavior indicators are encoded using a facial expression recognition network and a spatiotemporal graph convolutional network or a video Transformer.

[0108] The Conformer or wav2vec2 model is used to extract the emotional acoustic representation in the speech and language indicators, and keywords are extracted in combination with speech recognition;

[0109] The physiological indicators were encoded using a one-dimensional convolutional neural network and a bidirectional long short-term memory network.

[0110] The task context corresponding to the robot's own task log data is represented by an event embedding vector.

[0111] Specifically, in the algorithm processing flow, visual feature encoding uses facial expression recognition networks and spatiotemporal graph convolutional networks or video Transformers to encode skeletal and motion features; speech feature encoding uses Conformer or wav2vec2 to extract high-dimensional emotional acoustic representations, and combines lightweight speech recognition to extract discrete text keyword embeddings; physiological feature encoding uses a one-dimensional convolutional neural network combined with a bidirectional long short-term memory network to encode heart rate, heart rate variability, skin conductance, and respiratory waveforms into deep stress arousal representations; task context is directly represented by event embedding vectors. The technical effect of this step is that, for four types of heterogeneous data sources with completely different spatial dimensions and temporal scales—video streams, audio waveforms, continuous one-dimensional physiological signals, and discrete robot interaction events—the most efficient single-modal deep feature encoding network is customized and matched. This can extract the special external behavior and internal psychological features of children contained in each modality to the greatest extent, effectively avoiding the problem of deep feature representation collapse or long-tail feature submersion caused by direct and forceful feature splicing.

[0112] As one implementation method, this embodiment uses a cross-modal attention mechanism to perform quality-perceived fusion to obtain a fused state vector, the calculation formula of which is:

[0113]

[0114] in, For the first Quality-perceived attention weights for each modality For query vector, For the first The key vectors corresponding to each mode and It is a linear transformation matrix. For vector dimensions, For the first Quality score of each modality of data. This is the adjustment factor for the quality score;

[0115] The modal codes are fused according to the quality-aware attention weights to obtain the fused state vector.

[0116] In the fusion computation of the cross-modal attention mechanism, the input includes visual, speech, physiological, spatial, and task context encoding vectors. When calculating attention allocation, the real-time quality score of each modality, calculated through quality assessment, is directly introduced into the Softmax distribution as a bias term and multiplied by an adjustment coefficient. This weights or masks the features of each modality, ultimately outputting a fused state vector. The technical effect of this step is that, through the underlying quality score constraints of the mathematical formula, the system possesses self-cleaning and missing value compensation functions: when a modality's quality score drops sharply due to a temporary detachment of the wristband or a child tilting their head, causing a surge in sensor noise, packet loss, or occlusion, the cross-modal attention weights automatically decay to suppress the contribution weight of that modality in the final fused vector. Instead, the attention weights of other intact modalities are adaptively increased to dominate risk inference, significantly improving the fusion robustness in complex, unconstrained intervention training environments.

[0117] As one implementation method, this embodiment further includes the following before outputting the risk probability:

[0118] Construct an individualized risk map with a structure of antecedent event A, child state S, aura behavior P, risky behavior B, and intervention outcome C. The nodes of the individualized risk map include increased noise, task failure, covering ears, backing away, increased heart rate, screaming, and rushing to the door.

[0119] A graph attention network is used to calculate the contribution of each node to future risk.

[0120] In the specific implementation of graph modeling, the system constructs and maintains a temporal causal topological graph network that deeply simulates the individualized emotional evolution path of children. Graph network nodes store the historical trajectory of a specific individual as a linkage path between antecedents, states, precursors, risky behaviors, and intervention results. The graph attention network can dynamically calculate the causal transfer contribution of each node to the future outbreak of multiple types of risks based on the active nodes activated in the training environment at the current moment (such as nodes simultaneously detecting task failure, regression, and heart rate increase). The technical effect of this step is that by connecting scattered multimodal features through an explicit causal graph topological structure, the originally "black box" deep learning temporal prediction model acquires a high level of technical and logical interpretability. The generated report can clearly reconstruct the risk evolution path, facilitating therapists to accurately identify hidden antecedents leading to emotional escalation, thereby scientifically and quantitatively adjusting personalized rehabilitation courses.

[0121] Furthermore, the multiple types of risks include sensory overload, self-harm or aggression, escape, collision or fall, training interruption, and social withdrawal; the multiple preset time windows in the future include at least 30 seconds, 60 seconds, and 120 seconds in the future;

[0122] The formula for calculating the comprehensive risk score is as follows:

[0123]

[0124] in, For the current moment The overall risk score, This represents the total number of risk types. For the first Types of risks in the future time window The predicted probability of occurrence within the timeframe, For the first Weighting coefficients for risk classes For the first Uncertainty in the time-series prediction of risk-like conditions The penalty weighting coefficient is for uncertainty.

[0125] In this step, the multi-label predicted probabilities output by the network model are jointly weighted and calculated with the model's own predicted uncertainty, which is calculated through uncertainty estimation (such as Monte Carlo sampling or output entropy penalty). The risk weight coefficient in the formula is determined based on the severity and danger level of the specific child's historical risks (e.g., if the child has a history of severe self-injury scratching, the weight coefficient corresponding to the self-injury risk will be assigned a high value). This step not only enables the parallel output of high-precision multi-label risk probabilities across multiple future time horizons, but also effectively suppresses control malfunctions caused by "blind confidence" when the neural network model faces unknown and unfamiliar environmental disturbances or rare long-tail unusual body movements by linearly weighting the predicted uncertainty as a penalty term with the probability value in the formula. This greatly enhances the safety defense boundary of subsequent automated control decisions.

[0126] As one implementation method, the individualized dynamic risk threshold for children in this embodiment is updated online using the following formula:

[0127]

[0128] in, For the updated individualized risk thresholds for children, The preset initial global threshold, This represents the average fluctuation range of the child's stable period across a preset number of tasks. The success rate of the most recent robotic intervention for the children. and This is the dynamic adjustment coefficient.

[0129] When implementing adaptive threshold adjustments, the system records the fluctuation range of a specific child during a stable period in a preset number of normal tasks over a long period. During online updates, the system dynamically fine-tunes the current risk warning sensitivity by adding a positive adjustment term determined by the individual's baseline fluctuations to the initial global universal threshold and subtracting a feedback term determined by the recent intervention success rate. This step utilizes a strict mathematical closed loop to achieve adaptive negative feedback adjustment: when the system faces "highly sensitive" children with high baselines and drastic daily fluctuations in emotional expression, the formula automatically widens the risk threshold boundary appropriately to prevent frequent false alarms caused by the individual's high baseline of routine behavior, thus avoiding frequent interruptions to the continuity of normal teaching and training; conversely, when the recent intervention success rate declines continuously, indicating that the child may have developed frustration or tolerance for escalating emotions to the current interaction pattern, the formula automatically narrows the risk threshold boundary, increasing the warning sensitivity and completely eliminating the serious risk of underreporting for low-responsive or concealed children.

[0130] As one implementation method, the method of driving the robot to adaptively execute corresponding multi-level safety intervention actions in this embodiment includes:

[0131] When the overall risk score is less than a first preset threshold, the robot is controlled to continue the current task;

[0132] In this low-risk state, the first preset threshold is preferably configured as 0.35. When the calculated comprehensive risk score is below this threshold, it is determined that the child's current psychological arousal, external behavior, and environment are all within a safe and well-adaptable baseline range. At this time, the system controls the robot to maintain the execution speed, volume, and task arrangement logic of the original rehabilitation training script unchanged, and continues the current interaction process. The technical effect of this step is that by setting a quantified low-risk threshold boundary, the control system can avoid any unnecessary excessive intervention or protective actions when the child's state is stable, thereby maximizing the golden continuity and focus of rehabilitation training for children with special needs and ensuring rehabilitation efficiency.

[0133] When the overall risk score is greater than or equal to a first preset threshold and less than a second preset threshold, the robot is controlled to reduce the difficulty and intensity of the current task.

[0134] Under this moderate risk condition, the first preset threshold is preferably configured as 0.35, and the second preset threshold is preferably configured as 0.60. When the comprehensive risk score fluctuates between 0.35 and 0.60, it indicates that the child may have experienced some degree of emotional fluctuation or stress accumulation due to rule changes, waiting during transitions, or restrictions on preferred items. At this time, the system dynamically implements safety downgrade adjustments to the current task through control decisions. Specifically, the system controls the robot to automatically lower the speech rate and playback volume of the synthesized speech, reduce the screen display brightness, shorten the rule waiting time in the alternating collaboration, or reduce the current social or imitative interaction requirements, and replace the interface prompts on the screen with weakly stimulating graphical elements. The technical effect of this step is that it can utilize refined early fine-tuning and downgrade control to proactively reduce negative sensory stimuli and other task antecedents in the early stages when special needs children have not yet developed obvious destructive problem behaviors and are only showing early signs of accumulated emotional stress. This guides children to achieve autonomous emotional frustration regulation and greatly reduces the probability of further escalation of problem behaviors from the source.

[0135] When the comprehensive risk score is greater than or equal to the second preset threshold and less than the third preset threshold, the robot is controlled to pause the current task, the mobile chassis is controlled to retreat to a preset safe distance, and the robot is controlled to display prompt information to guide the child to express their needs.

[0136] In this high-risk state, the second preset threshold is preferably configured as 0.60, and the third preset threshold is preferably configured as 0.80. When the comprehensive risk score is between 0.60 and 0.80, it indicates that the child's stress and arousal level are approaching the critical collapse boundary, facing an imminent risk of behavioral loss of control. At this time, the control system immediately issues a pause command, controlling the robot to stop the currently executing audio-visual training script, game prompts, or limb movement, while driving the mobile chassis to actively move backward, retreating to a distance beyond the initial physical safety distance preset by the child's activity level (e.g., actively retreating 0.5 meters or maintaining a following distance of not less than 1.2 meters), creating physical space to alleviate the child's sensory anxiety. At the same time, the control robot's screen switches to display intuitive "Pause / Help / Discomfort" graphical help expression cards, supplemented by smooth and gentle voice prompts, guiding the child to express their current core needs through gestures, sign language, or touching the picture cards. The technical effect of this step is that, through the mandatory "pause task - chassis back distance - picture card guidance" three-in-one closed-loop decision, on the one hand, it quickly cuts off the sources of stimulation that cause sensory overload in special children at the physical space level, providing children with sufficient safe buffer space and preventing them from defensive scratching or bumping due to excessive spatial pressure; on the other hand, it positively transforms the negative and destructive emotions that may have evolved into self-harm or aggression into a behavioral reinforcement process of using multimodal picture cards to seek help and express discomfort in a standardized manner, thus achieving an organic combination of safety monitoring and proactive behavior correction.

[0137] When the comprehensive risk score is greater than or equal to the third preset threshold, or when self-harm or escape actions are detected, the mobile chassis and mechanical components are controlled to immediately stop moving and an alarm notification is sent to an external terminal.

[0138] In this emergency risk state, the third preset threshold is preferably configured as 0.80. When the comprehensive risk score is greater than or equal to 0.80, or when a child is detected directly by video and tactile sensors to make obvious self-harming movements such as covering their ears with their hands, or when a child is detected to be accelerating violently towards the exit or other escape or collision movements, the system immediately triggers the highest priority emergency stop decision, controlling the mobile chassis and mechanical components to immediately stop all physical displacement and mechanical movement, cutting off the potential physical threat posed by the moving components, and simultaneously activating a low-stimulation audible and visual alarm. The system also instantly sends structured data, including risk type, time window probability, and key contribution modality, concurrently through the network module to external terminals such as mobile phones or tablets held by caregivers or therapists, issuing real-time remote alarm notifications. At the same time, the system automatically locks and records multimodal data for 60 seconds before and after the event for subsequent review. The technical effect of this step is that it constructs a high-priority risk avoidance closed loop at the lowest level of robot physical control. In the extreme emergency moment when a child's behavior is out of control, the chassis and mechanical components can be completely stopped, thus completely eliminating the risk of secondary injury to the child caused by the robot's moving parts. There is no delay in notifying caregivers to intervene manually, effectively ensuring the absolute safety of rehabilitation training for children with special needs at the most critical moment.

[0139] As one implementation method, in this embodiment, when controlling the robot to reduce the difficulty of the current task or controlling the back distance of the mobile chassis, the comprehensive risk score and spatial motion data are embedded into a path cost function for path planning. The path cost function is:

[0140]

[0141] in, Candidate nodes in the path The overall cost of the comprehensive assessment The actual cost of the movement path. For heuristic estimation of path cost, The emotional and escape risk costs for the child when the robot reaches the candidate node. To predict the cost of a collision, The penalty for being near an exit or a pre-designated danger zone, These are the weighting coefficients for each item.

[0142] When the rehabilitation robot is a mobile robot with a movable chassis, and the control system drives the chassis to follow and accompany the child or perform retreat avoidance, the path planning algorithm nonlinearly integrates the physical cost of traditional movement paths with the child's real-time emotions and spatial escape risks. Among these, the penalty term... This indicates when the mobile robot reaches the candidate node. Estimated time At that time, the child's corresponding emotions and the cost of escaping risk; punishment items This represents the predicted cost of a physical collision between the chassis and boom-end mechanical components and a child or surrounding obstacle at a candidate node; penalty term. This represents the penalty incurred by a node for being too close to a room's safety exit or a physically hazardous area. The technical advantage of this step is that it completely breaks through the limitations of traditional path planning algorithms, which can only consider physical geometric obstacles. For the first time, it achieves path avoidance decision-making based on emotional and spatial risk perception: once the system detects increased stress, abnormal emotional arousal, or an escape tendency near the boundary, the path cost function automatically increases the cost of all candidate nodes around the robot and in the direction of the main entrance. This forces the mobile chassis to automatically select a safe path that, while geographically farther and slower, maintains the optimal monitoring distance from the child and adaptively bypasses the exit direction. From the perspective of physical chassis motion control, this effectively avoids the risk of serious collisions, falls, and accidental escapes between the child, caregiver, and robot.

[0143] Example 2

[0144] This embodiment provides a multimodal risk perception and safety monitoring device for children with autism based on robot interaction tasks. This device is used to implement the above-described method and includes:

[0145] The data acquisition module is used to synchronously collect multimodal raw data of the child in response to specific interactive tasks performed by the robot. This multimodal raw data includes at least visual behavioral data, speech expression data, physiological state data, spatial movement data, tactile or force feedback data, environmental stimulus data, and the robot's own task log data. The robot includes at least a mobile chassis and mechanical components. Figure 2As shown, the data acquisition module fully encompasses and interfaces with the robot interaction task layer and the multimodal acquisition layer in terms of hardware configuration, control logic, and data flow. Specifically, the data acquisition module includes a robot interaction task unit and a multimodal data acquisition unit. The robot interaction task unit controls the robot to execute specific interaction tasks, which in the robot interaction task layer are specifically manifested as name-calling and shared attention tasks, sensory comfort expression tasks, and alternating collaboration and safety boundary tasks. The multimodal data acquisition unit in the multimodal acquisition layer specifically includes multiple hardware channels for concurrently acquiring the multimodal raw data. The system employs peripheral sensors, specifically an RGBD camera, eye tracker, and microphone to collect data on the child's facial expressions, skeletal points of movement, gaze direction, and interactive voice. PPG, EDA, respiration, and force / tactile sensors collect data on the child's heart rate, skin conductance, respiratory rate, and tactile or force feedback data obtained from the robot's collision or touch sensors. Simultaneously, the system automatically records and reads robot logs and environmental data. The multiple data outputs of the data acquisition module are electrically connected to the corresponding inputs of the feature extraction module and the child's individual baseline library unit in the safety decision control module. This module architecture tightly and rigidly couples the structured interactive task layer at the robot's front end with the underlying omnidirectional multimodal acquisition layer hardware, ensuring that the collected multimodal raw data can be mapped in real-time and with high fidelity to the combined internal and external psychological arousal and behavioral performance of children with special needs under different social demands and sensory stimuli.

[0146] The feature extraction module is used to perform clock-aligned slicing of the multimodal raw data, with the robot's own event timestamps as the main axis, and extract the temporal feature indicators corresponding to each modality; such as Figure 2As shown, the feature extraction module corresponds to the edge fusion computing layer, and it specifically integrates a timestamp synchronization unit, a feature extraction unit, and a quality assessment and missing data compensation unit. The timestamp synchronization unit is used to perform the timestamp synchronization step, using the robot's own event timestamp as the main axis, and controls the concurrent acquisition of multiple heterogeneous data streams by various physical sensors to perform millisecond-level hard alignment. The feature extraction unit is used to perform the feature extraction step, and divides the aligned multimodal raw data stream into temporal feature index segments through a sliding window. The quality assessment and missing data compensation unit is used to perform the quality assessment and missing data compensation step, dynamically calculating the real-time quality score for missing or low-quality modal data caused by high-frequency body motion, motion artifacts, or head tilting occlusion, and using a masking mechanism to retain missing information in the current sliding window and perform data compensation, outputting a high-quality feature sequence to be sent to the backend network. The feature output end of the feature extraction module is connected to the feature input end of the risk perception and prediction module, and its feedback receiving end is reversely connected to the control feedback output end of the multi-level hierarchical intervention decision unit in the safety decision control module, so as to receive the feedback status after the system adaptively executes the control intervention command in real time. This module architecture utilizes the hardware and software collaborative processing of the edge fusion computing layer to efficiently complete millisecond-level alignment and noise reduction compensation of heterogeneous data at the edge near the sensor. This not only significantly reduces the bandwidth pressure of centralized data transmission, but also improves the system's mathematical robustness against noise interference and multi-sensor packet loss from the data pipeline architecture level.

[0147] The risk perception and prediction module is used to input the extracted temporal feature indicators into a task-guided temporal multimodal risk perception network, perform quality perception fusion through a cross-modal attention mechanism to obtain a fused state vector, and perform temporal prediction based on the fused state vector. It outputs the risk probability and temporal prediction uncertainty of multiple types of risks within multiple preset time windows in the future, and calculates a comprehensive risk score based on the risk probability and the temporal prediction uncertainty. Figure 2As shown, the risk perception and prediction module corresponds to the risk prediction and safety decision-making layer. Its internal hardware and algorithm architecture is specifically divided into a temporal multimodal attention fusion unit, a risk chain modeling unit, and a comprehensive risk calculation unit. The temporal multimodal attention fusion unit is used to perform temporal multimodal attention mechanism calculations, inputting the extracted temporal feature indicators into the task-guided temporal multimodal risk perception network, and introducing the quality score as a bias term into the attention matrix through a cross-modal attention mechanism to calculate a robust fusion state vector. The risk chain modeling unit is used to perform risk chain modeling steps, constructing and maintaining a five-stage causal temporal topology of antecedent event A, child state S, prodromal behavior P, risky behavior B, and intervention result C, and using a graph attention network to prospectively predict and output the risk probability and prediction uncertainty of multiple types of risks in multiple preset time windows in the future. The comprehensive risk calculation unit executes a weighted formula to calculate and output the comprehensive risk score at the current moment, which serves as the underlying quantitative basis for implementing the graded intervention strategy. The score output end of the risk perception and prediction module is electrically connected to the decision input end of the safety decision control module. This modular architecture organically integrates a temporal multimodal attention fusion mechanism with an explicit causal risk chain graph topology, constructing an intelligent reasoning brain with high clinical and logical interpretability at the core layer of the prediction network. This not only greatly improves the time foresight and classification accuracy of safety risk prediction, but also enables safety warnings to be accurately traced and associated with specific antecedent events of the task.

[0148] The safety decision control module is used to make control decisions based on the comprehensive risk score and the child's individualized dynamic risk threshold, driving the robot to adaptively execute corresponding multi-level safety intervention actions; such as Figure 2As shown, the safety decision control module fully encompasses and corresponds to the child's individual baseline database, the rehabilitation and safety execution layer, and the expert and therapist reporting layer. Internally, it is further subdivided into a child's individual baseline database unit, a multi-level hierarchical intervention decision unit, and a report output unit. The child's individual baseline database unit corresponds to the child's individual baseline database and is specifically used to store information entered during the initial interaction. It continuously collects and accumulates natural state data and task performance trajectories during subsequent rehabilitation and care processes, establishing a long-term, rolling, individualized historical behavioral baseline in the system backend, and dynamically performing personal threshold updates to output the child's current individualized dynamic risk threshold. The two input terminals of the multi-level hierarchical intervention decision unit are electrically connected to the score output terminal of the risk perception and prediction module and the threshold output terminal of the child's individual baseline database unit, respectively. Next, by comparing the comprehensive risk score with the individualized dynamic risk threshold in real time, the system outputs corresponding graded intervention strategy control commands to the mobile chassis and mechanical components at the rehabilitation and safety execution layer. This drives the robot to adaptively execute closed-loop defense actions such as robot stimulation reduction and retreat, task pause and demonstration assistance, and caregiver alarms. The control feedback output of the multi-level graded intervention decision unit is connected in reverse to the input of the feature extraction module to form an online feedback update closed loop. The report output unit is connected to the output of the multi-level graded intervention decision unit and is used to collect control logs and multimodal behavioral characteristics at each stage. It renders and outputs a visual report containing risk heatmaps, task indicator curves, and quantitative intervention suggestions at the expert and therapist report layer. This modular architecture truly realizes a complete closed-loop monitoring and control device for the behavior of children with special needs, consisting of specific interactive task induction, multimodal acquisition, edge alignment, intelligent perception and prediction, individual baseline calibration, hardware adaptive closed-loop execution, feedback and update, and quantitative report presentation. By utilizing the adaptive negative feedback adjustment of the child's individual baseline library unit, the system significantly reduces the fatigue false alarm rate under highly sensitive safety monitoring conditions. Through a graded strategy at the physical movement level, it completely eliminates the risk of personal injury in its infancy. At the same time, it replaces the traditional purely subjective assessment with an objective quantitative report layer, achieving a true organic unity between daily rehabilitation training and safety behavior monitoring for children with special needs.

[0149] Example 3

[0150] This embodiment provides a detailed execution flow and practical application scenario of a multimodal risk perception and safety monitoring method and device for autistic children based on robot interaction tasks. It comprehensively and thoroughly demonstrates and supplements the method steps of Embodiment 1 and the device structure of Embodiment 2. The safety monitoring method and device provided in this embodiment address the problems of existing autism rehabilitation robots in failing to identify safety risks in advance during real-world interaction, failing to correlate multimodal indicators with specific task antecedents, and failing to automatically adjust robot behavior based on risk prediction.

[0151] The safety monitoring device provided in this embodiment is constructed through the close collaboration of multiple underlying software and hardware sub-functional modules. Specifically, the device mainly integrates a rehabilitation robot interaction task module, a multimodal data acquisition module, a task event and time synchronization block, a multimodal feature extraction module, a temporal multimodal risk perception module, a robot safety decision-making module, and a personalized rehabilitation database module in its physical and logical architecture.

[0152] The rehabilitation robot's interactive task module includes a movable chassis, a head or screen expression unit, a speech synthesis unit, and an arm or light feedback unit; it is specifically used to drive and control the robot entity to sequentially or selectively perform the following tasks in the training site: T1 calling name and joint attention task, T2 sensory comfort expression task, T3 alternating cooperation and imitation task, T4 delayed gratification and help-seeking task, T5 safety boundary and spatial transfer task, and T6 natural companionship free play task.

[0153] The multimodal data acquisition module is equipped with red, green and blue depth cameras, wide-angle cameras, eye-tracking or gaze estimation units, microphone arrays, wearable or non-contact pulse wave and skin conductance and respiration sensors for children, floor pressure and distance sensors, robot collision or touch sensors, and environmental noise and illumination acquisition units at the physical hardware level, in order to capture high-dimensional heterogeneous behavioral, physiological and environmental signals in the human-computer interaction scene in all aspects.

[0154] The task event and time synchronization block is responsible for generating specific event codes in real time for each time the robot is called, limb pointing, audio-visual stimulus presentation, task stage switching, preference restrictions, movement trajectory planning and control prompts; and uses a unified system master clock to control the concurrently acquired video frames, audio frames, physiological sequences, distance sequences and robot logs to perform millisecond-level hard alignment.

[0155] The multimodal feature extraction module is specifically designed to extract facial motion units, gaze direction, head posture, body skeletal points, upper limb activity, repetitive motion index, and relative physical distance indicators between the child and the robot (i.e., child-robot distance) from videos; simultaneously, it extracts volume, fundamental frequency, speech rate, probability of screaming and crying, and semantic keywords for help and rejection from audio; extracts heart rate, heart rate variability indicators, peak skin conductance response (i.e., heart rate HR, heart rate variability HRV, peak skin conductance EDA), respiratory rate, and body movement intensity indicators (i.e., body movement intensity) from physiological data; and extracts temporal indicator features such as the number of prompts, the number of failures, and the intensity of stimuli from the robot's own task logs.

[0156] The temporal multimodal risk perception module employs cross-modal attention Transformer, temporal convolutional network (TCN), bidirectional long short-term memory network (LSTM), graph attention network (GAT), and uncertainty estimation algorithms for joint inference, outputting the probability, confidence score, and key contributing modalities that trigger different risk types within multiple future time windows.

[0157] The robot safety decision module closely follows the risk type, probability of occurrence, confidence level, and individualized dynamic risk threshold of the child as output above. Through the hardware driver layer, it adaptively issues task degradation, controls pause, reduces volume and brightness, controls the mobile chassis to retreat to a safe distance, controls adaptive path planning to avoid approaching the exit, uses screen image prompts, calls voice to play soothing statements, actively invites the child to express their needs for pause, help, or discomfort, and sends remote alarm notifications to external caregivers or therapists when necessary.

[0158] The personalized rehabilitation database module, with the explicit consent of the guardian, securely and in compliance with privacy protection standards for sensitive personal information and children's behavioral data, stores the child's natural state baseline data, task performance trajectory, specific risk warning chains, physical intervention responses, and expert review and annotation information. The module employs a localized encrypted storage architecture in its hardware deployment and incorporates a time-series data desensitization unit at the bottom of the data processing chain. This unit performs discrete desensitization and anonymization processing on concurrently collected facial biometrics, voice, and physiological sensitive indicators of the child. Furthermore, the module includes a retention period management module to monitor the retention time of behavioral data in real time. After exceeding a preset long-term rehabilitation monitoring period, a hardware erasure mechanism is triggered to physically destroy or offline anonymize and archive the original high-dimensional multimodal data. An online incremental learning mechanism continuously feeds back and updates the child's current individual risk threshold and the robot's specific task execution difficulty, achieving closed-loop adaptive iteration at the entire software and hardware level.

[0159] The complete control execution flow of this device is described in detail below with reference to specific embodiments and system process mechanisms:

[0160] When the system is first used or when interacting with newly admitted children with special needs, the rehabilitation therapist or on-site caregiver first enters the child's basic individual profile through the device's visual interface. The entered information includes the child's age, diagnosis, past problematic behaviors, sensory sensitivities, preferences, contraindicated stimuli, common expressions, and safety boundaries. Before the formal robot interaction task begins, the system controls the companion robot to perform a 3-5 minute resting companionship task with the child with special needs. During this resting companionship, the multimodal data acquisition module maintains a non-interventional working state, comprehensively collecting the child's basic facial expressions, spontaneous repetitive movements, heart rate / HRV, EDA, breathing, and natural speech in this stable state, thus forming an individual baseline. Simultaneously, the main control system adaptively sets an initial safe distance based on the child's height characteristics, limb mobility, and care guidelines. For example, in a desktop interaction scenario, the preset safe distance is no less than 0.8 meters; in a mobile following scenario, the preset initial safe distance is no less than 1.2 meters. These initial safe distance values ​​can be manually adjusted by the rehabilitation therapist according to actual safety monitoring needs.

[0161] like Figure 3 As shown, the mapping process planning architecture between robot interaction tasks and multimodal acquisition indicators in this embodiment is illustrated in detail. To avoid generalized acquisition and ensure that each extracted feature indicator has a clear causal explanation, this invention completely and organically decomposes and embeds screening assessment and real-time safety monitoring into specific robot interaction tasks and daily companionship processes that correspond to clear inducing purposes, acquisition indicators, and safety risk judgments. Specifically, the specific interaction tasks include one or more of the following: T1 calling name - joint attention, T2 sensory comfort expression, T3 alternating cooperation and imitation, T4 delayed gratification and request for help, T5 safety boundary and spatial transfer, and T6 natural companionship free play.

[0162] like Figure 3 As shown, when the system drives the rehabilitation companion robot to perform the T1 name-calling-shared attention task, the robot calls the child's name, points to the target object, and observes whether the child turns their head / looks / shares attention. Specifically, the robot calls the child's name from a distance of 1.2-1.8 meters, then points to the target toy or screen image, and gives the prompt "Look here / Let me see." During this task, the multimodal feature extraction module focuses on collecting and calculating specific indicators such as name-calling response latency, head turning angular velocity, first eye contact point, duration of shared gaze, number of times the target object is gazed at, and whether there is a verbal / gestural response. In terms of control logic, if the child exhibits specific characteristics such as persistent unresponsiveness, increased eye avoidance, and enhanced stereotyped upper limb movements during the interaction, the system automatically identifies and locks these as signs of the child's current social withdrawal or psychological risk caused by the pressure of a high-difficulty task.

[0163] like Figure 3 As shown, when the system drives the robot to perform the T2 sensory comfort expression task, the robot presents controllable audio-visual / tactile stimulation, demonstrating expressions such as "like / uncomfortable / pause". Specifically, the robot progressively presents low-intensity audio-visual or tactile simulations, actively demonstrating the expression "I think it's a bit loud, can I pause?", and allowing the child to choose to continue / be quieter / pause. The system simultaneously and in parallel collects data on whether the special needs child exhibits avoidance postures such as frowning / squinting / covering ears / stepping back, as well as specific indicators such as facial action unit (AU) intensity, heart rate changes, peak skin conductance (EDA), respiratory rate, probability of screaming / crying, and changes in distance from the stimulus source. This task is specifically designed to accurately and sensitively quantify the potential sensory overload, emotional breakdown, or sudden self-harm risk in children. Once the calculated combination of risk characteristics triggers an alert, the robot immediately and automatically reduces stimulation and prompts the child to express discomfort.

[0164] like Figure 3 As shown, when the system controls the robot to perform the T3 alternating collaboration and imitation task, specifically referring to building blocks / facial expression imitation / action imitation, it induces alternating waiting and social responses. The system specifically controls the robot and the child to take turns building blocks, passing a ball, or imitating actions, and during the interaction, the robot's main controller deliberately sets alternating waiting times and rule changes. The system comprehensively records and extracts specific indicators such as the child's waiting time, number of attempts to grab, imitation accuracy, smoothness of movement trajectory, frequency of repeated actions, recovery time after task errors, and verbal complaints / rejection words when facing rule changes. These features are used to accurately capture the escalation of frustration, potential aggressive behavior, or task avoidance risks experienced by children with special needs due to scene changes, rule complexity, or long waiting times, driving the robot's safety decision module to correspondingly reduce rule complexity or adaptively shorten waiting time.

[0165] like Figure 3 As shown, when the system controls the robot to perform the T4 delayed gratification and help-seeking task, it specifically refers to inducing requests for help and frustration adjustment when a child's preferred item is temporarily unavailable or a part of the task is missing. In actual interaction, the robot is specifically controlled to briefly place the preferred item into a transparent locking box, or to deliberately hide or omit a key assembly component in the current assembly interaction tool, thereby artificially creating a structured delayed gratification intervention environment to guide the child to actively seek help from the robot through voice, pictures, or gestures. The system simultaneously collects and evaluates specific indicators such as the child's help-seeking initiation probability, help-seeking latency, volume increase, grasping / slapping of the box, leaving the seat, crying, heart rate increase, and task abandonment rate. This task can dynamically assess the risk of emotional breakdown or destructive serious problem behaviors in children with special needs due to their inability to verbally express their higher-level core needs, prompting the robot's decision-making system to promptly present alternative expression boards and simultaneously send alarms to external caregivers through the network module.

[0166] like Figure 3 As shown, when the system controls the robot to perform the T5 safety boundary and spatial transfer task, it specifically involves the robot leading the child with special needs to a designated area, while comprehensively monitoring the risks of escape, collision, and close contact along the route. Specifically, the robot guides the child from the play area to the training mat or an area in the opposite direction of the doorway, deliberately setting safety boundaries and obstacles along the route. During this transition interaction phase, the multimodal feature extraction module continuously collects and monitors specific indicators such as the child-robot distance, speed approaching the doorway, number of boundary crossings, collision prediction distance, following rate, sudden acceleration, falls / balance, and the caregiver's distance. This step is specifically designed to prevent physical safety risks during the interactive transition, such as rushing towards the doorway to escape, falls, collisions with table corners, or excessively close proximity to the robot, driving the robot to automatically employ a risk-aware path planning algorithm and maintain a safe physical distance.

[0167] like Figure 3 As shown, when the system-controlled robot performs the T6 natural companionship free play task, it transitions to a low-initiative intelligent guardianship and companionship state, continuously accompanying the child in free play for 10-30 minutes at a time. During this companionship training phase, the robot maintains low-intervention companionship, continuously collecting spontaneous behaviors and risk warning chains, and only responding with low-frequency interactions when the child actively approaches, calls out, or when the risk increases. In this state, the system focuses on collecting specific indicators such as spontaneous social initiation, repetitive behavior baselines, emotional fluctuations, fatigue signals, environmental noise, natural speech, and abnormal approach / departure trajectories. The long-term channel data collected from this unstructured free task is directly and individually stored in the system backend, specifically used to form an individual baseline for that particular child in a natural state, minimizing the frequent false alarms and serious false negatives caused by the pressure of a single clinic assessment environment.

[0168] like Figure 3 As shown, when the above six standardized and non-standardized rehabilitation tasks are performed alternately, the feature extraction module uniformly and systematically outputs a multimodal feature index matrix to the subsequent model layer. This multimodal feature index matrix... Figure 3 The following data includes reaction latency, number of eye shifts and gaze duration, facial AU intensity and emotion transition, upper limb activity and stereotyped movement index, speech prosody / semantics / help words, HR / HRV / EDA / breathing, child-robot distance, boundary approach speed, and task completion rate and abandonment rate. This ensures that the entire process control decision has rigorous underlying feature support and consistency in full text literal reference.

[0169] Furthermore, considering the system architecture, the specific acquisition metrics corresponding to each modality are defined within the system as follows:

[0170] The visual behavior indicators specifically include face detection bounding box, 68 / 468 facial key points, FACS motion unit AU01 / AU04 / AU12 / AU25 intensity, blink frequency, gaze direction, co-attention switching, head posture, body skeletal points, hand key points, upper limb root mean square velocity, repetitive patting / shaking frequency, and postures of getting up from a seat and falling.

[0171] The specific speech and language metrics include sound pressure level, fundamental frequency (FO), energy, speech rate, pauses, probability of screaming / crying, emotional acoustic features, ASR transcription, rejection words / help words / preference words, semantic completeness, and response delay after robot prompts.

[0172] The specific physiological indicators include heart rate (HR), heart rate variability (HRV) (RMSSD, SDNN), skin conductance (EDA) level and peak SCR, respiratory rate, body movement intensity, and skin temperature changes, which are used to objectively assess children's stress, internal psychological arousal, and sensory overload levels at the most basic level.

[0173] The specific spatial and safety indicators include child-robot distance, child-exit / hazard distance, relative speed, collision time (TTC), number of boundary crossings, robot chassis speed, robot arm end distance, and changes in the center of pressure of the floor mat.

[0174] The specific interaction and task metrics include task stage, number of prompts, stimulus intensity, waiting time, number of failures, presentation status of rewards / preferences, number of pauses, number of times the child initiates a task, task completion rate, and recovery time.

[0175] The environmental indicators specifically include noise levels (decibels), light intensity, number of people in the room, stranger entry incidents, temperature and humidity, and background music / sudden noise incidents, etc., to fully explain the antecedents of safety risks to children with special needs.

[0176] like Figure 4 As shown, the internal processing pipeline and computing unit deployment architecture of the temporal multimodal risk perception algorithm model in this embodiment (implemented in engineering as a task-guided temporal multimodal risk perception network, abbreviated as T-MRAN deep model) are illustrated in detail. The algorithm process is strictly implemented through the following progressively layered temporal data flow computation steps:

[0177] Step 1 involves data preprocessing and clock alignment. Responding to specific interactive tasks performed by the front-end robot, the edge computing unit divides the multiple concurrent signals acquired by the front-end into a sliding window sequence with a length of 5 seconds and a scrolling step of 0.5 seconds, using the timestamp of the robot's own interaction events as the main axis. The input data stream precisely contains, both literally and physically, [the necessary parameters]. Figure 4The four concurrent signals described are: first, video / depth / eye-tracking signals, specifically including facial key points, facial action units (AU), gaze direction, skeletal points, gait, and relative distance; second, speech / text signals, specifically including Mel-frequency cepstral coefficients (MFCC), fundamental frequency F0, signal energy, speech rate, keywords, and help / rejection semantics; third, physiological / environmental signals, specifically including heart rate (HR), heart rate variability (HRV), electrical conductance of skin (EDA), respiration, noise, illumination, and crowding; and fourth, robot context data output and recorded by the robot itself, specifically including task stage, number of prompts, stimulus intensity, and failure events. To eliminate motion artifact noise caused by frequent body movements of children with special needs or loose wearable sensors, the quality assessment and missing information compensation submodule in the feature extraction module dynamically calculates real-time quality scores for data streams experiencing head tilting occlusion or brief interruptions in physiological signals, and uses a masking mechanism to retain missing information in the current sliding window, achieving high-fidelity quality weight allocation and missing information compensation.

[0178] Step 2 is the single-modal feature encoding step. The sliced ​​temporal feature index fragments are synchronously input into a parallel multi-channel dedicated single-modal encoder, which is specifically included in... Figure 4 In the architecture shown, specifically, the single-modal encoder corresponding to the video / depth / eye-tracking signal input employs a facial expression recognition network and a spatiotemporal graph convolutional network (ST-GCN) or a video Transformer (corresponding to...). Figure 4 The ST-GCN / Video Transformer in the encoding process encodes key skeletal points and body movement gait features; speech, language, and text signals are input into a single-modal encoder, specifically using a Conformer or wav2vec2 model to extract high-dimensional emotional acoustic representations, and combining lightweight ASR to extract discrete text keywords and semantic embeddings for help and rejection; physiological and environmental signals are input into the single-modal encoder, specifically using a one-dimensional convolutional neural network combined with a bidirectional long short-term memory network (corresponding to...). Figure 4 The Conformer / 1D-CNN single-modal encoder form in the model more rigorously corresponds to the 1D-CNN (joint bidirectional long short-term memory network) encoding continuous one-dimensional heart rate, heart rate variability, skin conductance and respiratory deformation waveform features; the robot's own task log context data is sent to the single-modal encoder and directly represented by event embedding vectors, and then uniformly transformed into task context feature vectors.

[0179] Step 3 is the cross-modal quality-aware fusion step. The extracted single-modal encoding vectors from each path are input into the core layer of the cross-modal attention fusion network based on Cross-Modal Transformer. This fusion unit performs fusion calculations through a cross-modal attention mechanism, directly incorporating the real-time single-modal quality score output from the previous quality assessment as a bias term into the softmax distribution attention weight allocation and attention matrix calculation, performing fusion calculations with quality weights and missing value compensation.

[0180]

[0181] in, For the first Quality-perceived attention weights for each modality For query vector, For the first The key vectors corresponding to each mode and It is a linear transformation matrix. For vector dimensions, For the first Quality score of each modality of data. This is the adjustment coefficient for the quality score. Based on the calculated quality-perceived attention weights, the codes of each modality are fused to obtain a high-dimensional fused state vector representing the child's current interactive stress and arousal state. Through this formula constraint, when the quality score of a certain channel drops sharply due to a brief detachment of the wristband or a large angle of facial deflection, the cross-modal attention weights automatically decay and suppress the contribution weight of the missing modality, instead adaptively increasing the weights of other intact channels. This greatly ensures the robustness of the algorithm in complex, unconstrained rehabilitation environments.

[0182] Step 4 is the risk chain graph modeling step. To endow the entire deep model with advanced interpretability and causal tracking capabilities, the system constructs risk chain graph modeling at the algorithm layer, specifically building an individualized risk chain graph with antecedent-state-behavior-effect graph topology (specifically corresponding to...). Figure 4 The risk graph includes antecedent, state, behavior, and consequence diagrams. Nodes in this risk graph specifically include on-site triggering events such as increased noise, task failure, covering ears, backing away, increased heart rate, screaming, and rushing towards the door. The system employs either a Graph Attention Network (GAT) or a Graph Sampling Aggregation Network (GraphSAGE, specifically corresponding to...). Figure 4 The GAT / GraphSAGE algorithm in the database calculates the causal transfer contribution of each active node to future security risks and the causal chain graph node allocation, connecting scattered high-dimensional multimodal features into a causal precursor chain with high-level logical persuasiveness.

[0183] Step 5 involves the calculation of future risk time series prediction and comprehensive score. Based on the fusion output of the fused state vector and causal risk chain graph network, it utilizes a bidirectional long short-term memory network (BiLSTM), a temporal convolutional network (TCN), or a survival analysis module (Survival Head, specifically corresponding to...) Figure 4 The model, constructed using BiLSTM / TCN / Survival Head, predicts future risks by assigning multi-type risk labels to six specific risk categories: sensory overload, self-harm / aggression, escape, collision / fall, training interruption, and social withdrawal. The network processing layer can predict the probability of the output occurring within multiple preset time windows in parallel, specifically including the probability of occurrence within short time windows such as 30 seconds, 60 seconds, and 120 seconds (specifically corresponding to...). Figure 4 The probability of 30 / 60 / 120 seconds in the model (e.g., the probability of 30 / 60 / 120 seconds). Based on this, the system combines units including the child baseline and individualized calibration, utilizing the child baseline B_i and EMA for online updates (specifically corresponding to...). Figure 4 Individualized calibration is performed using the child baseline B_i and EMA online updates; and the confidence output of the current judgment result of the output model is obtained through the uncertainty estimation module, utilizing Monte Carlo Dropout sampling (MCDropout) or uncertainty entropy penalty (specifically corresponding to...). Figure 4 The MCDropout / entropy penalty and confidence output are used to output high-confidence results. Current time step. The final comprehensive risk score is calculated and output using the following weighted formula:

[0184]

[0185] in, For the current moment The comprehensive risk score, with risk levels including low / medium / high / emergency risk labels (specifically corresponding to...). Figure 4 (Low / medium / high / urgent and multiple risk labels). The total number of preset risk types, For the first Types of risks in the future time window The predicted probability of occurrence within the timeframe, For the first The weighting coefficients for risk categories are specifically determined based on the severity of the historical risky behaviors and the level of destructive hazard for that particular child individual. For the first Uncertainty in the time-series prediction of risk-like conditions The penalty weighting coefficient for uncertainty is used to calculate the value for subsequent implementation of intervention strategies (specifically including destimulation, pause, withdrawal, prompting for help, and alarm, respectively). Figure 4 The underlying quantitative basis for (reducing stimulation, pausing, distancing, prompting for help, and issuing alarms) in the process.

[0186] Step 6 involves individualized calibration and online threshold update. Addressing the industry pain point of significant heterogeneity in external behavior and internal physiological arousal during emotional escalation in autistic children, the system dynamically and adaptively adjusts the sensitivity of current risk decision-making warnings using the following formula:

[0187]

[0188] in, For the updated individualized risk thresholds for children, The initial global threshold preset by the system. The standard deviation is the range of the average fluctuation of the indicators of the child during a stable period in the past preset number of tasks. The success rate of intervention when the child faced robot safety intervention actions in the recent period. and The dynamic adjustment coefficient allows for online fine-tuning of the warning sensitivity to address behavioral differences among children with special needs, effectively preventing frequent false alarms from highly sensitive children or missed alarms from low-responsive children. This update mechanism utilizes negative feedback adjustment to adaptively match and output corresponding closed-loop physical control intervention strategies at the backend.

[0189] like Figure 5 The diagram illustrates in detail the risk perception safety closed-loop and robot control diagram of this invention. After obtaining multi-type risk scores output by the multimodal model, this closed-loop control mechanism directly translates them into corresponding hardware-level hierarchical intervention strategy control commands in the robot safety decision module, achieving online, fully closed-loop adaptive drive between robot task execution, child reaction collection, risk scoring, and safety strategies. The entire control closed loop revolves around a central area integrated control rule example, resulting in high-frequency online adaptive closed-loop control linkages.

[0190] The specific hardware hierarchical control decision-making control rules are designed as follows:

[0191] When the risk score output by the multimodal model triggers a low-risk label, that is, when the risk score output based on the multimodal model at the current time (specifically corresponding to...) Figure 5 When the comprehensive risk score calculated by the multimodal model output R_k(t+Δ) meets the threshold for recovery to a low-risk state, the system determines that the current site is in a safe and stable state, and the control rules (specifically corresponding to...) are applied. Figure 5(Example of control rules in the text) directly issues instructions to the robot to continue the task; specifically, it drives the front-end robot in the task execution state to normally issue prompts / move / demonstrate; in this state, the system continuously and non-invasively records, monitors and analyzes the child's reaction in real time. The child's reaction at the bottom layer of the data chain specifically includes the linkage changes in the state of action, voice, physiological changes, etc. collected by the multimodal peripheral sensor array. When the comprehensive risk score meets the medium risk range, that is, meets the threshold range for reducing difficulty, the safety decision control module triggers the corresponding difficulty reduction strategy, controls the robot to adaptively execute safety strategy actions to reduce the difficulty and stimulation intensity of the current interactive task, specifically including adjusting the speech synthesis speed and playback volume through the control circuit, reducing the screen display brightness, shortening the waiting time in the alternation collaboration, and providing picture cards for the child to make a choice, etc., as active destimulation actions.

[0192] When the comprehensive risk score reaches the high-risk range, i.e., when the pause and prompt expression range is met, the safety decision control module triggers the corresponding safety policy. The control robot immediately controls the currently executing audio and video script and motion components to pause the current interactive task and execute the pause and prompt expression control logic. Specifically, it drives the mobile chassis to move backward until it is backed up and maintained beyond the preset physical safety distance, opening up physical space to quickly eliminate the feeling of physical space oppression and achieve adaptive back-off control. At the same time, the control robot's expression screen switches to an intuitive help expression interface and plays a smooth and calming soothing voice to actively guide and prompt for help, inducing the special needs child to correctly express their current needs.

[0193] When the comprehensive risk score triggers the emergency risk label, that is, when the emergency stop label is met (stop movement and notify the caregiver), or when obvious self-harming actions are captured directly and with high sensitivity by external sensors, or when dangerous escape or collision actions such as a child accelerating violently towards the exit are detected, the system determines that the scene has entered an emergency risk state. It immediately triggers the highest priority emergency stop safety strategy, controls the mobile chassis and mechanical components to immediately stop physical movement and mechanical displacement, completely cuts off the physical threat source from the moving parts, activates a low-stimulation audible and visual alarm, and executes the stop movement and notify the caregiver without delay. It also sends a remote alarm notification to the tablet, mobile phone or institutional management system held by the caregiver or therapist, and automatically locks and records the multimodal raw data for 60 seconds before and after the event for subsequent expert review.

[0194] In an optional mobile robot control embodiment, when the rehabilitation robot is a wheeled mobile robot with a free-moving chassis, and the system controls the mobile chassis to perform path planning while following a caregiver or retreating to avoid danger, the algorithm further embeds the calculated comprehensive risk score and the spatial motion data extracted by the sensors directly as penalty terms into the path cost function, thereby driving the chassis to perform adaptive path planning with emotion and escape boundary perception capabilities.

[0195]

[0196] in, Candidate nodes in the movement path The overall cost of the comprehensive assessment The actual physical path cost from the starting point to the current candidate node. To heuristically estimate the physical path cost from the current candidate node to the target node, The estimated time for the robot to reach the candidate node according to the current planned path. The corresponding emotions and risk-avoidance value of children; Predicted collision costs for physical collisions with a child's body or interactive physical obstacles; This is the penalty incurred by the node for being too close to the room's main exit or a pre-defined dangerous physical area. The weighted system corresponds to each term. This path planning formula enables the robot's master path search algorithm to adaptively select an absolutely safe path that, while geometrically farther and slower, can physically detour and intercept the exit direction when the robot encounters emotional disturbances, increased internal psychological pressure, or approaches the physical boundary of the space. This effectively avoids the risk of accidental collisions and personal injury during escape at the level of controlling the physical movement of the chassis.

[0197] Furthermore, the risk perception safety solution of this invention not only relies on single, short-term structured task tests, but also establishes a long-term, personalized rehabilitation database module containing massive amounts of real-world interaction data through multi-channel data collection in natural states. In specific research implementation and dataset construction, the system can, under the premise of strictly obtaining written informed consent from guardians and ethical approval from the medical ethics committee, collect continuous, long-term interaction data between multiple children with autism spectrum disorder and rehabilitation robots in various real-world interactive scenarios such as intervention training rooms in rehabilitation institutions, resource classrooms in inclusive education schools, and family living rooms across regions. During data collection, the dataset construction for each child must fully cover all scenarios, including the aforementioned structured teaching tasks, natural free play companionship, waiting with rotation rules, transportation during transfers, spontaneous requests for preferred objects, presentation of progressively stronger sensory stimuli, and daily companionship. The entire process employs a three-pronged mechanism: the robot automatically generates coarse annotations from its own interactive context event flow; on-site caregivers quickly click on markers via a mobile app; and professional behavioral therapists conduct offline high-precision secondary verification to generate accurate risk warning labels. Furthermore, an active learning algorithm prioritizes extracting long-tail behavioral fragments with extremely high model classification uncertainty from the fused features and pushes them to experts for review. In natural companionship mode, the system continuously records the individualized historical behavioral baseline of the child with special needs, including their most realistic normal heart rate range, frequency of typical stereotyped repetitive movements, specific expression patterns, physiological tolerance thresholds for sounds, lights, or waiting, and the impact of different robot limb movements on the child. During model parameter training, the system first uses massive amounts of data from a spectral group of children to pre-train a network with generalized weights to give it basic generalization ability. Subsequently, the system uses an incremental fine-tuning algorithm to fine-tune the child's current specific historical data with a small number of samples. This ensures that the entire perception and monitoring system possesses both excellent out-of-the-box initial generalization performance and can perfectly and specifically adapt to the unique and hidden emotional breakdowns or self-harming outbursts of each autistic child.

[0198] To enable those skilled in the art to more clearly understand the closed-loop operation of this device, this embodiment describes in detail a safety monitoring process in a rehabilitation institution: After a child with special needs and autism spectrum disorder enters the institution's intervention training room, the system-controlled robot first performs a 2-minute free companionship task. During this silent state, the system automatically collects and records the child's natural heart rate, repetitive movements, and gaze distribution in a stable state, and automatically establishes an individual baseline in the individualized rehabilitation database module. Subsequently, the robot automatically switches to perform a name-calling joint attention task, calling the child's name at a physical distance of 1.5 meters. The child turns their head 2.1 seconds after the first name call, and does not turn their head after the second name call, but their gaze accurately looks at the target toy pointed to by the robot; the system background then quantifies and records the joint attention stability index as moderate.

[0199] Then, when the interactive flow adaptively switched to the sensory comfort expression task, the robot played a low-intensity sound. The child briefly frowned, but was then able to independently click the "continue" option on the screen. When the sound intensity increased, the multimodal data acquisition module sensitively captured and monitored a combination of multimodal features, including the child covering their ears, moving back 0.6 meters, a rapid increase in EDA (Electronic Distress Ability), and an increase in volume. Based on this high-dimensional multimodal slice feature matrix, the task-guided temporal multimodal risk perception network predicted a sensory overload risk of 0.72 within 60 seconds over several preset time windows. This comprehensive risk score directly triggered high-risk classification control. The robot control system immediately issued an interrupt command, forcibly cutting off the current high-volume hardware stimulus source, controlling the mobile chassis to actively move back 0.5 meters to quickly increase the sense of spatial oppression, and popping up a graphical interface for pausing, lowering the volume, and continuing to ask for help, playing a soothing voice to guide the child in expressing their needs. With clear prompts, the child selected "quiet," and the system detected that the child's overall risk score quickly dropped to 0.38. The robot then switched back to a low-stimulation training state, thus successfully ensuring that the entire rehabilitation and teaching task could continue safely and uninterrupted.

[0200] Following this, the interaction flow shifted to a delayed gratification task. The robot placed a preferred toy car into a transparent box and guided the child to say "Help me open it." The child failed to utter a proper help-seeking phrase on their first attempt, and exhibited increased heart rate and upper limb activity at a lower level. Based on this, the network model highly assessed the child's risk of emotional breakdown at 0.65, triggering tiered regulation. The control system then instructed the robot to shorten the waiting time and demonstrate "Help me," while simultaneously displaying a help-seeking card on the screen. With clear prompting, the child touched the help-seeking card, and the robot immediately activated its arm-end mechanical components to automatically open the box and reinforce the help-seeking behavior. After the interaction, the system automatically generated a structured, quantitative rehabilitation and safety report, including task completion rates, risk heatmaps, key contribution modalities, and robot adaptive intervention control records. The report clearly showed that the child was sensitive to changes in sound intensity during sensory tasks and required earlier help-seeking prompts during delayed gratification tasks. Based on this, the rehabilitation therapist planned the next stage of the course: reducing the escalation rate of sound stimulation and increasing training in "pause / help" expressions, achieving long-term, highly unified daily rehabilitation training and safety monitoring at the most fundamental level.

[0201] Through the complete hardware and software module design, multi-dimensional feature perception based on specific tasks, depth map attention model derivation, and adaptive execution control that fully corresponds to the accompanying drawings, this monitoring method and device can achieve the following significant, direct, and reproducible technical effects in actual deployment:

[0202] The special needs children's rehabilitation robot has been successfully transformed from a traditional static tool that could only execute fixed training scripts and present test tasks into an adaptive closed-loop drive system that can identify the child's risk status and safely adjust the task difficulty and the intensity of audio-visual stimulation in a closed loop during continuous human-computer interaction.

[0203] By fully embedding multimodal data collection into specific structured tasks with clear interactive causal meaning and behavioral induction logic, comprehensive quantification of key high-level indicators such as reaction latency, gaze direction avoidance, co-attention switching rate, deep autonomic psychological stress, and spatial boundary collision risk has been achieved, greatly improving the objectivity, reproducibility, and clinical logical interpretability of behavioral assessment and safety warning.

[0204] An online update mechanism for individualized dynamic thresholds based on adaptive negative feedback adjustment for children was introduced, and individual baseline accumulation of long-term data from both structured paradigm tasks and natural free play companionship was perfectly integrated. This enabled the monitoring device to adaptively adjust the sensitivity of the early warning system according to the behavioral heterogeneity of different children, completely overcoming the defects of frequent false positives or serious missed alarms caused by the traditional unified hard-coded threshold scheme, significantly reducing alarm fatigue of caregivers, and improving the robustness of the system in multiple environments.

[0205] The algorithm risk score output by the high-dimensional neural network model at the edge fusion computing layer is seamlessly and without delay transformed into graded intervention action strategies such as chassis movement deceleration and retreat to avoid danger, adaptive unloading of physical stimuli, multimodal chart help-seeking behavior induction, and complete emergency stop of moving parts at the hardware level. This constructs a high-priority physical control closed loop at the lowest level of robot physical control, which significantly reduces the probability of destructive behavior occurring and escalating. At the physical level, it effectively avoids the risk of serious collisions, falls, and injuries between special children, on-site caregivers, and equipment, and achieves a high degree of unity between the continuity of special education training and the safety of closed care.

[0206] By utilizing the visualized individualized risk chain diagram reports accumulated from long-term databases, doctors, therapists, and parents can clearly and intuitively present the most underlying and hidden technical or interactive antecedent events that trigger emotional fluctuations in a specific child, as well as the most effective robot adaptive adjustment and intervention control methods. This provides scientific and highly reliable objective data support for the precise adjustment, optimization, and evolution of personalized rehabilitation intervention courses and behavioral training programs.

[0207] This embodiment also provides a system, specifically: a multimodal risk perception and safety monitoring system for children with autism based on robot interaction tasks, including a rehabilitation companion robot body, a peripheral sensor array and an edge fusion computing chip;

[0208] The rehabilitation companion robot body includes at least a movable chassis, multi-degree-of-freedom mechanical parts, and audio-visual interaction hardware devices for human-computer dialogue;

[0209] The peripheral sensor array includes a multi-channel concurrent physical sensor device for acquiring visual, speech, physiological, spatial physical motion, and environmental indicators;

[0210] The edge fusion computing chip is electrically connected to the peripheral sensor array and the rehabilitation companion robot body to store and run the method described above in order to drive the movable chassis and the multi-degree-of-freedom mechanical components to adaptively perform defensive and risk-avoidance actions.

[0211] The algorithm model, sensor type, number of tasks, risk threshold and intervention strategy described in the above embodiments can be replaced or adjusted according to specific equipment conditions, child's age, rehabilitation goals and rehabilitation training monitoring standards. As long as they still realize the technical concept of "multimodal risk perception, future risk prediction and safety closed-loop intervention in robot interaction tasks", they should all fall within the protection scope of this invention.

[0212] The algorithm model, sensor type, number of tasks, risk threshold and intervention strategy described in the above embodiments can be replaced or adjusted according to specific equipment conditions, child's age, rehabilitation goals and rehabilitation training monitoring standards. As long as they still realize the technical concept of "multimodal risk perception, future risk prediction and safety closed-loop intervention in robot interaction tasks", they should all fall within the protection scope of this invention.

Claims

1. A method for multimodal risk perception and safety monitoring of autistic children based on robot interaction tasks, characterized in that, Includes the following steps: In response to a specific interactive task performed by the robot, multimodal raw data of the child is collected synchronously. The multimodal raw data includes at least two or more of the following: visual behavior data, speech expression data, physiological state data, spatial motion data, tactile or force feedback data, environmental stimulus data, and the robot's own task log data. The robot includes at least a mobile chassis and mechanical components. Using the robot's own event timestamps as the main axis, the multimodal raw data is clock-aligned and sliced, and the temporal feature indicators corresponding to each modality are extracted respectively. The extracted temporal feature indicators are input into a task-guided temporal multimodal risk perception network, and quality perception fusion is performed through a cross-modal attention mechanism to obtain a fused state vector. Based on the fused state vector, time series prediction is performed, and the risk probability and time series prediction uncertainty of multiple types of risks within multiple preset time windows in the future are output. A comprehensive risk score is calculated based on the risk probability and the time series prediction uncertainty. Control decisions are made based on the comprehensive risk score and the child's individualized dynamic risk threshold, driving the robot to adaptively execute corresponding multi-level safety intervention actions.

2. The method for multimodal risk perception and safety monitoring of autistic children based on robot interaction tasks according to claim 1, characterized in that, Clock-aligned slicing of the original multimodal data includes: The multimodal raw data is divided into sliding windows of preset length and scrolling according to preset step size; The quality of the missing modal data is scored, and a masking mechanism is used to preserve the missing information in the sliding window.

3. The method for multimodal risk perception and safety monitoring of autistic children based on robot interaction tasks according to claim 1, characterized in that, Before inputting the task-guided temporal multimodal risk perception network, the method further includes unimodal encoding of the temporal feature indicators corresponding to each modality: The skeletal and motion features in the visual behavior indicators are encoded using a facial expression recognition network and a spatiotemporal graph convolutional network or a video Transformer. A speech feature extraction model based on a self-attention mechanism is used to extract the emotional acoustic representation in the speech and language indicators, and keywords are extracted by combining speech recognition. The physiological indicators were encoded using a one-dimensional convolutional neural network and a bidirectional long short-term memory network. The task context corresponding to the robot's own task log data is represented by an event embedding vector.

4. The method for multimodal risk perception and safety monitoring of autistic children based on robot interaction tasks according to claim 1, characterized in that, Quality perception fusion is performed through a cross-modal attention mechanism to obtain a fused state vector, the calculation formula of which is as follows: ; in, For the first Quality-perceived attention weights for each modality For query vector, For the first The key vectors corresponding to each mode and It is a linear transformation matrix. For vector dimensions, For the first Quality score of each modality of data. This is the adjustment factor for the quality score; The modal codes are fused according to the quality-aware attention weights to obtain the fused state vector.

5. The method for multimodal risk perception and safety monitoring of autistic children based on robot interaction tasks according to claim 1, characterized in that, Before outputting the risk probability, it also includes: Construct an individualized risk map with a structure of antecedent event A, child state S, aura behavior P, risky behavior B, and intervention outcome C. The nodes of the individualized risk map include increased noise, task failure, covering ears, backing away, increased heart rate, screaming, and rushing to the door. A graph attention network is used to calculate the contribution of each node to future risk.

6. The method for multimodal risk perception and safety monitoring of autistic children based on robot interaction tasks according to claim 1, characterized in that, The various types of risks include sensory overload, self-harm or aggression, escape, collision or fall, training interruption, and social withdrawal; the multiple preset time windows in the future include at least 30 seconds, 60 seconds, and 120 seconds in the future; The formula for calculating the comprehensive risk score is as follows: ; in, For the current moment The overall risk score, This represents the total number of risk types. For the first Types of risks in the future time window The predicted probability of occurrence within, For the first Weighting coefficients for risk classes For the first Uncertainty in the time-series prediction of risk-like conditions The penalty weighting coefficient is for uncertainty.

7. The method for multimodal risk perception and safety monitoring of autistic children based on robot interaction tasks according to claim 1, characterized in that, The individualized dynamic risk threshold for children is updated online using the following formula: ; in, The updated individualized risk thresholds for children. The preset initial global threshold, This represents the average fluctuation range of the child's stable period across a preset number of tasks. The success rate of the most recent robotic intervention for the children. and This is the dynamic adjustment coefficient.

8. The method for multimodal risk perception and safety monitoring of autistic children based on robot interaction tasks according to claim 1, characterized in that, The process of driving the robot to adaptively execute corresponding multi-level safety intervention actions includes: When the overall risk score is less than a first preset threshold, the robot is controlled to continue the current task; When the overall risk score is greater than or equal to a first preset threshold and less than a second preset threshold, the robot is controlled to reduce the difficulty and intensity of the current task. When the comprehensive risk score is greater than or equal to the second preset threshold and less than the third preset threshold, the robot is controlled to pause the current task, the mobile chassis is controlled to retreat to a preset safe distance, and the robot is controlled to display prompt information to guide the child to express their needs. When the overall risk score is greater than or equal to the third preset threshold, or when self-harm or escape actions are detected, the mobile chassis and mechanical components are controlled to immediately stop moving and an alarm notification is sent to an external terminal.

9. The method for multimodal risk perception and safety monitoring of autistic children based on robot interaction tasks according to claim 8, characterized in that, When controlling the robot to reduce the difficulty of the current task or controlling the back distance of the mobile chassis, the comprehensive risk score and spatial motion data are embedded into a path cost function for path planning. The path cost function is: ; in, Candidate nodes in the path The overall cost of the comprehensive assessment The actual cost of the movement path. For heuristic estimation of path cost, The emotional and escape risk costs for the child when the robot reaches the candidate node. To predict the cost of a collision, The penalty for being near an exit or a pre-designated danger zone, , , These are the weighting coefficients for each item.

10. A multimodal risk perception and safety monitoring device for autistic children based on robot interaction tasks, characterized in that, The apparatus is used to implement the method according to any one of claims 1 to 9, and the apparatus comprises: The data acquisition module is used to synchronously acquire multimodal raw data of children in response to specific interactive tasks performed by the robot. The multimodal raw data includes at least two or more of the following: visual behavior data, speech expression data, physiological state data, spatial movement data, tactile or force feedback data, environmental stimulus data, and the robot's own task log data. The robot includes at least a mobile chassis and mechanical components. The feature extraction module is used to perform clock-aligned slicing of the multimodal raw data with the robot's own event timestamp as the main axis, and extract the time-series feature indicators corresponding to each modality respectively. The risk perception and prediction module is used to input the extracted temporal feature indicators into a task-guided temporal multimodal risk perception network, perform quality perception fusion through a cross-modal attention mechanism to obtain a fusion state vector, perform temporal prediction based on the fusion state vector, output the risk probability and temporal prediction uncertainty of multiple types of risks in multiple preset time windows in the future, and calculate a comprehensive risk score based on the risk probability and the temporal prediction uncertainty. The safety decision control module is used to make control decisions based on the comprehensive risk score and the child's individualized dynamic risk threshold, and drive the robot to adaptively execute corresponding multi-level safety intervention actions.