An edge device end-side child mental health assessment intervention full-closed loop implementation method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-30
- Publication Date
- 2026-08-11
AI Technical Summary
现有系统依赖云端算力与云端数据存储,测评、干预、数据分析全流程需要联网运行,在基层、农村等网络条件薄弱的场景无法落地使用;同时儿童敏感隐私数据需要上传云端处理与存储,存在数据泄露、滥用的重大风险
1.该边缘设备端侧儿童心理健康测评干预全闭环实现方法,通过将双量表AI测评、个性化干预方案生成、干预效果评估三大核心模型量化压缩后部署于ARM边缘智慧屏,实现儿童心理健康测评、干预、效果迭代、复测的全流程端侧离线运行,彻底摆脱云端算力与网络依赖,既解决了基层、农村等网络薄弱场景的落地难题,又通过AES-256加密本地存储与四级权限管理,从根源规避儿童隐私数据泄露风险,打破传统测训脱节的行业痛点,凭借轻量化模型适配与无需专业人员操作的特性,实现家庭、学校、社区场景的普惠化落地,一套系统覆盖健康儿童与孤独症行为风险儿童的差异化需求,具备极强的实用性与社会价值。
Smart Images

Figure CN122552167A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of children's mental health assessment and intervention technology, specifically to a closed-loop implementation method for children's mental health assessment and intervention on the edge device side. Background Technology
[0002] Current child mental health assessment and developmental intervention services have gradually shifted from offline, manual models to digital models. However, existing digital systems and technological solutions have the following shortcomings: The existing system relies on cloud computing power and cloud data storage. The entire process of assessment, intervention, and data analysis requires network connectivity, making it unusable in scenarios with weak network conditions, such as at the grassroots level and in rural areas. At the same time, children's sensitive privacy data needs to be uploaded to the cloud for processing and storage, posing a significant risk of data leakage and misuse. Summary of the Invention
[0003] To address the shortcomings of existing technologies, this invention provides a closed-loop implementation method for edge device-based children's mental health assessment and intervention, thus solving the problems mentioned in the background section.
[0004] To achieve the above objectives, the present invention provides the following technical solution: a closed-loop method for edge device-based children's mental health assessment and intervention, comprising the following steps: S1. Edge AI Model Deployment: The trained dual-scale AI assessment model for child development, personalized intervention plan generation model, and intervention effect evaluation model are quantified and compressed, and then deployed on the NPU or CUDA computing unit of the ARM edge smart screen. All AI models support real-time inference on the edge and do not require cloud computing power support. S2, Dual-Scale AI Assessment on the Edge: Through the motion sensing and audio acquisition devices of the edge smart screen, multi-dimensional data is collected during the assessment of children on the edge. The dual-scale AI assessment model deployed on the edge calculates and outputs the assessment results and risk levels of the Gesell Developmental Scale and the CARS Childhood Autism Behavior Screening Scale in real time. S3. Personalized intervention plan generation on the edge: Based on the edge assessment results, the personalized plan generation model deployed on the edge is used to automatically generate a personalized intervention plan on the edge by inputting the child's age, developmental characteristics, and risk level. This includes a list of UE5 motion-sensing courses for the corresponding training dimensions, training frequency, difficulty gradient, and adaptation mode. S4. Edge-side intervention training and data collection: During the process of children completing the UE5 somatosensory intervention course on the edge smart screen according to the intervention plan, the edge device collects all-dimensional training data of children in real time, including task completion rate, action achievement rate, attention duration, emotional changes, and social interaction performance, and stores the data locally throughout the process. S5. End-to-end intervention effect evaluation and program iteration: Through the intervention effect evaluation model deployed on the end-to-end, based on the collected training data, the intervention effect is quantitatively evaluated, the improvement of children's abilities and remaining shortcomings are identified, and the intervention program is automatically optimized and adjusted on the end-to-end to achieve dynamic iteration of the intervention program. S6. End-to-End Retesting and Management: Based on the child's intervention cycle and risk level, the end-to-end automatically pushes retesting tasks, completes the retesting through gamified assessment tasks, updates the child's developmental score and risk level, and regenerates the intervention plan based on the retest results, forming a closed loop on the end-to-end throughout the entire process; at the same time, the end-to-end completes the encrypted storage of all process data, the generation of growth reports, and the recording of compliance audit logs.
[0005] Furthermore, in step S2, the multi-dimensional data collection includes the full-dimensional collection of limb movement data, language feature data, and visual behavior data. All data collection, feature extraction, and evaluation calculation processes are completed on the device side without cloud data transmission.
[0006] Furthermore, in step S4, the UE5 motion-sensing intervention course is a lightweight 3D motion-sensing course adapted to edge device operation. The course content corresponds one-to-one with the evaluation indicators, realizing targeted training for the weakest dimension.
[0007] Furthermore, in step S5, the intervention effect evaluation model quantifies the ability improvement rate of each training dimension of the child in real time. When the ability improvement rate of the child is lower than the preset threshold for two consecutive training cycles, the difficulty, adaptation mode and training frequency of the intervention course are automatically adjusted.
[0008] Furthermore, in step S6, the retesting cycle is automatically set according to the risk level, specifically: the retesting cycle is 3 months for low-risk children, 1 month for medium-risk children, and 2 weeks for high-risk children.
[0009] Furthermore, in step S6, all data in the process is stored locally using AES-256 encryption, and a four-level hierarchical access control system is set up for parents, teachers, intervention specialists, and administrators. Unalterable compliance audit logs are generated, and the storage period is ≥3 years.
[0010] Furthermore, the multi-dimensional data acquisition in step S2 also includes non-contact end-to-end extraction of physiological signals using a binocular camera, specifically including the following sub-steps: S2-1, Binocular video stream preprocessing: The left-channel RGB video frame and the right-channel depth video frame are simultaneously acquired by a 120FPS global shutter binocular camera. The device side locates the optimal ROI region of the child's forehead and cheeks in real time, and calculates the three-dimensional spatial coordinates of the ROI region based on the depth frame. S2-2, Binocular Depth Motion Compensation: Based on binocular depth frames, the translation, rotation, and jitter amplitude of the child's face and head are calculated in real time. Spatial alignment compensation is performed on RGB video frames to eliminate motion interference and obtain a clean ROI image sequence. S2-3, End-side rPPG signal separation and noise reduction: RGB three-channel light intensity change time-series signal is extracted pixel by pixel in the ROI region, blood flow-related rPPG raw signal is extracted by ICA blind source separation, and heart rate and respiration corresponding filtered signals are obtained by wavelet filtering and adaptive bandpass filtering, respectively. S2-4, Physiological index calculation: Perform FFT frequency domain transformation on the heart rate filter signal to calculate the instantaneous heart rate; extract the amplitude modulation wave from the respiratory filter signal to calculate the respiratory rate; S2-5. Emotional arousal and stress index quantification: Integrating heart rate variability, respiratory fluctuations, facial micro-movements, and limb tension characteristics, the terminal calculates emotional arousal from 0 to 100; Based on heart rate abnormalities, respiratory disturbances, emotional arousal thresholds, and behavioral stress characteristics, the terminal outputs a stress index of 0 to 5 levels. The entire process of extracting, calculating, and quantifying the aforementioned physiological signals is completed on the device side, without cloud data transmission or cloud computing.
[0011] Furthermore, in steps S2-3, the frequency bands of the adaptive bandpass filter are set as follows: heart rate passband 0.8–3Hz, respiratory passband 0.15–0.5Hz.
[0012] This invention provides a closed-loop method for edge device-based children's mental health assessment and intervention, which has the following beneficial effects: 1. This edge device-based closed-loop implementation method for children's mental health assessment and intervention quantifies and compresses three core models—dual-scale AI assessment, personalized intervention plan generation, and intervention effect evaluation—and deploys them on an ARM edge smart screen. This enables offline operation of the entire process—from assessment and intervention to effect iteration and retesting—completely eliminating reliance on cloud computing power and networks. It solves the implementation challenges in grassroots and rural areas with weak networks, and avoids the risk of children's privacy data leakage through AES-256 encrypted local storage and four-level access control. It overcomes the industry pain point of traditional testing and training being disconnected. With its lightweight model adaptability and lack of professional personnel operation, it achieves universal implementation in families, schools, and communities. A single system covers the differentiated needs of healthy children and children at risk of autism, possessing strong practicality and social value.
[0013] 2. This edge device-based method for implementing a closed-loop intervention for children's mental health assessment introduces non-contact physiological signal extraction technology using binocular cameras. Through algorithms such as rPPG blind source separation, binocular depth motion compensation, and cross-modal feature fusion, it achieves precise edge-side quantification of heart rate, respiratory rate, emotional arousal, and stress index. This not only compensates for the insufficient accuracy of traditional single-behavioral modality assessments and reduces the scoring error of dual scales, but also eliminates the need for wearable devices, significantly improving the usage compliance of young children and children with special needs. At the same time, the deep integration of physiological indicators and behavioral characteristics allows personalized intervention plans to accurately adapt to children's real-time emotional and stress states, effectively solving problems such as emotional breakdown and low compliance during training, thus improving the intervention effect. Furthermore, all physiological signal extraction, calculation, and fusion processes are completed offline on the edge, without compromising the privacy compliance and hardware compatibility of the original closed loop, further strengthening the technical barriers and practical value of the invention. Attached Figure Description
[0014] Figure 1 This is a schematic diagram of the steps in the closed-loop implementation method for edge device-side children's mental health assessment and intervention according to the present invention. Detailed Implementation
[0015] The embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and should not be construed as limiting the scope of the invention.
[0016] like Figure 1 As shown, the present invention provides a technical solution: a closed-loop implementation method for edge device-based children's mental health assessment and intervention, the specific steps of which are as follows: S1, Edge AI Model Deployment The trained AI assessment model for dual-scale child development, personalized intervention plan generation model, intervention effect evaluation model, physiological feature extraction model, and physiological-behavioral cross-modal fusion model are INT8 quantized and compressed, and deployed on the NPU or CUDA computing unit of the ARM edge smart screen; the ARM edge smart screen is equipped with an RK3588 or Jetson Orin Nano main control chip and a 120FPS global shutter dual-lens camera; After quantization and compression, the physiological feature extraction model has a single-frame inference time of ≤60ms and a peak memory usage of ≤70MB; the cross-modal fusion model has an inference time of ≤40ms. All models support real-time offline inference on the device side without the need for cloud computing power. S2, End-to-End Dual-Scale AI Assessment Children complete gamified assessment tasks themed around the Octonauts IP adventure on a motion-sensing smart screen. The end-device device completes the entire assessment process through three modules, without any human intervention or cloud data transmission. Data acquisition module: Children's limb movement data are collected by a 120FPS global shutter binocular camera, children's language feature data are collected by a 4-array microphone, and children's visual behavior data are collected by an edge AI vision model. All data is processed in the device's local memory and is not written to external storage or uploaded to the cloud. It also uses the binocular camera and audio acquisition device of the edge smart screen to simultaneously complete the collection of behavioral data and non-contact physiological signals on the edge. After cross-modal fusion, the data is input into the dual-scale AI assessment model and outputs the assessment results, risk level, emotional arousal level, and stress index. Specifically, it includes the following sub-steps: S2-1, Binocular video stream preprocessing: The binocular camera synchronously outputs the left-channel RGB video frame and the right-channel depth video frame. The edge-side lightweight face detection model locates the optimal ROI region for the physiological signals of the forehead and cheeks, and calculates the three-dimensional spatial coordinates of the ROI region based on the depth frame. The specific left-side RGB frame sequence is set as follows: The right-path depth frame sequence is: in, For timestamps, Total number of frames, frame rate ; The edge-side lightweight face detection model uses a quantized version of the MTCNN model, and the edge-side inference outputs the face bounding box. The coordinates of the top-left pixel. Width; For height; The optimal ROI region rules are as follows: Forehead ROI: Cheek ROI: Total ROI after merger: The 3D spatial coordinates of the ROI are calculated as follows: Based on stereo depth frames ( (pixel coordinates), combined with the binocular intrinsic parameter matrix extrinsic parameter matrix Formula for calculating three-dimensional coordinates: in, This represents the pixel depth value. This is the intrinsic parameter matrix of the binocular camera (end-side pre-calibration). The ROI region is defined by its three-dimensional world coordinates. S2-2, Binocular Depth Motion Compensation: Based on binocular depth frames, the translation, rotation and jitter amplitude of the child's head are calculated in real time, spatial alignment compensation is performed on RGB video frames to eliminate motion interference and output a clean ROI image sequence; Head motion posture estimation uses the ICP iterative closest point algorithm to calculate the head rigidity transformation matrix. : in, For rotation matrices (corresponding to pitch, yaw, and roll); The translation vector (corresponding to) Directional displacement); The jitter amplitude quantification formula includes translation jitter amplitude: This indicates the total amplitude of the translational jitter; , , They represent along in three-dimensional space. Translational jitter component along the axial direction; This also includes the amplitude of rotational jitter: in, Rotation matrix The trace, i.e., the rotation matrix The sum of the elements on the main diagonal; This refers to the angle amplitude of the rotational jitter. RGB frame space alignment compensation for original pixels Perform an inverse transform to obtain the compensated pixels. : The compensated output sequence of ROIs with no motion interference is as follows: ;in, Indicates at time After depth motion compensation, the output is a sequence of regions of interest (ROI) images without motion interference. This sequence removes interference caused by device or scene motion, providing a stable image basis for subsequent operations such as physiological index calculation. Represents the moment The raw image data acquired by the left camera in the binocular camera system is one of the important input sources for motion compensation and subsequent processing. This is the motion parameter matrix, which contains various parameters describing the motion of the image, such as translation, rotation, and scaling. Through these parameters, The function can accurately target The transformation is performed to achieve the purpose of motion compensation; The function is an image warping operation function that transforms an image according to given motion parameters to achieve motion compensation. This function restores a moving image to a relatively stable state by rearranging the pixels in the original image according to a specific mapping relationship. S2-3, End-side rPPG signal separation and noise reduction: RGB three-channel light intensity change time-series signal is extracted pixel by pixel in the ROI region. Blood flow-related rPPG raw signal is extracted by ICA blind source separation. After wavelet filtering and adaptive bandpass filtering, heart rate filtered signal (0.8–3Hz) and respiratory filtered signal (0.15–0.5Hz) are obtained. The specific steps for extracting RGB three-channel light intensity signals are as follows: Time series of average light intensity in ROI region: in, For at any time The mean of the red channel signal extracted from the region of interest (ROI) of the left camera, similarly, and The green channel and blue channel are respectively located at the time. The mean signal extracted from the region of interest of the left camera; This represents the total number of pixels in the ROI. The left camera is located at coordinates The intensity value of the red channel at that pixel; and These correspond to the intensity values of the green and blue channels, respectively. The ICA blind source separation formula is as follows: Observation signal matrix ; It is a dimension of The observed signal matrix, This indicates that the matrix has 3 rows, corresponding to red... ,green ,blue Three color channels at different times The signal value, The number of sampling points indicates the number of times the signal is sampled within a certain period of time; Separation model: Unmixing: ; in, It is a mixed matrix; This is the unmixing matrix; That is, the original signal vector containing three components, where The rPPG (remote photoplethysmography) signal component. For noise components, the maximum pulse component is taken as the original rPPG signal. By analyzing the changes of each component over time Signals with pulse wave characteristics were selected as the original rPPG signals. S2-4. Calculation of physiological indicators: Perform FFT frequency domain transformation on the heart rate filter signal to calculate the instantaneous heart rate, with an error ≤3 beats / min; extract the amplitude modulation wave from the respiratory filter signal to calculate the respiratory rate, with an error ≤2 beats / min; and calculate the heart rate variability simultaneously. The FFT frequency domain transform formula is as follows: in, Indicates signal The frequency domain transform result; That is, Fast Fourier Transform; This is the original time-domain signal; It is a frequency variable; This is a discrete frequency index, with values ranging from 0 to... ; It is the sampling frequency; take the main frequency. , It is the dominant frequency obtained from frequency domain analysis, which corresponds to the main frequency components of heartbeat; Instantaneous heart rate calculation formula: (Unit: bpm) This indicates the instantaneous heart rate, expressed in heartbeats per minute. Respiratory rate calculation formula: (Unit: times / minute) Represents respiratory rate, measured in breaths per minute; It is the dominant frequency obtained from the frequency domain analysis of the respiratory signal, reflecting the main frequency characteristics of respiratory motion; Heart rate variability (HRV) calculation: Normal cardiac cycle sequence is , The sequence records the time interval (usually in seconds) between adjacent normal heartbeats. Indicates the first The duration of a heartbeat cycle, The total number of central dynamic periods of the sequence; Time-domain metrics SDNN: SDNN, or standard deviation of all normal sinus intervals (NN intervals), is used to measure the overall magnitude of heart rate variability. yes The average value of the sequence; this formula quantifies the dispersion of the cardiac cycle by calculating the square root of the sum of squares of the deviations of each cardiac cycle duration from the average duration. Time-domain metric RMSSD: RMSSD is the root mean square of the difference between adjacent normal heartbeats, which focuses on reflecting the rapid changes in heart rate variability. The formula captures the short-term fluctuations of the heart cycle by calculating the square root of the sum of the squares of the differences in the duration of adjacent cardiac cycles. S2-5. Quantification of Emotional Arousal and Stress Index: Integrating standardized heart rate values, respiratory fluctuation values, facial micro-movements, and limb tension characteristics, calculates emotional arousal from 0 to 100; based on abnormal heart rate, disordered respiratory rhythm, emotional arousal threshold, and behavioral stress characteristics, outputs a stress index of 0 to 5. The feature standardization formula is as follows: in, These are the original feature values, i.e., the original data to be standardized; This is the minimum value in the original dataset, used to determine the lower limit of the data range; The maximum value in the original dataset is used to determine the upper limit of the data range; The standardized feature values are then mapped from the original data to... using a formula. The interval facilitates subsequent processing and analysis; The formula for measuring emotional arousal (0–100) is as follows: Let the standardized heart rate be 1. respiratory fluctuation value Facial micro-movement feature value Limb tension characteristic value The weights of each feature are as follows: , , , And satisfy The formula for calculating emotional arousal is: The rules for the stress index grading model (0–5 levels) are as follows: Let the abnormal heart rate index be (Normal is 0, abnormal is 1), respiratory rhythm disorder index is (Normal is 0, abnormal is 1), the emotional arousal threshold is Behavioral stress characteristic value Stress index The calculation process is as follows: when and and hour, ; when or and hour, (Mild pressure); when and and hour, (Mild pressure); when or and and hour, (Moderate pressure); when and and and hour, (Severe stress); when and and and ( )hour, (Severe stress); , These are used to determine the stress index. Two thresholds for the level, and Through these two thresholds and The comparison allows for the assessment of the stress index. More detailed classification judgment; S2-6, Cross-modal feature fusion: The behavioral feature vector (limbs, language, vision) and the physiological feature vector (heart rate, respiration, emotion, stress) are concatenated to generate a 192-dimensional cross-modal fusion feature vector; The eigenvector is defined as follows: Behavioral feature vector: (48-dimensional features of limbs + 40-dimensional features of language + 40-dimensional features of vision) Physiological feature vector: ( 16-dimensional features + 16-dimensional features + 16-dimensional features + (16-dimensional features) Represents the heart rate feature vector. Represents the respiratory rate feature vector. This represents the eye movement feature vector, which records the relevant parameters of eye movement; The pulse wave conduction index is used to measure the relevant characteristics of pulse wave conduction in blood vessels; The feature splicing formula is as follows: The formula for normalizing end-side features is as follows: in, The mean; Standard deviation; S2-7, Dual-scale fusion scoring: Input cross-modal fusion features into the dual-scale AI assessment model, and output Gesell developmental results, developmental level, CARS autism behavioral risk level in real time, generating a physiological-behavioral dual-dimensional assessment report; S3, Generation of Personalized Intervention Plans at the Endpoint Based on the assessment results generated on the device, a personalized intervention plan matching the child's individual situation is automatically generated through a locally deployed personalized intervention plan generation model: For healthy children, a "growth progression" program is developed, which focuses on strengthening the children's weak developmental dimensions and improving their overall quality at the same time. The program includes 2-3 training sessions per week, each lasting 15-20 minutes. For children at high risk of autism, a "low-sensitivity adaptive" intervention program is generated, which automatically activates a low-sensory stimulation mode, sets training difficulty in a graded manner, focuses on improving the core dimensions of social, language and motor control, and sets a short training plan of 3-4 times a week, 10-15 minutes each time. Once the solution is generated, it directly connects to the device's local UE5 course library, allowing children to access the corresponding training courses with one click without needing to download resources online. S4. End-sided intervention training and data collection During the UE5 motion-sensory course training, which is conducted according to the intervention plan, the edge device collects multi-dimensional training data in real time, including: the completion rate of each task, the rate of achieving the required action, the duration of sustained focus, the emotional change curve, the frequency of social interaction response, the number of stereotyped actions and other core indicators. All collected data is encrypted using the AES-256 encryption algorithm and stored in the device's local encrypted partition. Only authorized parents can view and export the data, and there is no data upload to the cloud throughout the process. S5. Evaluation of End-Stage Intervention Effectiveness and Program Iteration After each preset training cycle (2 weeks), the intervention effect evaluation model deployed on the device automatically quantifies the improvement rate of each training dimension of the child's ability based on the collected full training data, and generates a visual training effect report. When the child's single-dimensional ability improvement rate is lower than the preset threshold of 20% for two consecutive training cycles, the system automatically adjusts the intervention plan, changes the appropriate training courses, reduces the task difficulty, and adjusts the adaptation mode and training frequency. When the child's ability improves significantly and continuously meets the target, the system gradually increases the training difficulty to achieve step-by-step ability growth. The entire evaluation and plan iteration process is completed locally on the device, without cloud intervention or manual operation. S6, End-side Retesting and Management The system automatically pushes retesting tasks according to the child's developmental risk level: low-risk children are retested at 3 months, medium-risk children at 1 month, and high-risk children at 2 weeks. The retesting is completed through gamified assessment tasks on the device, updating the child's assessment results and risk level. Based on the retesting results, an optimized intervention plan is regenerated, forming a closed loop on the device throughout the entire process of "assessment - plan generation - intervention training - effect evaluation - retesting - plan optimization".
[0017] Meanwhile, the system generates a full-cycle growth profile for children on the client side, recording assessment and training data throughout the entire process; it sets up a four-level hierarchical access control system for parents, teachers, intervention specialists, and administrators, with different roles only able to view content within their corresponding permissions and not allowed to exceed their authority; it also generates tamper-proof operation audit logs, with a storage period of ≥3 years, fully complying with the relevant laws and regulations on the protection of minors' personal information.
[0018] In summary, this edge device-based closed-loop method for children's mental health assessment and intervention, by quantifying and compressing three core models—dual-scale AI assessment, personalized intervention plan generation, and intervention effect evaluation—and deploying them on an ARM edge smart screen, enables offline operation of the entire process of children's mental health assessment, intervention, effect iteration, and retesting. It completely eliminates reliance on cloud computing power and networks, solving the implementation challenges in grassroots and rural areas with weak networks. Furthermore, through AES-256 encrypted local storage and four-level access control, it fundamentally avoids the risk of children's privacy data leakage, breaking the industry pain point of traditional disconnect between testing and training. With its lightweight model adaptability and the fact that it requires no professional personnel to operate, it achieves universal implementation in family, school, and community scenarios. A single system covers the differentiated needs of healthy children and children at risk of autism, possessing strong practicality and social value. By introducing non-contact physiological signal extraction technology using binocular cameras, and employing algorithms such as rPPG blind source separation, binocular depth motion compensation, and cross-modal feature fusion, this invention achieves precise on-device quantification of heart rate, respiratory rate, emotional arousal, and stress index. This not only compensates for the insufficient accuracy of traditional single-behavioral modality assessments and reduces the scoring error of dual scales, but also eliminates the need for wearable devices, significantly improving compliance for young children and children with special needs. Furthermore, the deep integration of physiological indicators and behavioral characteristics allows personalized intervention plans to accurately adapt to children's real-time emotional and stress states, effectively addressing issues such as emotional breakdowns and low compliance during training, thus enhancing intervention effectiveness. All physiological signal extraction, calculation, and fusion processes are completed offline on the device, preserving the privacy compliance and hardware compatibility of the original closed loop, further strengthening the invention's technological barriers and practical value.
[0019] The embodiments of the present invention are given for illustrative and descriptive purposes only, and are not intended to be exhaustive or to limit the invention to the forms disclosed. Many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described in order to better illustrate the principles and practical application of the invention, and to enable those skilled in the art to understand the invention and to design various embodiments with various modifications suitable for a particular purpose.
Claims
1. A closed-loop method for edge device-based children's mental health assessment and intervention, characterized in that: Includes the following steps: S1. Edge AI Model Deployment: The trained dual-scale AI assessment model for child development, personalized intervention plan generation model, and intervention effect evaluation model are quantified and compressed, and then deployed on the NPU or CUDA computing unit of the ARM edge smart screen. All AI models support real-time inference on the edge and do not require cloud computing power support. S2, Dual-Scale AI Assessment on the Edge: Through the motion sensing and audio acquisition devices of the edge smart screen, multi-dimensional data is collected during the assessment of children on the edge. The dual-scale AI assessment model deployed on the edge calculates and outputs the assessment results and risk levels of the Gesell Developmental Scale and the CARS Childhood Autism Behavior Screening Scale in real time. S3. Personalized intervention plan generation on the edge: Based on the edge assessment results, the personalized plan generation model deployed on the edge is used to automatically generate a personalized intervention plan on the edge by inputting the child's age, developmental characteristics, and risk level. This includes a list of UE5 motion-sensing courses for the corresponding training dimensions, training frequency, difficulty gradient, and adaptation mode. S4. Edge-side intervention training and data collection: During the process of children completing the UE5 somatosensory intervention course on the edge smart screen according to the intervention plan, the edge device collects all-dimensional training data of children in real time, including task completion rate, action achievement rate, attention duration, emotional changes, and social interaction performance, and stores the data locally throughout the process. S5. End-to-end intervention effect evaluation and program iteration: Through the intervention effect evaluation model deployed on the end-to-end, based on the collected training data, the intervention effect is quantitatively evaluated, the improvement of children's abilities and remaining shortcomings are identified, and the intervention program is automatically optimized and adjusted on the end-to-end to achieve dynamic iteration of the intervention program. S6. End-to-End Retesting and Management: Based on the child's intervention cycle and risk level, the end-to-end automatically pushes retesting tasks, completes the retesting through gamified assessment tasks, updates the child's developmental score and risk level, and regenerates the intervention plan based on the retest results, forming a closed loop on the end-to-end throughout the entire process; at the same time, the end-to-end completes the encrypted storage of all process data, the generation of growth reports, and the recording of compliance audit logs.
2. The method for implementing a closed-loop approach to edge device-based children's mental health assessment and intervention according to claim 1, characterized in that: In step S2, the multi-dimensional data collection includes the full-dimensional collection of limb movement data, language feature data, and visual behavior data. All data collection, feature extraction, and evaluation calculation processes are completed on the device side without cloud data transmission.
3. The method for implementing a closed-loop approach to edge device-based children's mental health assessment and intervention according to claim 1, characterized in that: In step S4, the UE5 motion-sensing intervention course is a lightweight 3D motion-sensing course adapted to edge device operation. The course content corresponds one-to-one with the evaluation indicators to achieve targeted training of the weaker dimensions.
4. The method for implementing a closed-loop approach to edge device-based children's mental health assessment and intervention according to claim 1, characterized in that: In step S5, the intervention effect evaluation model quantifies the ability improvement rate of children in each training dimension in real time. When the ability improvement rate of children is lower than the preset threshold for two consecutive training cycles, the difficulty, adaptation mode and training frequency of the intervention course are automatically adjusted.
5. The method for implementing a closed-loop approach to edge device-based children's mental health assessment and intervention according to claim 1, characterized in that: In step S6, the retesting period is automatically set according to the risk level, specifically: 3 months for low-risk children, 1 month for medium-risk children, and 2 weeks for high-risk children.
6. The method for implementing a closed-loop approach to edge device-based children's mental health assessment and intervention according to claim 1, characterized in that: In step S6, all data in the process is encrypted and stored locally using AES-256, and a four-level hierarchical access control system is set up for parents, teachers, intervention specialists, and administrators. Unalterable compliance audit logs are generated, and the storage period is ≥3 years.
7. The method for implementing a closed-loop approach to edge device-based children's mental health assessment and intervention according to claim 1, characterized in that: The multi-dimensional data acquisition in step S2 also includes non-contact end-to-end extraction of physiological signals using a binocular camera, specifically including the following sub-steps: S2-1, Binocular video stream preprocessing: The left-channel RGB video frame and the right-channel depth video frame are simultaneously acquired by a 120FPS global shutter binocular camera. The device side locates the optimal ROI region of the child's forehead and cheeks in real time, and calculates the three-dimensional spatial coordinates of the ROI region based on the depth frame. S2-2, Binocular Depth Motion Compensation: Based on binocular depth frames, the translation, rotation, and jitter amplitude of the child's face and head are calculated in real time. Spatial alignment compensation is performed on RGB video frames to eliminate motion interference and obtain a clean ROI image sequence. S2-3, End-side rPPG signal separation and noise reduction: RGB three-channel light intensity change time-series signal is extracted pixel by pixel in the ROI region, blood flow-related rPPG raw signal is extracted by ICA blind source separation, and heart rate and respiration corresponding filtered signals are obtained by wavelet filtering and adaptive bandpass filtering, respectively. S2-4. Calculation of physiological indicators: Perform FFT frequency domain transformation on the filtered heart rate signal to calculate the instantaneous heart rate; The respiratory rate is calculated by extracting the amplitude modulation wave from the respiratory filter signal; S2-5. Quantification of emotional arousal and stress index: Integrating heart rate variability, respiratory fluctuations, facial micro-movements, and limb tension characteristics, the emotional arousal is quantified from 0 to 100 by the end-side calculation. Based on heart rate abnormalities, respiratory disturbances, emotional arousal thresholds, and behavioral stress characteristics, the edge outputs a stress index of 0-5 levels. The entire process of extracting, calculating, and quantifying the aforementioned physiological signals is completed on the device side, without cloud data transmission or cloud computing.
8. The method for implementing a closed-loop approach to edge device-based children's mental health assessment and intervention according to claim 7, characterized in that: In steps S2-3, the frequency bands of the adaptive bandpass filter are set as follows: heart rate passband 0.8–3Hz, respiratory passband 0.15–0.5Hz.