An interactive autism child training system, method, electronic device, readable storage medium

The interactive autism training system, which combines TOF radar and multimodal sensors, solves the problems of monotonous training and inaccurate assessments, realizes the scientific and systematic nature of personalized training programs, and improves training effectiveness and data support.

CN121789911BActive Publication Date: 2026-05-08GUANGZHOU61LEARN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU61LEARN INFORMATION TECH CO LTD
Filing Date
2026-03-05
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Traditional training equipment for children with autism suffers from problems such as a boring training process, poor interactivity, lack of objective assessment, and inability to capture the details of movements, making it difficult to quantify the training effect and implement personalized programs accurately.

Method used

Employing TOF radar, a multimodal sensor fusion camera group, a gesture practice recognition module, an entity interaction behavior recognition module, and an edge-side intelligent computing module, this system enables personalized training through multimodal feedback. By combining TOF radar and cameras to automatically collect interaction and action data during the training process, it provides a scientific and systematic training solution.

Benefits of technology

It enhances the attractiveness and focus of the training process, promotes the transfer of skills from cognition to practice, provides reliable data support, and improves the scientific and systematic nature of training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789911B_ABST
    Figure CN121789911B_ABST
Patent Text Reader

Abstract

The application discloses an interactive autistic child training system, method, electronic device and readable storage medium. The interactive autistic child training system comprises a training configuration module, a TOF radar, a multi-modal sensor fusion camera group, a gesture exercise recognition module, an entity interaction behavior recognition module and a prompt module. The training configuration module is used for configuring training actions / training elements corresponding to autistic life self-care skills. The TOF radar is used for collecting gesture exercise data of palms and knuckles. The gesture exercise recognition module is used for judging the gesture exercise data. The multi-modal sensor fusion camera group is used for collecting entity interaction behavior data of a trainee. The entity interaction behavior recognition module is used for judging the entity interaction behavior data. The prompt module is used for outputting an object according to a personalized ability model of the trainee. The application is used to improve the scientificity and systematicness of training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of human-computer interaction devices, specifically relating to an interactive training system, method, electronic device, and readable storage medium for children with autism. Background Technology

[0002] Children with autism spectrum disorder (ASD) typically face significant difficulties in learning and developing self-care skills, requiring systematic, highly repetitive, and structurally rigorous training equipment to effectively guide the treatment process and monitor patient adherence. Current technologies suffer from at least the following problems:

[0003] (1) Traditional training adopts one-on-one manual guidance, which relies heavily on the trainer's personal experience. The teaching process is relatively boring, making it difficult to maintain the interest and focus of autistic children in the long term. In addition, there is a lack of objective and consistent means of evaluating training effectiveness.

[0004] (2) Some existing technologies use a simple video viewing learning method, which has poor interactivity, and children are in a passive receiving state, lacking a way to transfer skills from cognition to practice;

[0005] (3) Some existing interactive projection training methods mostly use simple touch or basic motion sensing technology. The interaction dimension is single, which cannot effectively capture and quantify the fine motor details of children when operating real objects. It is also difficult to systematically record and quantify the entire training process, thus failing to accurately guide the treatment process and detect patient compliance for personalized training programs. Summary of the Invention

[0006] The purpose of this application is to provide an interactive training system, method, electronic device, and readable storage medium for children with autism, so as to improve the scientific and systematic nature of the training.

[0007] In the first aspect, this application provides an interactive training system for children with autism, including a training configuration module, a TOF radar, a multimodal sensor fusion camera group, a gesture practice recognition module, an entity interaction behavior recognition module, a prompting module, and an edge-side intelligent computing module.

[0008] The training configuration module is used to build a personalized ability model based on the trainee's cognitive level, motor ability, and perceptual sensitivity, and to configure training actions / training elements corresponding to autism self-care skills according to the personalized ability model. It also supports dynamic adaptive parameter adjustment of training difficulty, training rhythm, and feedback mode.

[0009] TOF radar is used to perform dynamic background suppression, palm point cloud segmentation and 3D world coordinate system registration on the hand area of ​​the trainee, and to collect hand gesture practice data of the palm and fingers.

[0010] The gesture practice recognition module is used to extract 3D spatiotemporal features, reduce feature dimensionality and match similarity in gesture practice data, determine the matching degree between the data and the configured training actions / training elements, and evaluate the proficiency of the trainees through a multi-dimensional quantitative model.

[0011] A multimodal sensing fusion camera group is used to collect the full-body fine skeletal point three-dimensional spatiotemporal coordinate data stream of the training object, multi-scale enhanced image data of the entity object, and limb-entity interaction contact data as entity interaction behavior data;

[0012] The entity interaction behavior recognition module is used to perform spatiotemporal correlation feature fusion and interaction logic reasoning on entity interaction behavior data, determine its matching degree with the configured training actions / training elements, and evaluate the trainee's action proficiency.

[0013] The prompting module is used to output training guidance objects and recognition result feedback objects that are linked by visual, auditory and tactile multimodal senses based on the personalized ability model of the trainee;

[0014] The edge-side intelligent computing module adopts a heterogeneous computing architecture to achieve local real-time preprocessing of all sensor data, low-latency inference of AI models, privacy desensitization of training data, and collaborative scheduling of multiple modules.

[0015] Preferably, the prompting module includes an AR spatial perspective projection unit, a multi-channel spatial audio unit, and a haptic feedback unit;

[0016] The AR spatial perspective projection unit is used to anchor the visual objects of training actions / training elements onto the physical entity interaction area based on spatial calibration and planar reprojection algorithms, and realize the virtual-real fusion display. At the same time, it outputs the recognition result feedback object with dynamic visual enhancement. The visual enhancement strategy is adaptively adjusted according to the visual perception sensitivity of the training object.

[0017] The multi-channel spatial audio unit is used to output training-guided auditory objects that match the spatial position of visual objects using customized TTS speech synthesis and spatial sound image localization technology. It also supports adapting the frequency, volume, speech rate and feedback interval of the audio according to the auditory perception threshold of the trainee, and outputs recognition result feedback objects with voice emotion.

[0018] The tactile feedback unit is used to output tactile feedback when the trainee makes matching / non-matching actions through a vibration module with vibration frequency / intensity graded, thereby realizing a multimodal feedback closed loop of vision-auditory-tactile feedback.

[0019] More preferably, the TOF radar is a solid-state TOF radar with a sampling frame rate of ≥120fps and a point cloud resolution of ≥640×480.

[0020] The TOF radar is equipped with a point cloud preprocessing unit, which uses Euclidean clustering and region growing algorithms to achieve precise segmentation of the palm and knuckle regions. It collects the three-dimensional spatiotemporal coordinate sequence of 21 fine skeletal points of the trainee's palm as corresponding gesture practice data, and eliminates the inter-frame drift of hand movements through inter-frame point cloud registration.

[0021] The gesture practice recognition module pre-builds a 3D spatiotemporal feature template library of training actions. It uses an improved 3DCNN+LSTM hybrid network to extract spatiotemporal features from the 3D spatiotemporal coordinate sequence of fine skeletal points on the palm, generating a high-dimensional gesture feature vector. Then, it uses a two-dimensional matching algorithm of Euclidean distance and cosine similarity, combined with dynamic distance threshold constraints, to determine the matching degree between the gesture feature vector and the corresponding training action features in the feature template library. It also supports differentiated recognition and matching of fine gestures with one hand and two hands.

[0022] More preferably, the multimodal sensing fusion camera group includes a motion-sensing depth camera, a high-definition visible light camera, and an infrared interactive contact camera;

[0023] A motion-sensing depth camera acquires a three-dimensional spatiotemporal coordinate sequence data stream of ≥25 refined skeletal points on the whole body of the trainee, and smooths and completes the trajectories of the skeletal points through Kalman filtering;

[0024] High-definition visible light cameras acquire multi-scale fused image data of physical objects, which are then preprocessed by low-light enhancement and motion blur removal to serve as visual data of the physical objects.

[0025] Infrared interactive contact cameras collect data on the contact area, contact timing, and contact pressure trend between the trainee's limbs and physical objects.

[0026] The entity interaction behavior recognition module uses a cross-modal feature fusion network to deeply fuse the spatiotemporal features of limb skeleton points, the visual features of entity objects, and the interaction contact features between limbs and entities to generate a joint feature vector of limb-entity interaction. Then, it uses an attention mechanism to focus on key interaction features and judge the matching degree between the joint feature vector and the configured training actions / training elements. It also supports triple judgment of the temporal logic, contact accuracy, and action amplitude of limb-entity interaction.

[0027] More preferably, the entity interaction behavior recognition module includes a limb trajectory modeling unit with a spatiotemporal attention mechanism, an entity object recognition unit with few-shot learning, and a limb-entity interaction logic reasoning unit;

[0028] The limb trajectory modeling unit is used to allocate attention weights to the three-dimensional spatiotemporal coordinate sequence data stream of the whole body's fine skeletal points, focus on key skeletal points related to training movements, extract the motion trajectory, angle change, and velocity / acceleration features of key skeletal points through trajectory feature decoupling, calculate the limb movement trajectory of the trainee, and determine the degree of overlap between the trajectory and the set training movements in terms of spatial form, movement sequence, and movement amplitude.

[0029] The entity object recognition unit is used to extract dual features of visual and semantic features from multi-scale fused image data of entity objects. Based on the few-shot learning algorithm, it accurately identifies scarce entity samples in autism self-care training and judges the placement posture, interaction state, and similarity between entity objects and training elements.

[0030] The limb-entity interaction logic reasoning unit is used to perform interaction logic reasoning based on limb trajectory features and entity object state features using a Bayesian network to determine whether the limb movements of the trainee have completed the interaction behavior that meets the training requirements for the target entity object.

[0031] More preferably, the limb spatial coordinate data stream is a microsecond-level timestamp three-dimensional spatiotemporal coordinate sequence data stream of ≥25 refined skeletal points on the whole body of the trainee, and the tracking accuracy of the skeletal points is ≤0.5mm;

[0032] The limb trajectory modeling unit is based on an improved dynamic time warping algorithm with feature weighting and time window constraints. It combines a hidden Markov model to perform trajectory feature modeling and limb movement trajectory calculation on the spatial coordinate sequence data stream of skeletal points, and realizes probabilistic inference of movement time sequence to determine whether the limb movement trajectory coincides with the set training movement.

[0033] The entity recognition unit adopts a lightweight YOLOv9-Lite visual deep learning network that combines transfer learning and few-sample fine-tuning. The visual deep learning network is pre-trained and fine-tuned on a dedicated entity dataset for training self-care skills for children with autism, and supports incremental updates of model parameters on the edge.

[0034] The proficiency of the movement is judged based on a four-dimensional quantitative evaluation model. The four-dimensional quantitative evaluation model integrates four dimensions: the proportion of matching accuracy when the trainee performs the movement according to the training movement, the change in the smoothness of the matching movement within a set time period, the temporal compliance of the movement execution, and the accuracy of the interaction between the limb and the entity. In addition, the LSTM time series prediction model is used to predict the trend of movement proficiency. At the same time, the proficiency evaluation threshold is adaptively and dynamically adjusted according to the trainee's training curve.

[0035] Preferably, it also includes a reporting module, a medical record module, and a privacy computing module;

[0036] The reporting module is used to perform in-depth analysis of training data based on the judgment results of the gesture practice recognition module and / or entity interaction behavior recognition module, through machine learning clustering analysis and trend fitting algorithms. It generates a training report containing quantitative training data, proficiency trend analysis, interaction behavior characteristics, and intelligent diagnostic conclusions of training effect. Based on the diagnostic conclusions and combined with the trainee's personalized ability model, it generates customized optimization suggestions for subsequent training programs, including adaptive adjustments to training content, training difficulty, and training duration.

[0037] The medical record module is used to store the basic information, personalized ability model, training report and training data of the trainees in a blockchain-based encrypted manner, and supports authorized hierarchical access on multiple terminals and standardized synchronization of training data in multiple centers.

[0038] The privacy computing module is used to perform differential privacy desensitization and federated learning processing on all training data, enabling joint training and optimization of models across multiple devices and centers.

[0039] Secondly, this application provides an interactive training method for children with autism, applicable to any of the aforementioned interactive training systems for children with autism, comprising:

[0040] A personalized ability model is established based on the trainee's cognitive level, motor skills, and perceptual sensitivity. Training actions / elements corresponding to autism self-care skills are configured according to the personalized ability model. At the same time, dynamic adaptive parameter adjustment of training difficulty, training rhythm, and feedback mode is supported.

[0041] Dynamic background suppression, palm point cloud segmentation, and 3D world coordinate system registration were performed on the hand area of ​​the trainees, and hand gesture practice data of palm and finger joints were collected.

[0042] 3D spatiotemporal feature extraction, feature dimensionality reduction and similarity matching are performed on gesture practice data to determine its matching degree with the configured training actions / training elements, and the proficiency of the trainees' actions is evaluated through a multi-dimensional quantitative model.

[0043] Collect the full-body refined skeletal point three-dimensional spatiotemporal coordinate data stream of the training subject, multi-scale enhanced image data of the entity object, and limb-entity interaction contact data as entity interaction behavior data;

[0044] The spatiotemporal correlation features of limb-entity interaction behavior data are fused and interaction logic reasoning is performed to determine the matching degree with the configured training actions / training elements, and to evaluate the proficiency of the trainees' actions.

[0045] Based on the individualized ability model of the trainee, the system outputs a training guidance object and a recognition result feedback object that are linked by visual, auditory and tactile multimodal interaction.

[0046] It handles local real-time preprocessing of all sensor data, low-latency inference of AI models, privacy desensitization of training data, and collaborative scheduling of multiple modules.

[0047] Thirdly, this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the computer program to realize the interactive autism training system of the first aspect.

[0048] Fourthly, this application provides a readable storage medium storing computer instructions for enabling a computer to implement the interactive autism training system of the first aspect.

[0049] Compared with the prior art, the advantages of this application are as follows:

[0050] By enhancing the visual environment and combining it with Time-of-Flight (TOF) radar to achieve precise human-computer interaction, the training process is effectively made more attractive and engaging for children with autism. The assessment of training effectiveness through motor proficiency enhances the scientific rigor of the training. The virtual cognitive training interaction process is closely integrated with the real-world skill acquisition process through unified projection guidance and data collection, effectively promoting the transfer of skills training from cognition to practice. The automatic collection of interaction and movement data during training via TOF radar and cameras, using multimodal data, provides reliable data support for evaluating training effectiveness, monitoring changes in patient compliance, and developing personalized follow-up treatment plans, thus improving the scientific and systematic nature of the training. Attached Figure Description

[0051] Figure 1 This is a schematic diagram of the framework structure of an interactive training system for children with autism according to this application.

[0052] Figure 2 This is a flowchart illustrating an interactive training method for children with autism according to this application.

[0053] Figure 3 This is a schematic diagram of the frame structure of an electronic device according to this application. Detailed Implementation

[0054] Please refer to the diagrams, where the same component symbols represent the same components. The principles of this application are illustrated by way of example implementation in a suitable operating environment. The following description is based on the specific embodiments of this application exemplified, and should not be construed as limiting other specific embodiments not detailed herein.

[0055] Traditional training methods for children with autism spectrum disorder to improve their daily living skills often rely on manual guidance or simple interactive methods. These methods suffer from problems such as a monotonous training process, poor interactivity, difficulty in maintaining children's interest, lack of objective assessment, and inability to capture the details of movements. Consequently, the training effect is difficult to quantify, and personalized programs are difficult to implement accurately.

[0056] In this regard, such as Figure 1 As shown in the figure, this application discloses an interactive training system for children with autism, including a training configuration module, a TOF radar, a multimodal sensor fusion camera group, a gesture practice recognition module, an entity interaction behavior recognition module, a prompting module, and an edge-side intelligent computing module.

[0057] The training configuration module is used to build a personalized ability model based on the trainee's cognitive level, motor skills, and perceptual sensitivity, and to configure training actions / training elements corresponding to autism self-care skills according to the personalized ability model. It also supports dynamic adaptive adjustment of training difficulty, training rhythm, and feedback mode.

[0058] The training configuration module can be implemented in the following ways. For example, a professional trainer can manually input the trainee's various ability assessment results, such as their cognitive level, motor skills, and sensory sensitivity recorded through questionnaires or observation. Based on this information, the trainer can then manually select a preset training template in the system and adjust the combination of training movements or elements. Adjustments to training difficulty, pace, and feedback patterns can also be made manually in real-time by the trainer based on experience during the training process. Alternatively, the system can provide a series of training courses with fixed difficulty levels. After completing a course, the trainer can manually select the next difficulty level based on the trainee's performance and manually adjust the frequency and intensity of feedback.

[0059] TOF radar, or Time-of-Flight radar, is used to perform dynamic background suppression, palm point cloud segmentation, and 3D world coordinate system registration on the hand area of ​​the trainee, and to collect hand gesture practice data of the palm and fingers.

[0060] Time-of-Flight (TOF) radar can be deployed in various ways. For example, a commercially available general-purpose TOF sensor can be used, placed in front of the trainee, and its built-in firmware can be used for background removal and point cloud data output. Hand point cloud segmentation can be achieved by setting a fixed depth threshold or region boundary, while the alignment criteria of the 3D world coordinate system can be roughly calibrated using preset sensor position and attitude parameters. Thus, the acquired hand and finger point cloud data can be directly used as gesture training data. Alternatively, a single-point or linear array TOF sensor can be used to acquire depth information of the hand region through scanning. These discrete depth points are then stitched together to form a hand point cloud, followed by basic background removal and coordinate transformation.

[0061] The gesture practice recognition module is configured to process gesture practice data collected by TOF radar, and to perform 3D spatiotemporal feature extraction, feature dimensionality reduction and similarity matching on the gesture practice data, to determine the matching degree between the data and the configured training actions / training elements, and to evaluate the trainee's proficiency in the actions through a multi-dimensional quantification model.

[0062] The implementation of the gesture practice recognition module can include the following methods. For example, traditional machine learning algorithms, such as Support Vector Machines (SVM) or decision trees, can be used to extract features from the collected gesture practice data. These features can be the relative positions, angles, or movement speeds of key hand points. Feature dimensionality reduction can be achieved through methods such as Principal Component Analysis (PCA). Similarity matching can be performed by calculating Euclidean distance or Manhattan distance to determine the closeness of the current gesture to a preset template. The evaluation of movement proficiency can be based on binary judgment (match / mismatch) or a preset fixed score range. As another implementation method, a rule-based expert system can be used, which predefines a series of geometric features and temporal rules for gestures, and then compares the collected data with these rules to determine the degree of matching.

[0063] Multimodal sensor fusion camera group refers to a camera combination that integrates multiple sensing technologies (such as depth perception, visible light imaging, infrared sensing, etc.) to collect the three-dimensional spatiotemporal coordinate data stream of the full-body fine skeletal points of the training object, multi-scale enhanced image data of the entity object, and limb-entity interaction contact data as entity interaction behavior data.

[0064] The implementation of a multimodal sensor fusion camera array can include the following methods. For example, a separate depth camera can be used to collect full-body skeletal point data of the training subject, a separate visible light camera can be used to collect image data of the object, and a separate pressure sensor or touch sensor can be used to collect contact data between the limb and the object. These data are collected separately, synchronized using timestamps, and then used as entity interaction behavior data. As another implementation, multiple independent sensors, such as multiple infrared sensor arrays, can be used to detect the proximity or contact between the limb and the object, while a standard RGB camera captures the object image. Skeletal point data is extracted from the RGB image through manual labeling or image processing algorithms.

[0065] The entity interaction behavior recognition module is configured to process entity interaction behavior data collected by a multimodal sensing fusion camera group. It is used to perform spatiotemporal correlation feature fusion and interaction logic reasoning on the entity interaction behavior data, determine its matching degree with the configured training actions / training elements, and evaluate the training subject's action proficiency.

[0066] The implementation of the entity interaction behavior recognition module can include the following methods. For example, feature concatenation can be used to linearly combine limb skeleton point data, entity image features, and contact data to form a comprehensive feature vector. Interaction logic reasoning can be performed through preset logical rules, such as "if the hand skeleton point is above the entity and there is a contact signal, then it is judged as a grasping action." Matching degree judgment can be performed by comparing with preset static feature templates. The evaluation of action proficiency can be based on the number of times the action is completed or whether it is successfully completed. As another implementation method, a finite state machine-based model can be used, predefining a series of interaction states and state transition conditions, and then judging whether the trainee completes the interaction according to the preset path based on real-time data.

[0067] The prompting module is used to output training guidance objects and recognition result feedback objects that are linked by visual, auditory and tactile multimodal senses, based on the individualized ability model of the trainee.

[0068] The implementation of the prompting module can include the following methods. For example, it can use a separate display to output visual guidance images, a separate speaker to play pre-recorded voice prompts, and a vibration motor to vibrate when a specific event occurs. These feedback methods can be preset according to a personalized ability model. For example, for trainees with high visual perception sensitivity, the brightness of the visual prompts can be manually reduced; for trainees with high auditory perception sensitivity, the volume of the auditory prompts can be manually reduced. The linkage of feedback can be achieved through timing control; for example, an auditory prompt can be played immediately after a visual prompt appears. As another implementation method, single-modal feedback can be used, such as providing guidance and feedback solely through text or images on the screen, or solely through voice prompts.

[0069] The edge-side intelligent computing module adopts a heterogeneous computing architecture to achieve local real-time preprocessing of all sensor data, low-latency inference of AI models, privacy desensitization of training data, and collaborative scheduling of multiple modules.

[0070] The implementation of edge-side intelligent computing modules can include the following approaches. For example, a single general-purpose processor (such as a CPU) can be used to process all sensor data. Data preprocessing can be achieved through sequentially executed software algorithms, and AI model inference can be performed using lightweight models or in the cloud. Privacy desensitization of training data can be achieved through anonymization, such as removing personally identifiable information. Cooperative scheduling of multiple modules can be achieved through the basic process management functions provided by the operating system. As another implementation approach, an embedded microcontroller can be used, with limited computing power, primarily responsible for data acquisition and filtering, while most complex computation and inference tasks are uploaded to a remote server for processing.

[0071] This embodiment, through the aforementioned system, can provide customized training content and dynamic adaptive parameter tuning based on the trainee's personalized ability model. It utilizes multi-sensor fusion technology to capture and recognize gestures and entity interaction behaviors, and provides timely and personalized feedback through a multimodal prompting module. At the same time, it leverages the edge-side intelligent computing module to ensure the real-time performance and privacy of data processing, thereby effectively solving the problems of low training efficiency, inaccurate evaluation, insufficient interactivity, and lack of personalized support in existing technologies.

[0072] In some implementations, the prompting module includes an AR spatial perspective projection unit, a multi-channel spatial audio unit, and a haptic feedback unit. The AR spatial perspective projection unit, based on spatial calibration and planar reprojection algorithms, anchors the visual objects of training actions / training elements onto the physical entity interaction area, achieving a virtual-real fusion display. It also outputs a recognition result feedback object with dynamic visual enhancement, the visual enhancement strategy adaptively adjusting according to the visual perception sensitivity of the trainee. The multi-channel spatial audio unit employs customized TTS speech synthesis and spatial sound image localization technology to output a training guidance auditory object matching the spatial position of the visual object. It supports adapting the audio frequency, volume, speech rate, and feedback interval according to the trainee's auditory perception threshold, while also outputting a recognition result feedback object with emotional voice. The haptic feedback unit uses a vibration module with graded vibration frequency / intensity to output haptic feedback when the trainee performs matching / non-matching actions, achieving a multimodal feedback closed loop of vision, hearing, and touch.

[0073] An AR spatial perspective projection unit is a display technology component that overlays virtual information onto a real physical environment. Its core lies in using spatial calibration and planar reprojection algorithms to accurately anchor virtual training actions or visual objects—such as virtual items, operation step instructions, or action demonstrations—to the physical interaction area where the trainee performs actual operations. This allows the trainee to train in a virtual-real blended environment, such as seeing virtual bowl and chopstick placement instructions on a real table. Furthermore, the unit can output recognition result feedback objects with dynamic visual enhancements. For example, when the trainee completes an action, the virtual object may flash, change color, or display special effects to enhance the intuitiveness of the feedback. The visual enhancement strategy is not static but adaptively adjusted according to the trainee's visual perception sensitivity. For example, for children with high visual sensitivity, a softer visual effect may be used to avoid overstimulation. Implementation methods can include, but are not limited to, desktop projection systems based on projectors, transparent displays, or spatial augmented reality systems combined with depth sensors.

[0074] The multi-channel spatial audio unit is designed to provide directional and immersive auditory feedback. Employing customized TTS (Text-to-Speech) speech synthesis technology, it generates natural and clear voice commands or feedback information based on training content. It also supports flexible adjustment of audio frequency, volume, speech rate, and feedback intervals according to the trainee's auditory perception threshold, ensuring effective information reception without causing auditory discomfort. Combined with spatial image localization technology, the unit can output training guidance auditory objects that match the spatial location of visual objects. For example, when a virtual "cup" appears to the trainee's left front, a corresponding voice prompt will also come from the left front, enhancing the correlation between hearing and vision. Simultaneously, the unit can also output recognition result feedback objects with emotionally charged voices, such as using variations in tone and speech rate to express encouragement or prompts, enhancing the friendliness of the feedback. This can be achieved through a multi-speaker array system, simulating sound source location using sound field synthesis technology, or realizing spatial audio effects in headphones using binaural rendering technology.

[0075] The haptic feedback unit provides tactile information to the trainee through physical vibration. This unit typically contains one or more vibration modules that output tactile feedback at different frequencies and / or intensities according to a preset strategy. For example, when the trainee successfully completes a movement, they may feel a gentle, short vibration; while when the movement is mismatched or requires correction, they may feel a more pronounced or continuous vibration. This hierarchical vibration feedback mechanism allows for rich expressiveness of tactile information. The haptic feedback unit outputs tactile feedback instantly when the trainee performs a matching or mismatched movement, working in conjunction with visual and auditory feedback to achieve a multimodal feedback loop of visual-auditory-tactile, providing the trainee with immediate, intuitive, and comprehensive sensory information. Implementation methods can include linear resonant actuators (LRAs) or eccentric rotating mass (ERM) motors integrated into wearable devices (such as wristbands or vests), or vibrators integrated into training surfaces or physical props.

[0076] In some implementations, the TOF radar is a solid-state TOF radar with a sampling frame rate of ≥120fps and a point cloud resolution of ≥640×480. The TOF radar is equipped with a point cloud preprocessing unit, which achieves accurate segmentation of the palm and knuckle regions through Euclidean clustering and region growing algorithms. It collects the three-dimensional spatiotemporal coordinate sequence of 21 fine skeletal points of the trainee's palm as the corresponding gesture practice data, and eliminates the inter-frame drift of hand movements through inter-frame point cloud registration. The gesture practice recognition module pre-builds a 3D spatiotemporal feature template library of training actions. It extracts spatiotemporal features from the three-dimensional spatiotemporal coordinate sequence of fine skeletal points of the palm through an improved 3DCNN+LSTM hybrid network to generate a high-dimensional gesture feature vector. Then, it uses a two-dimensional matching algorithm of Euclidean distance and cosine similarity, combined with dynamic distance threshold constraints, to determine the matching degree between the gesture feature vector and the corresponding training action features in the feature template library, and supports differentiated recognition and matching of single-hand / double-hand fine gestures.

[0077] Specifically, the TOF radar is designed as a solid-state area-array TOF radar with a high sampling frame rate and high point cloud resolution. A sampling frame rate of ≥120fps means that the TOF radar can capture depth images at a rate of at least 120 frames per second, effectively capturing rapidly changing hand movement details, reducing motion blur, and ensuring that even fast or minute hand movements are clearly recorded. Simultaneously, a point cloud resolution of ≥640×480 ensures sufficiently dense depth information in the hand region, providing high-quality raw data for subsequent fine segmentation and skeletal point extraction. The solid-state area-array design ensures the stability and reliability of the device, making it suitable for long-term, high-frequency training environments.

[0078] To further improve data quality, the TOF radar is equipped with a point cloud preprocessing unit. This unit achieves precise segmentation of the palm and knuckle regions through Euclidean clustering and region growing algorithms. Euclidean clustering groups adjacent points into the same cluster based on spatial distance between point clouds, thus separating the hand point cloud from the background. The region growing algorithm, building upon this, further refines the identification and segmentation of precise palm and knuckle regions by setting certain growth criteria (such as normal direction, color, or depth similarity), ensuring that subsequent analysis focuses solely on the target hand structure, improving processing efficiency and accuracy. Based on this, the system collects the three-dimensional spatiotemporal coordinate sequence of 21 fine skeletal points from the trainee's palm as corresponding gesture practice data. These 21 fine skeletal points typically cover key joint positions such as the palm, finger roots, middle fingers, and fingertips, providing a comprehensive and detailed description of the hand's posture and movement trajectory. Collecting the three-dimensional spatiotemporal coordinate sequence of these skeletal points means not only recording the hand's position in space but also its dynamic information over time, providing a rich data foundation for the recognition of complex gestures. Furthermore, the system can eliminate inter-frame drift in hand movements through inter-frame point cloud registration technology. This technique aims to address the minute displacements or tremors that may occur in the trainee's hand between consecutive frames, which can lead to discontinuities or inaccuracies in skeletal point data. By aligning and correcting the point cloud data of adjacent frames, this "drift" phenomenon can be effectively eliminated, ensuring the temporal coherence and spatial consistency of gesture training data, thereby improving the accuracy of subsequent feature extraction and recognition.

[0079] In terms of gesture recognition, the gesture practice recognition module pre-constructs a 3D spatiotemporal feature template library for training actions. This template library stores feature representations of standard or desired training actions. These features are professionally designed and extracted, comprehensively reflecting the spatial morphology and temporal evolution of the actions. During recognition, the real-time gesture features of the trainee are compared with the features in this template library. To extract more discriminative features from the acquired 3D spatiotemporal coordinate sequence of fine hand skeletal points, the system employs an improved 3DCNN+LSTM hybrid network for spatiotemporal feature extraction, generating a high-dimensional gesture feature vector. The improved 3DCNN (3D Convolutional Neural Network) can effectively capture the local spatial features and structural information of hand skeletal points in 3D space. LSTM (Long Short-Term Memory) network excels at processing sequential data and can learn and memorize the temporal dependencies of gesture actions. Combining the two to form a 3DCNN+LSTM hybrid network can simultaneously extract fine spatial features and dynamic temporal features of gestures, generating high-dimensional gesture feature vectors containing rich semantic information, thus more comprehensively and accurately representing complex gestures.

[0080] To accurately determine the matching degree between gesture feature vectors and corresponding training action features in the feature template library, the system employs a two-dimensional matching algorithm combining Euclidean distance and cosine similarity, along with dynamic distance threshold constraints. Euclidean distance measures the absolute distance between two feature vectors in multidimensional space, reflecting their similarity. Cosine similarity focuses on the directional consistency of feature vectors and is insensitive to changes in vector scale. The two-dimensional matching algorithm comprehensively evaluates the similarity between real-time gestures and template gestures from different perspectives. Dynamic distance threshold constraints flexibly adjust the matching strictness based on factors such as training progress and trainee performance, making the matching process more adaptable and robust. Furthermore, the system supports differentiated recognition and matching of single-handed and two-handed fine motor gestures. For single-handed or two-handed operations that may be involved in training children with autism, the system can distinguish and recognize different types of gestures. For single-handed gestures, the system focuses on the fine motor skills of one hand; for two-handed gestures, it considers the coordination, relative position, and synchronization between both hands simultaneously, performing more complex matching to ensure comprehensive coverage and accurate evaluation of various training actions.

[0081] In some implementations, the multimodal sensing fusion camera group includes a motion-sensing depth camera, a high-definition visible light camera, and an infrared interactive contact camera. The motion-sensing depth camera acquires a three-dimensional spatiotemporal coordinate sequence data stream of ≥25 refined skeletal points on the whole body of the trainee, and uses Kalman filtering to smooth and complete the skeletal point trajectories. The high-definition visible light camera acquires multi-scale fused image data of the entity object, which is then preprocessed with low-light enhancement and motion blur removal as the entity object's visual data. The infrared interactive contact camera acquires data on the contact area, contact timing, and contact pressure trend between the trainee's limbs and the entity object. The entity interaction behavior recognition module deeply fuses the spatiotemporal features of limb skeletal points, the visual features of the entity object, and the limb-entity interaction contact features through a cross-modal feature fusion network to generate a joint feature vector of limb-entity interaction. Then, it focuses on key interaction features through an attention mechanism to determine the matching degree between the joint feature vector and the configured training actions / training elements, and supports triple judgment on the temporal logic, contact accuracy, and action amplitude of the limb-entity interaction.

[0082] Specifically, the multimodal sensor fusion camera group, as the core component for data acquisition, integrates different types of sensors to comprehensively and meticulously capture various interactive information of the trainee during the training process. The configuration of this camera group is designed to overcome the limitations of a single sensor in complex interactive scenarios, ensuring the richness and accuracy of the data.

[0083] A motion-sensing depth camera is used to acquire a sequence of 3D spatiotemporal coordinates for ≥25 finely detailed skeletal points across the entire body of a training subject. This camera can acquire 3D positional information of key body parts in real time, forming continuous skeletal point trajectory data. Kalman filtering technology is used to smooth the acquired skeletal point trajectory data, effectively removing sensor noise and transient jitter. It also fills in missing skeletal point information caused by occlusion or data loss, ensuring the continuity and accuracy of the skeletal point trajectory data and providing high-quality input for subsequent motion recognition.

[0084] High-definition visible light cameras are used to acquire multi-scale fused image data of physical objects. These cameras can capture visual information about objects in the training scene, including their appearance, color, and texture. To address motion blur issues caused by low light or rapid movement of the training object in real-world training environments, the acquired image data undergoes low-light enhancement and motion blur removal preprocessing. Low-light enhancement improves image visibility under low-light conditions, while motion blur removal restores image details lost due to motion, thus ensuring the clarity and usability of the visual data of the physical objects.

[0085] Infrared interactive contact cameras are used to collect data on the contact area, timing, and pressure trends between a trainee's limbs and a physical object. Utilizing infrared technology, this camera can accurately detect the specific location of contact, the timing of contact, and the pressure changes during contact. This detailed contact data is crucial for assessing the accuracy, force, and timing of a trainee's manipulation of the object, especially in training daily living skills requiring fine motor skills.

[0086] After receiving multimodal data from different cameras, the entity interaction behavior recognition module processes it through a cross-modal feature fusion network. This network deeply fuses the spatiotemporal features of limb skeleton points provided by the depth-sensing camera, the visual features of the entity object provided by the high-definition visible light camera, and the interactive contact features between the limb and the entity provided by the infrared interactive contact camera. This fusion is not a simple data superposition, but rather learns the intrinsic correlation between different modal data through a complex neural network structure, generating a joint feature vector that comprehensively represents limb-entity interaction behavior. Based on this, the module further introduces an attention mechanism, enabling the recognition system to intelligently focus on key interaction features most relevant to the current training action / training element, such as the movement of specific skeleton points, visual changes in specific entity regions, or specific contact patterns, thereby improving the accuracy and robustness of recognition. Finally, the module can determine the matching degree between the generated joint feature vector and the pre-configured training action / training element, and supports triple judgment of the temporal logic, contact accuracy, and movement amplitude of the limb-entity interaction, ensuring a comprehensive evaluation of the trainee's actions.

[0087] In some implementations, the entity interaction behavior recognition module includes a limb trajectory modeling unit with a spatiotemporal attention mechanism, an entity object recognition unit with few-shot learning, and a limb-entity interaction logic reasoning unit.

[0088] The entity interaction behavior recognition module is a core component of the interactive autism training system. Its main function is to receive and process entity interaction behavior data from a multimodal sensor fusion camera group, and then determine whether the trainee's actual operation conforms to the preset training actions or training elements. To achieve accurate recognition and evaluation of complex interactive behaviors, this module is further subdivided into three closely cooperating sub-units to handle key tasks such as limb movement, entity recognition, and interactive logic reasoning, respectively.

[0089] The limb trajectory modeling unit with a spatiotemporal attention mechanism is specifically responsible for in-depth analysis of the trainee's limb movements. Its core lies in the use of a spatiotemporal attention mechanism, meaning the system can intelligently identify and focus on the most critical skeletal points in the training movement and their changes over time, rather than processing all skeletal point data equally. For example, when training a "cup-holding" motion, the attention mechanism might focus more on the movement trajectories of the wrist and fingers, while giving less weight to minor swaying of the legs or torso. Through trajectory feature decoupling technology, this unit can decompose complex limb movements into independent trajectories, angle changes, and basic features such as velocity / acceleration, thereby more accurately calculating the trainee's limb movement trajectory. Finally, the unit compares the calculated trajectory with preset training movements, judging the degree of overlap in spatial form (e.g., hand posture), movement timing (e.g., grasping timing), and movement amplitude (e.g., the height of raising the hand).

[0090] The entity recognition unit with few-shot learning focuses on recognizing and judging the status of entities involved in the training process. In the self-care training of autistic children, there are many types of entities, and it may be difficult to obtain a large number of samples for model training of certain specific training props or daily necessities. To this end, this unit introduces a few-shot learning algorithm, which enables it to accurately identify these scarce entity samples even with only a small amount of sample data. For example, the system can accurately identify an uncommon type of tableware even if it has only seen a few pictures. In addition, this unit not only identifies the category of entity objects, but also performs dual feature extraction of visual and semantic features from multi-scale fused image data of entity objects to further determine their real-time placement posture (e.g., whether a cup is upright or upside down), interaction status (e.g., whether a door is open or closed), and whether the entity object is the same as or similar to the elements required for current training, providing comprehensive entity information for subsequent interaction logic reasoning.

[0091] The limb-entity interaction logic reasoning unit is the "brain" that connects limb movements with entity states and ultimately determines whether the interaction meets the training requirements. It receives limb trajectory features from the limb trajectory modeling unit and entity object state features from the entity object recognition unit, and performs interaction logic reasoning based on a Bayesian network. As a probabilistic graphical model, the Bayesian network effectively handles uncertainty and models causal relationships between different events. For example, in training for "opening a door," this unit comprehensively considers a series of limb movements and entity state changes, such as "whether the hand touched the doorknob," "whether the doorknob was turned," and "whether the door was pushed open," and judges whether the trainee's limb movements have completed the interaction behavior that meets the training requirements for the target entity object according to preset logic rules. This reasoning mechanism enables the system to understand the deeper meaning of the interaction behavior, rather than simply recognizing isolated actions or objects.

[0092] In some implementations, the limb spatial coordinate data stream is set as a microsecond-timestamped three-dimensional spatiotemporal coordinate sequence data stream of at least 25 finely detailed skeletal points throughout the trainee's body, with a skeletal point tracking accuracy of no more than 0.5 mm. This means that the system can capture the dynamic changes of the trainee's skeletal points throughout the body with extremely high temporal resolution and spatial accuracy, which is crucial for identifying subtle motor impairments or non-standard movements that may exist in children with autism, providing an extremely detailed and reliable raw data foundation for subsequent movement analysis.

[0093] The limb trajectory modeling unit is based on an improved dynamic time warping algorithm with feature weighting and time window constraints. It combines a Hidden Markov Model (HMM) to model trajectory features and calculate limb movement trajectories from the spatial coordinate sequence data stream of skeletal points, and performs probabilistic inference of the movement sequence to determine whether the limb movement trajectory overlaps with the set training movement. Specifically, feature weighting allows the system to assign different weights to different skeletal points based on their importance in a specific movement; for example, in a grasping movement, the weight of hand skeletal points is higher than that of leg skeletal points. Time window constraints help focus on the key stages of the movement when comparing movement sequences of different lengths or speeds. The improved dynamic time warping algorithm can effectively handle differences in the speed and rhythm of the trainee's movements, achieving flexible matching with standard movement templates. The HMM further probabilistically models the temporal structure of the movement, identifying temporal features such as the sequence and duration of movements, thus more accurately determining whether the trainee's limb movement trajectory overlaps with the preset training movement, even if the movement has some degree of variation.

[0094] The entity recognition unit employs a lightweight YOLOv9-Lite visual deep learning network combined with transfer learning and few-shot fine-tuning. YOLOv9-Lite, as a highly efficient object detection model, enables low-latency entity recognition on edge devices. Considering the potential scarcity of certain entity samples in self-care training for autistic children, transfer learning and few-shot fine-tuning techniques allow the model to accurately identify these scarce entity samples even with limited labeled data. Furthermore, the visual deep learning network is pre-trained and fine-tuned using a dedicated entity dataset for self-care training of autistic children, ensuring high accuracy in recognizing objects specific to the training scene. Support for incremental updates of model parameters on the edge means the system can continuously learn and optimize the model locally based on new training data or scene changes, without frequent uploads to the cloud, improving the system's adaptability and real-time performance.

[0095] Motor proficiency is assessed using a four-dimensional quantitative evaluation model. This model integrates four dimensions: the percentage of matching accuracy when the trainee performs the training movement, the change in the smoothness of the matched movement over a set time period, the temporal compliance of the movement execution, and the accuracy of limb-entity interaction. The percentage of matching accuracy measures the similarity between the movement and the standard template; the change in movement smoothness assesses the smoothness and coherence of the movement; the temporal compliance judges whether the order and rhythm of the movement steps are correct; and the accuracy of limb-entity interaction focuses on the precision of the interaction between the movement and the physical object. By integrating these four dimensions, a comprehensive and objective quantitative evaluation of the trainee's motor proficiency can be achieved. Furthermore, the trend of motor proficiency changes is predicted using an LSTM time series prediction model, which can anticipate the trainee's learning progress and potential difficulties. Simultaneously, the proficiency evaluation threshold is adaptively and dynamically adjusted based on the trainee's training curve, ensuring that the evaluation criteria can be personalized according to individual learning progress, avoiding a one-size-fits-all evaluation approach, and thus providing more accurate and humanized training feedback.

[0096] In some implementations, the interactive autism training system also includes a reporting module, a medical record module, and a privacy computing module.

[0097] The reporting module, a key analytical component of the system, goes beyond real-time training feedback, providing a comprehensive and in-depth evaluation and guidance of the training process. This module integrates machine learning clustering analysis and trend fitting algorithms to deeply analyze the judgment results generated by the gesture practice recognition module and / or entity interaction behavior recognition module. Specifically, machine learning clustering analysis can identify the performance patterns of trainees across different training actions or elements, for example, discovering common difficulties in fine gesture training or behavioral deviations in specific entity interaction tasks. The trend fitting algorithm is used to predict the long-term development trend of trainees' motor proficiency, thus providing trainers with forward-looking guidance. The training report generated by this module not only includes quantitative training data (such as number of completions, matching degree, and proficiency score) but also deeply analyzes proficiency change trends and interaction behavior characteristics, providing intelligent diagnostic conclusions. Based on these diagnostic conclusions and combined with the trainee's personalized ability model, the reporting module can automatically generate customized optimization suggestions for subsequent training programs, including adaptive adjustments to training content, difficulty, and duration, ensuring that the training program always matches the trainee's actual progress and needs.

[0098] The medical record module addresses the critical needs for secure, standardized storage and management of training data. It employs blockchain-based encrypted storage technology to securely store trainees' basic information, personalized ability models, generated training reports, and raw training data. The application of blockchain technology ensures the immutability and traceability of the data, providing a highly reliable record of the trainees' training history. Encrypted storage further protects the confidentiality of sensitive personal information, preventing unauthorized access. Furthermore, the module supports multi-terminal, tiered access authorization, meaning different user roles (such as trainers, parents, and doctors) can access corresponding data according to their permission levels, enabling granular data management and access control. Simultaneously, it supports standardized synchronization of training data across multiple centers, which is crucial for sharing and integrating training data across different training institutions or locations for more macro-level analysis and research, while ensuring consistency in data format and semantics, laying the foundation for cross-institutional collaboration.

[0099] The privacy-preserving computation module is designed to balance data-driven model optimization with stringent privacy requirements. It performs differential privacy anonymization and federated learning processing on all training data. Differential privacy anonymization is an advanced privacy-preserving technique that adds mathematically defined noise to the original data, making it difficult to deduce sensitive information about any specific individual from the aggregated results, even when the data is shared or analyzed, thus achieving a balance between data availability and privacy protection. Federated learning allows a global model to be trained collaboratively on multiple devices or centers (e.g., different training institutions) without directly sharing the original training data. Each local device trains its model using its local data and sends model updates (rather than the original data) to a central server for aggregation, enabling joint training and optimization of models across multiple devices / centers. This approach significantly reduces the risk of data leakage while leveraging abundant, dispersed data resources to improve the accuracy and generalization ability of AI models, promoting the intelligent upgrade of the entire training system.

[0100] Compared to existing technologies, this embodiment achieves precise human-computer interaction by enhancing the visual environment and combining it with TOF radar, effectively improving the attractiveness and focus of the training process for children with autism. It also enhances the scientific rigor of the training by assessing the training effect through motor proficiency evaluation. By closely integrating the virtual cognitive training interaction process with the real-world skill acquisition process through unified projection guidance and data collection, it effectively promotes the transfer of skills training from cognition to practice. Furthermore, by automatically collecting interaction and movement data during the training process using TOF radar and a camera, and using multimodal data, it provides reliable data support for evaluating training effectiveness, monitoring changes in patient compliance, and developing personalized follow-up treatment guidance plans, thereby improving the scientific rigor and systematic nature of the training.

[0101] like Figure 2 As shown, based on the same inventive concept, and corresponding to any of the above embodiments, this application also discloses an interactive training method for children with autism, including the following steps.

[0102] S1. Establish a personalized ability model based on the trainee's cognitive level, motor skills, and perceptual sensitivity, and configure training actions / training elements corresponding to autism self-care skills according to the personalized ability model. At the same time, it supports dynamic adaptive parameter adjustment of training difficulty, training rhythm, and feedback mode.

[0103] S2. Perform dynamic background suppression, palm point cloud segmentation and 3D world coordinate system registration on the hand area of ​​the trainee, and collect hand gesture practice data of palm and finger joints.

[0104] S3. Perform 3D spatiotemporal feature extraction, feature dimensionality reduction and similarity matching on the gesture practice data, determine its matching degree with the configured training actions / training elements, and evaluate the trainee's action proficiency through a multi-dimensional quantitative model.

[0105] S4. Collect the full-body refined skeletal point three-dimensional spatiotemporal coordinate data stream of the training object, multi-scale enhanced image data of the entity object, and limb-entity interaction contact data as entity interaction behavior data;

[0106] S5. Perform spatiotemporal correlation feature fusion and interaction logic reasoning on the entity interaction behavior data, determine its matching degree with the configured training actions / training elements, and evaluate the trainee's action proficiency.

[0107] S6. Based on the individualized ability model of the trainee, output the training guidance object and the recognition result feedback object of the visual-auditory-tactile multimodal linkage;

[0108] S7 performs local real-time preprocessing of all sensor data, low-latency inference of AI models, privacy desensitization of training data, and collaborative scheduling of multiple modules.

[0109] It should be noted that the method of this disclosure embodiment can be implemented by a single device, such as a computer, mobile terminal, or server. The method of this embodiment can also be applied in a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may implement only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.

[0110] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0111] The system described above is used to execute the corresponding interactive autism training system in any of the foregoing embodiments, and has the beneficial effects of the corresponding system embodiments, which will not be repeated here.

[0112] like Figure 3As shown, based on the same inventive concept, corresponding to any of the above embodiments, this application also discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor implements the computer program, it implements the above-mentioned interactive autism training system for children.

[0113] Specifically, the device includes: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected within the device via the bus 1050.

[0114] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), GPU (Graphics Processing Unit), or one or more integrated circuits, to implement relevant programs and achieve the technical solutions provided in the embodiments of this specification.

[0115] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and called by the processor 1010. The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touch screens, microphones, various sensors, etc., and output devices may include displays, projectors, speakers, vibrators, indicator lights, etc.

[0116] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB (Universal Serial Bus), network cable, etc.) or wireless means (such as mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).

[0117] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.

[0118] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.

[0119] The electronic devices described above are used to implement the corresponding interactive autism training system in any of the foregoing embodiments, and have the beneficial effects of the corresponding system embodiments, which will not be repeated here.

[0120] Based on the same inventive concept, corresponding to any of the above-described embodiments, this application also discloses a non-transitory computer-readable storage medium that stores computer instructions for enabling a computer to implement the interactive autism training system described above.

[0121] The computer-readable storage medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium, which can be used to store information accessible by a computing device. The computer instructions stored in the storage medium of the above embodiments are used to enable the computer to implement the interactive autism training system for children as described in any of the above embodiments, and have the beneficial effects of the corresponding system embodiments, which will not be repeated here.

[0122] The above description is merely a preferred embodiment and the technical principles employed in this application. This application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions that can be made by those skilled in the art will not depart from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of this application. The scope of this application is determined by the scope of the claims.

Claims

1. An interactive training system for children with autism, characterized in that, It includes a training configuration module, a TOF radar, a multimodal sensor fusion camera group, a gesture practice recognition module, an entity interaction behavior recognition module, a prompting module, and an edge-side intelligent computing module; The training configuration module is used to establish a personalized ability model based on the trainee's cognitive level, motor ability, and perceptual sensitivity, and to configure training actions / training elements corresponding to autism self-care skills according to the personalized ability model. It also supports dynamic adaptive parameter adjustment of training difficulty, training rhythm, and feedback mode. TOF radar is used to perform dynamic background suppression, palm point cloud segmentation and 3D world coordinate system registration on the hand area of ​​the trainee, and to collect hand gesture practice data of the palm and fingers. The gesture practice recognition module is used to extract 3D spatiotemporal features, reduce feature dimensionality and match similarity in gesture practice data, determine the matching degree between the data and the configured training actions / training elements, and evaluate the proficiency of the trainees through a multi-dimensional quantitative model. A multimodal sensing fusion camera group is used to collect the full-body fine skeletal point three-dimensional spatiotemporal coordinate data stream of the training object, multi-scale enhanced image data of the entity object, and limb-entity interaction contact data as entity interaction behavior data; The entity interaction behavior recognition module is used to perform spatiotemporal correlation feature fusion and interaction logic reasoning on entity interaction behavior data, determine its matching degree with the configured training actions / training elements, and evaluate the trainee's action proficiency. The prompting module is used to output training guidance objects and recognition result feedback objects that are linked by visual, auditory and tactile multimodal senses based on the personalized ability model of the trainee; The edge-side intelligent computing module adopts a heterogeneous computing architecture to achieve local real-time preprocessing of all sensor data, low-latency inference of AI models, privacy desensitization of training data, and collaborative scheduling of multiple modules. The entity interaction behavior recognition module includes a limb trajectory modeling unit with spatiotemporal attention mechanism, an entity object recognition unit with few-shot learning, and a limb-entity interaction logic reasoning unit. The limb trajectory modeling unit is used to allocate attention weights to the three-dimensional spatiotemporal coordinate sequence data stream of the whole body's fine skeletal points, focus on key skeletal points related to training movements, extract the motion trajectory, angle change, and velocity / acceleration features of key skeletal points through trajectory feature decoupling, calculate the limb movement trajectory of the trainee, and determine the degree of overlap between the trajectory and the set training movements in terms of spatial form, movement sequence, and movement amplitude. The entity object recognition unit is used to extract dual features of visual and semantic features from the multi-scale fused image data of entity objects. Based on the few-shot learning algorithm, it accurately identifies scarce entity samples in autism self-care training and judges the placement posture, interaction state, and similarity between entity objects and training elements. The limb-entity interaction logic reasoning unit is used to perform interaction logic reasoning on limb trajectory features and entity object state features based on Bayesian network, and to determine whether the limb movements of the trainee have completed the interaction behavior that meets the training requirements for the target entity object.

2. The interactive autism training system according to claim 1, characterized in that, The prompt module includes an AR spatial perspective projection unit, a multi-channel spatial audio unit, and a haptic feedback unit; The AR spatial perspective projection unit is used to anchor the visual objects of training actions / training elements onto the physical entity interaction area based on spatial calibration and planar reprojection algorithms, and realize the virtual-real fusion display. At the same time, it outputs the recognition result feedback object with dynamic visual enhancement. The visual enhancement strategy is adaptively adjusted according to the visual perception sensitivity of the training object. The multi-channel spatial audio unit is used to output training-guided auditory objects that match the spatial position of visual objects using customized TTS speech synthesis and spatial sound image localization technology. It also supports adapting the frequency, volume, speech rate and feedback interval of the audio according to the auditory perception threshold of the trainee, and outputs recognition result feedback objects with voice emotion. The tactile feedback unit is used to output tactile feedback when the trainee makes matching / non-matching actions through a vibration module with vibration frequency / intensity graded, thereby realizing a multimodal feedback closed loop of vision-auditory-tactile feedback.

3. The interactive autism training system according to claim 2, characterized in that, The TOF radar is a solid-state TOF radar with a sampling frame rate of ≥120fps and a point cloud resolution of ≥640×480. The TOF radar is equipped with a point cloud preprocessing unit, which uses Euclidean clustering and region growing algorithms to achieve precise segmentation of the palm and knuckle regions. It collects the three-dimensional spatiotemporal coordinate sequence of 21 fine skeletal points of the trainee's palm as corresponding gesture practice data, and eliminates the inter-frame drift of hand movements through inter-frame point cloud registration. The gesture practice recognition module pre-builds a 3D spatiotemporal feature template library of training actions. It uses an improved 3DCNN+LSTM hybrid network to extract spatiotemporal features from the 3D spatiotemporal coordinate sequence of fine skeletal points on the palm, generating a high-dimensional gesture feature vector. Then, it uses a two-dimensional matching algorithm of Euclidean distance and cosine similarity, combined with dynamic distance threshold constraints, to determine the matching degree between the gesture feature vector and the corresponding training action features in the feature template library. It also supports differentiated recognition and matching of fine gestures with one hand and two hands.

4. The interactive autism training system according to claim 2, characterized in that, The multimodal sensing fusion camera group includes a motion-sensing depth camera, a high-definition visible light camera, and an infrared interactive contact camera; A motion-sensing depth camera acquires a three-dimensional spatiotemporal coordinate sequence data stream of ≥25 refined skeletal points on the whole body of the trainee, and smooths and completes the trajectories of the skeletal points through Kalman filtering; High-definition visible light cameras acquire multi-scale fused image data of physical objects, which are then preprocessed by low-light enhancement and motion blur removal to serve as visual data of the physical objects. Infrared interactive contact cameras collect data on the contact area, contact timing, and contact pressure trend between the trainee's limbs and physical objects. The entity interaction behavior recognition module uses a cross-modal feature fusion network to deeply fuse the spatiotemporal features of limb skeleton points, the visual features of entity objects, and the interaction contact features between limbs and entities to generate a joint feature vector of limb-entity interaction. Then, it uses an attention mechanism to focus on key interaction features and judge the matching degree between the joint feature vector and the configured training actions / training elements. It also supports triple judgment of the temporal logic, contact accuracy, and action amplitude of limb-entity interaction.

5. The interactive autism training system for children according to claim 4, characterized in that, The limb spatial coordinate data stream is a microsecond-level timestamp three-dimensional spatiotemporal coordinate sequence data stream of ≥25 refined skeletal points on the whole body of the trainee, with a skeletal point tracking accuracy of ≤0.5mm; The limb trajectory modeling unit is based on an improved dynamic time warping algorithm with feature weighting and time window constraints. It combines a hidden Markov model to perform trajectory feature modeling and limb movement trajectory calculation on the spatial coordinate sequence data stream of skeletal points, and realizes probabilistic inference of movement time sequence to determine whether the limb movement trajectory coincides with the set training movement. The entity object recognition unit adopts a lightweight YOLOv9-Lite visual deep learning network that combines transfer learning and few-sample fine-tuning. The visual deep learning network is pre-trained and fine-tuned on a dedicated entity dataset for training self-care skills for children with autism, and supports incremental updates of model parameters on the edge. The proficiency of the movement is judged based on a four-dimensional quantitative evaluation model. The four-dimensional quantitative evaluation model integrates four dimensions: the proportion of matching accuracy when the trainee performs the movement according to the training movement, the change in the smoothness of the matching movement within a set time period, the temporal compliance of the movement execution, and the accuracy of the interaction between the limb and the entity. In addition, the LSTM time series prediction model is used to predict the trend of the movement proficiency, and the proficiency evaluation threshold is adaptively and dynamically adjusted according to the training curve of the trainee.

6. The interactive autism training system according to claim 1, characterized in that, It also includes a reporting module, a medical record module, and a privacy computing module; The reporting module is used to perform in-depth analysis of training data based on the judgment results of the gesture practice recognition module and / or entity interaction behavior recognition module, through machine learning clustering analysis and trend fitting algorithms. It generates a training report containing quantitative training data, proficiency trend analysis, interaction behavior characteristics, and intelligent diagnostic conclusions of training effect. Based on the diagnostic conclusions and combined with the trainee's personalized ability model, it generates customized optimization suggestions for subsequent training programs, including adaptive adjustments to training content, training difficulty, and training duration. The medical record module is used to store the basic information, personalized ability model, training report and training data of the trainees in a blockchain-based encrypted manner, and supports authorized hierarchical access on multiple terminals and standardized synchronization of training data in multiple centers. The privacy computing module is used to perform differential privacy desensitization and federated learning processing on all training data, enabling joint training and optimization of models across multiple devices and centers.

7. An interactive training method for children with autism, characterized in that, The interactive autism training system for children according to any one of claims 1-6 comprises: A personalized ability model is established based on the trainee's cognitive level, motor skills, and perceptual sensitivity. Training actions / elements corresponding to autism self-care skills are configured according to the personalized ability model. At the same time, dynamic adaptive parameter adjustment of training difficulty, training rhythm, and feedback mode is supported. Dynamic background suppression, palm point cloud segmentation, and 3D world coordinate system registration were performed on the hand area of ​​the trainees, and hand gesture practice data of palm and finger joints were collected. 3D spatiotemporal feature extraction, feature dimensionality reduction and similarity matching are performed on gesture practice data to determine its matching degree with the configured training actions / training elements, and the proficiency of the trainees' actions is evaluated through a multi-dimensional quantitative model. Collect the full-body refined skeletal point three-dimensional spatiotemporal coordinate data stream of the training subject, multi-scale enhanced image data of the entity object, and limb-entity interaction contact data as entity interaction behavior data; The spatiotemporal correlation features of limb-entity interaction behavior data are fused and interaction logic reasoning is performed to determine the matching degree with the configured training actions / training elements, and to evaluate the proficiency of the trainees' actions. Based on the individualized ability model of the trainee, the system outputs a training guidance object and a recognition result feedback object that are linked by visual, auditory and tactile multimodal interaction. It handles local real-time preprocessing of all sensor data, low-latency inference of AI models, privacy desensitization of training data, and collaborative scheduling of multiple modules.

8. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the computer program to implement the interactive autism training system as described in any one of claims 1-6.

9. A readable storage medium, characterized in that, The readable storage medium stores computer instructions for causing the computer to implement the interactive autism training system as described in any one of claims 1-6.

Citation Information

Patent Citations

  • System and method for assisting pairing training of autistic child by applying fine gesture recognition apparatus

    CN107168525A

  • Autism intervention training meta universe system and learning monitoring and personalized recommendation method

    CN117238448A