Skeleton-Based Autism Recognition with Hybrid Deep Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing ASD screening methods are time-consuming and lack measurable indicators, and traditional deep learning approaches struggle to accurately capture the complexity of social interactions and cognitive behaviors in young children due to limitations in temporal dynamic modeling and reliance on single biomarkers.

Innovation Solution

A hybrid deep learning system utilizing a Parent-Child Dyad Grouping Game protocol to create a comprehensive dataset, combined with a Two-stream Graph Attention Long Short-Term Memory (2sG-ALSTM) network architecture for recognizing autism by analyzing skeletal keypoint movements of children and parents through a high-resolution network and graph convolutional networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional auxiliary clinical screening methods are used, then ASD detection can be performed, but the process is time-consuming and lacks measurable indicators

Engineering Contradiction:
Improvemeasurable indicatorsVSAvoidtime-consuming
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces traditional manual clinical observation and assessment methods with an automated computer vision system that uses deep learning models to analyze video data. The system automatically extracts skeletal keypoints, processes motion data through LSTM networks, and generates ASD screening results without requiring lengthy clinician observation, thus providing measurable indicators efficiently

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service ASD screening by allowing parents to record their children playing with blocks at home using a camera or smartphone. The automated processing system then analyzes the video data independently without requiring clinician involvement in the observation process, significantly reducing time loss while maintaining measurement precision

Inventive Principle:
Principle #25Self-service

2Device complexity

If a single biomarker is used for ASD detection, then the detection process is simplified, but the complexity of social interactions and cognitive behaviors cannot be fully captured

Engineering Contradiction:
Improvedetection processVSAvoidcapture complexity of behaviors
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent employs a multi-modal composite approach by integrating multiple biomarkers including skeletal motion data, facial expressions, and interaction patterns. The system processes various types of behavioral data simultaneously through a unified deep learning framework, capturing the complexity of social interactions and cognitive behaviors more comprehensively than single biomarker methods while maintaining systematic processing

Inventive Principle:
Principle #40Composite materials

Solution Approach 2:

The system segments the complex behavioral analysis into distinct components: skeletal keypoint extraction, motion pattern recognition, interaction quality assessment, and temporal dynamic modeling. Each component is processed separately through specialized modules before being integrated for final ASD classification, allowing the system to handle complexity systematically without overwhelming the detection process

Inventive Principle:
Principle #1Segmentation

3Productivity

If traditional deep learning approaches are used, then ASD screening can be performed, but temporal dynamic modeling of complex actions shows limitations

Engineering Contradiction:
ImproveASD screening capabilityVSAvoidtemporal dynamic modeling
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements dynamic temporal modeling using Long Short-Term Memory (LSTM) networks that specifically capture the temporal dynamics of children's actions during block play. The system models the evolution of motor behaviors, interaction patterns, and engagement levels over time, providing precise measurement of temporal dynamics that traditional static deep learning approaches cannot achieve while maintaining screening productivity

Inventive Principle:
Principle #15Dynamics

4Productivity

If 3D Convolutional Neural Network (3DCNN) is used, then action recognition can be performed, but it cannot achieve higher performance due to difficulty in extracting useful information from background

Engineering Contradiction:
Improveaction recognition capabilityVSAvoidbehavioral activity recognition accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent extracts only the essential skeletal keypoint data from video frames, removing the distracting background information that confuses 3DCNN models. By converting video data into skeletal pose sequences and focusing computation on relevant motion patterns rather than entire video frames, the system achieves higher precision in behavioral activity recognition while maintaining efficient processing

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20250265816A1System, non-transitory storage medium and electronic device for recognizing autism based on hybrid deep learning
Publication Date: 2025.08.21 SHANDONG UNIV
  • US20250265816A1 patent drawing
  • US20250265816A1 patent drawing
  • US20250265816A1 patent drawing

AI summary

A system for recognizing autism based on hybrid deep learning includes a data acquisition module, a skeleton keypoint extraction module and a recognition and classification module. The data acquisition module is configured for obtaining a dataset based on a parent-child dyad block game protocol. The skeleton keypoint extraction module is configured for identifying a plurality of skeleton keypoints of a target and a position of each of the plurality of skeleton keypoints in the video data based on a high-resolution network to generate a skeleton sequence. The recognition and classification module is configured for classifying the child into autism spectrum disorder (ASD) children and typically developing (TD) children by inputting the skeleton sequence in a graph form into a Two-stream Graph Attention Long Short-Term Memory (2sG-ALSTM) network architecture.