Improved YOLOv11-based autism engraving behavior detection method
By improving the YOLOv11 algorithm combined with ADOS-2 standards, optimizing network structure and data enhancement technology, the automatic identification of stereotyped behaviors of autism is solved, the time-consuming and subjective problems of traditional diagnostic methods are improved, detection accuracy and efficiency are improved, and early diagnosis and treatment are supported.
Patent Information
- Application Number
- CN202510373127.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-01
AI Technical Summary
Traditional autism diagnosis methods rely on clinician observation and judgment, are time-consuming and susceptible to subjective factors, making it difficult to accurately identify and analyze the stereotyped behaviors of children with ASD, especially when the behavior is complex or concealed.
The improved YOLOv11 algorithm is adopted, combined with ADOS-2 diagnostic standards, and through data enhancement and network structure optimization, the automatic identification and analysis of autism stereotyped behavior is achieved, including the ADown downsampling module, the SCDown downsampling module and the SCSA attention mechanism to improve detection accuracy and efficiency.
It improves the accuracy and efficiency of detection of stereotyped behaviors of autism, can identify and assist diagnosis in early stages, provide objective basis for clinicians, and helps parents to detect abnormal behaviors in a timely manner in daily life, and promotes early intervention and treatment.
Smart Images

Figure CN120236332A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision deep learning, and particularly relates to a method for detecting stereotyped behaviors of autism based on improved YOLOv11. Background Art
[0002] Autism Spectrum Disorder (ASD) is a complex neurodevelopmental disorder, and its core symptoms include social interaction disorders, language and non-verbal communication disorders, as well as stereotyped repetitive behaviors and interests. In the diagnosis process of ASD, identifying and analyzing stereotyped behaviors is crucial because these behavior patterns are not only a significant sign of ASD but also an important basis for evaluating the severity of the child's condition and formulating treatment plans.
[0003] Currently, the gold standard for autism diagnosis is the semi-structured test commonly used by psychologists - the second edition of the Autism Diagnostic Observation Schedule (ADOS-2). ADOS-2 aims to evaluate children's performance in social interaction, communication, and stereotyped behaviors through a series of standardized activities and interactions. ADOS-2 observes that the stereotyped behaviors of ASD children usually manifest as a series of repetitive and mechanical actions or interests, which often lack functionality and are difficult to be interrupted or changed. These stereotyped behaviors may include, but are not limited to: repetitive hand movements, body movements, excessive fascination with specific objects, and self-stimulating behaviors. However, the traditional ADOS-2 test relies on the observation and judgment of clinicians, which not only takes a long time but may also be affected by the subjective factors of observers, resulting in limitations in the accuracy and consistency of the diagnosis results. In addition, as ASD children grow older and their condition develops, stereotyped behaviors may become more complex and concealed, posing greater challenges to clinicians' diagnosis.
[0004] To address these problems, researchers have been exploring the use of advanced computer vision and artificial intelligence technologies to assist in the diagnosis of ASD. These methods can automatically extract and analyze key behavioral characteristics by analyzing the behavioral performance of ASD children in videos, accurately and quickly identify stereotyped behaviors of autism, and provide more objective, accurate, and timely diagnostic basis for clinicians. In recent years, with the rapid development of computer vision technology, object detection and behavior recognition methods based on deep learning have been widely used in the medical field. As an advanced real-time object detection algorithm, YOLOv11 has the advantages of fast detection speed and high accuracy, providing new possibilities for the early detection of stereotyped behaviors of autism. Summary of the Invention
[0005] Based on the stereotyped behaviors in the ADOS-2 scale for the identification and early screening and diagnosis of children with autism spectrum disorder, the present invention proposes an autism stereotyped behavior detection method based on improved YOLOv11, focusing on the stereotyped behaviors of autism mentioned in ADOS-2, aiming to automatically identify and analyze key features such as the stereotyped behaviors of ASD children through computer vision technology to assist clinicians in making rapid and accurate diagnoses. This method combines the clinical experience of ADOS-2 and the advantages of the YOLOv11 algorithm, can automatically detect and analyze the stereotyped behavior characteristics in videos or pictures, improve the objectivity, accuracy and efficiency of diagnosis, and provide strong support for the early intervention and treatment of ASD children.
[0006] An autism stereotyped behavior detection method based on improved YOLOv11 of the present invention includes the following steps:
[0007] 1) Obtain autism stereotyped behavior videos, including the following methods:
[0008] 1.1 Download the public dataset SSBD of autism stereotyped behaviors from the Internet. The SSBD dataset only contains 3 types of autism stereotyped behavior videos, namely: butterfly hands, head banging, and spinning. The names of these three stereotyped behaviors are: "ArmFlapping", "HeadBanging", "Spinning";
[0009] 1.2 On the basis of the 3 types of stereotyped behavior videos included in the SSBD dataset, in order to expand the types and quantities of videos, collect 6 types of autism stereotyped behavior videos from the Internet, namely: butterfly hands, biting hands, pinching hands, walking on tiptoe, head banging, and spinning, and name them in sequence: "ArmFlapping", "Biting", "Finger", "Toe", "HeadBanging", "Spinning", and screen and integrate all the videos with the videos in the SSBD dataset;
[0010] 2) Establish an autism stereotyped behavior dataset, including the following steps:
[0011] 2.1 Process the videos collected in step 1). In order to ensure the complexity of the data for better training of the model, intercept pictures containing different backgrounds and 6 types of stereotyped behaviors in the videos as much as possible, and use the YOLO module in labelImg for annotation. The annotation labels are: "ArmFlapping", "Biting", "Finger", "Toe", "HeadBanging", "Spinning";
[0012] 2.2 Data augmentation operations such as random cropping, flipping, and color jittering are adopted to expand the diversity of the dataset. Specifically, the random cropping operation can change the cropping area of the image, enhancing the model's recognition ability for different parts of the object; image flipping helps the model learn the left-right symmetric object features; while color jittering simulates different lighting and environmental conditions by adjusting the brightness, contrast, and saturation of the image, making the dataset more abundant. For the augmented autism stereotypical behavior dataset, it is divided into a training set and a validation set according to a ratio of 7:3, and model training begins. Since issues related to the portrait rights and privacy of people are involved, the children appearing in the images are mosaicked.
[0013] 3) Autism stereotypical behavior detection based on the improved YOLOv11 includes the following steps:
[0014] YOLOv11 is an efficient, robust, and lightweight object detection system suitable for real-time object detection tasks. Through an improved network structure and optimization strategies, it significantly improves the inference speed and adaptability to complex scenarios while maintaining high accuracy. The C3k2 module is introduced in YOLOv11, which is an improvement of the C2f module in YOLOv8, enhancing the processing ability of shallow features. At the same time, the C2PSA mechanism is introduced to further optimize the processing flow of the feature map. Two depthwise separable convolutions DWConv are added to the detection head of YOLOv11, improving the perception ability of details and classification accuracy.
[0015] 3.1 In the Backbone network of YOLOv11, the original convolutional layers except the first layer are replaced with ADown convolutional layers. The input image first passes through the Backbone part, and features are gradually extracted and the spatial resolution is reduced through alternately connected ADown downsampling modules and C3k2 feature enhancement modules. The ADown module contains a 1×1 convolutional layer for channel compression (compression rate 0.5), a 3×3 deformable convolutional layer with a stride of 2 for downsampling operations, and finally, through batch normalization and the SiLU activation function, the nonlinear expression of features is enhanced.
[0016] 3.2 Replace the Neck network of YOLOv11 with a CCFM structure integrated with the SCDown module.
[0017] The feature map enters the Neck part, which contains the SCDown joint downsampling module and the CCFM structure for cross-layer feature splicing. The SCDown module performs spatial downsampling through a 3×3 convolution and channel downsampling through a 1×1 convolution to reduce feature redundancy. The CCFM structure enhances the correlation between features through cross-channel feature modulation, improving the model's representation ability for complex behaviors.
[0018] 3.3 Introduce the SCSA attention mechanism into the large target detection layer in front of the detection head;
[0019] SCSA consists of a shareable multi-semantic space attention (SMSA) module and a progressive channel self-attention (PCSA) module. By integrating multi-semantic information and effectively guiding channel recalibration, performance improvement is achieved;
[0020] SMSA module: First, use multi-scale, depth-shared one-dimensional convolutions to extract spatial information at different semantic levels from four independent sub-features, and accelerate model convergence through group normalization. Then, input the feature map modulated by SMSA into PCSA;
[0021] PCSA module: Combine progressive compression and channel-specific single-head self-attention mechanisms, and use the input-aware single-head self-attention mechanism to alleviate the semantic differences between different sub-features in SMSA and promote information fusion;
[0022] 3.4 Input the data into the model for detection to obtain the detection results;
[0023] The model can accurately and quickly identify which of the 6 types of stereotyped behaviors appears in the picture or video, and mark the specific location where the behavior occurs in the picture; The detection method of stereotyped behaviors can only play an auxiliary role in the process of autism diagnosis, and the final diagnosis still depends on the professional diagnosis of doctors;
[0024] 4) Model evaluation:
[0025] Model evaluation metrics include: mAP@0.5, frames per second (FPS), number of parameters (Params), floating-point operation count (FLOPs), confusion matrix, precision, recall, and F1 score;
[0026] Among them: mAP@0.5 indicates the value of mAP when the IOU threshold is 0.5. mAP is used to evaluate the average value of the average precision (AP) of multiple categories. The larger the value of mAP, the better the performance of the detection model; frames per second (FPS) indicates the number of frames that the model can process per second; the number of parameters (Params) and the number of floating-point operations (FLOPs) are important indicators to measure the computational complexity of the model. The fewer the number of parameters, the more lightweight the model; the confusion matrix is a summary of the target recognition prediction results, and the number of correct and incorrect predictions is obtained by counting; precision is the proportion of actual positive examples among the samples predicted as positive examples, indicating the accuracy of the model in positive example predictions; recall is the proportion of samples that are successfully predicted as positive examples among the actual positive examples, indicating the model's ability to recognize positive examples; the F1 score is the harmonic mean of precision and recall, with a value between 0 and 1, where 1 indicates the best performance and 0 indicates the worst performance; the mathematical expressions of Precision, Recall, and F1 are respectively:
[0027]
[0028] Among them: TP (True Positive) is the number of positive classes predicted as positive classes, that is, the number of positive class samples correctly predicted; TN (True Negative) is the number of negative classes predicted as negative classes, that is, the number of negative class samples correctly predicted; FN (False Negative) is the number of positive classes predicted as negative classes, that is, the number of positive class samples incorrectly predicted; FP (False Positive) is the number of negative classes predicted as positive classes, that is, the number of negative class samples incorrectly predicted.
[0029] The beneficial effects of the present invention are as follows: By combining the SCSA attention mechanism, the detection accuracy of stereotyped behaviors is significantly improved, and the recognition accuracy of stereotyped behaviors is enhanced. While ensuring the accuracy, by combining the ADown downsampling module and the SCDown joint downsampling module, the behavior recognition network can more efficiently and accurately recognize the stereotyped behaviors of autistic children, providing an auxiliary diagnosis tool for doctors. During the process of doctors' evaluation and diagnosis of autistic children, the model can assist in detection. When stereotyped behaviors are detected, it can attract the doctor's key attention; or when the amplitude of the autistic children's stereotyped behaviors is very small and not obvious enough to be easily ignored, it can also play a good auxiliary role. Even in daily life, when parents are suspicious of the abnormal behaviors of their children, the model can be used for detection. When the model determines that stereotyped behaviors occur, it can arouse the parents' attention and prompt them to go to the hospital in time to seek professional diagnosis and help from doctors, so as to intervene and treat in the early stage of discovery. It can be seen that the method for detecting autistic stereotyped behaviors based on the improved YOLOv11 plays a positive and important role in the process of assisting in the early diagnosis and treatment of autism. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a flowchart of the method for detecting autistic stereotyped behaviors based on the improved YOLOv11;
[0031] Figure 2 is a structural diagram of the network for detecting autistic stereotyped behaviors based on the improved YOLOv11;
[0032] Figure 3 is a recognition result diagram of the network for detecting autistic stereotyped behaviors based on the improved YOLOv11;
[0033] Among them: (a) is the ArmFlapping action, (b) is the Biting action, (c) is the Finger action, (d) is the Toe action, (e) is the HeadBanging action, and (f) is the Spinning action. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] The present invention will be described below with reference to the accompanying drawings.
[0035] As Figure 1 shown, a method for detecting autistic stereotyped behaviors based on the improved YOLOv11 of the present invention includes the following steps:
[0036] 1) Obtain videos of autistic stereotyped behaviors, including the following methods:
[0037] 1.1 Download the public dataset of autism stereotyped behaviors SSBD from the Internet. The SSBD dataset only contains 3 types of autism stereotyped behavior videos, namely: butterfly hands, head banging, and spinning. The names of these three stereotyped behaviors are: "ArmFlapping", "HeadBanging", "Spinning";
[0038] 1.2 On the basis of the 3 types of stereotyped behaviors included in the SSBD dataset, collect 6 types of autism stereotyped behavior videos from the Internet, namely: butterfly hands, biting hands, pinching hands, walking on tiptoe, head banging, and spinning, and name them in sequence as: "ArmFlapping", "Biting", "Finger", "Toe", "HeadBanging", "Spinning", and screen and integrate all the videos with the videos in the SSBD dataset;
[0039] 3) Establish an autism stereotyped behavior dataset, including the following steps:
[0040] 2.1 Process the videos collected in step 1), intercept pictures containing different backgrounds and 6 types of stereotyped behaviors in the videos, and use the YOLO module in labelImg for annotation. The annotation labels are: "ArmFlapping", "Biting", "Finger", "Toe", "HeadBanging", "Spinning";
[0041] 2.2 Adopt data augmentation operations such as random cropping, flipping, and color jittering to expand the diversity of the dataset. Divide the enhanced autism stereotyped behavior dataset into a training set and a validation set according to a ratio of 7:3, start model training, and perform mosaic processing on the children appearing in the images;
[0042] 3) Autism stereotyped behavior detection based on the improved YOLOv11, the network structure is as Figure 2 shown, including the following steps:
[0043] 3.1 In the Backbone network of YOLOv11, replace the original standard convolutional layers except the first layer with ADown convolutional layers. The input image first passes through the Backbone part, and gradually extracts features and reduces the spatial resolution through the alternately connected ADown downsampling module and C3k2 feature enhancement module; The ADown module contains a 1×1 convolutional layer, a 3×3 deformable convolutional layer, with a stride of 2, and finally enhances the non-linear expression of the features through batch normalization and the SiLU activation function;
[0044] 3.2 Replace the Neck network of YOLOv11 with a CCFM structure integrated with the SCDown module;
[0045] The feature map enters the Neck part, which contains the SCDown combined downsampling module and the CCFM structure for cross-layer feature splicing; the SCDown module performs spatial downsampling through 3×3 convolution and channel downsampling through 1×1 convolution; the CCFM structure enhances the correlation between features through cross-channel feature modulation;
[0046] 3.3 Introduce the SCSA attention mechanism in the large object detection layer before the detection head;
[0047] SCSA consists of a shareable multi-semantic space attention (SMSA) module and a progressive channel self-attention (PCSA) module, which integrates multi-semantic information and effectively guides channel recalibration;
[0048] SMSA module: First, use multi-scale, depth-shared one-dimensional convolution to extract spatial information at different semantic levels from four independent sub-features, and accelerate model convergence through group normalization. Then, input the feature map modulated by SMSA into PCSA;
[0049] PCSA module: Combine progressive compression and channel-specific single-head self-attention mechanism, and use the input-aware single-head self-attention mechanism to alleviate the semantic differences between different sub-features in SMSA and promote information fusion;
[0050] 3.4 Input the data into the model for detection, and the detection results are as Figure 3 shown;
[0051] The model can accurately and quickly identify which of the 6 types of stereotyped behaviors appears in the picture or video, and mark the specific location where the behavior occurs in the picture;
[0052] 4) Model evaluation:
[0053] The model evaluation metrics include: mAP@0.5, frames per second (FPS), number of parameters (Params), floating-point operation count (FLOPs), confusion matrix, precision, recall, and F1 score;
[0054] Among them: mAP@0.5 indicates the value of mAP when the IOU threshold is 0.5. mAP is used to evaluate the average value of the average precision AP of multiple categories. The larger the value of mAP, the better the performance of the detection model; Frames Per Second (FPS) represents the number of frames that the model can process per second; The number of parameters (Params) and the number of floating-point operations (FLOPs) are important indicators to measure the computational complexity of the model. The fewer the number of parameters, the more lightweight the model; The confusion matrix is a summary of the target recognition prediction results, and the number of correct and incorrect predictions is obtained by counting; Precision refers to the proportion of actual positive examples among the samples predicted as positive examples, indicating the accuracy of the model in positive example predictions; Recall refers to the proportion of samples that are actually positive examples and are successfully predicted as positive examples by the model, indicating the model's ability to identify positive examples; The F1 score is the harmonic mean of precision and recall, with a value between 0 and 1, where 1 represents the best performance and 0 represents the worst performance; The mathematical expressions for Precision, Recall, and F1 are respectively:
[0055]
[0056] Among them: TP is the number of positive classes predicted as positive classes, that is, the number of correctly predicted positive class samples; TN is the number of negative classes predicted as negative classes, that is, the number of correctly predicted negative class samples; FN is the number of positive classes predicted as negative classes, that is, the number of incorrectly predicted positive class samples; FP is the number of negative classes predicted as positive classes, that is, the number of incorrectly predicted negative class samples.
Claims
1. A method for detecting stereotyped behaviors of autism based on improved YOLOv11, characterized in that The following steps are involved: 1) Obtain videos of autism stereotyped behaviors, including the following methods: 1.1 Download the public dataset of autism stereotyped behaviors SSBD from the Internet. The SSBD dataset contains only three types of autism stereotyped behavior videos: butterfly hand, head banging, and spinning. The names of these three types of stereotyped behaviors are: "ArmFlapping", "HeadBanging", and "Spinning"; 1.2 Based on the three types of stereotyped behavior videos included in the SSBD dataset, six types of stereotyped behavior videos of autism were collected from the Internet, namely: butterfly hand, biting hand, pinching hand, tiptoe walking, head banging, and spinning, and named in sequence: "ArmFlapping", "Biting", "Finger", "Toe", "HeadBanging", and "Spinning", and all videos were screened and integrated with the videos in the SSBD dataset; 2) Establishing autism stereotyped behavior dataset, including the following steps: 2.1 Process the videos collected in step 1), capture pictures containing different backgrounds and 6 stereotyped behaviors in the videos, and use the YOLO module in labelImg to annotate them. The annotated labels are: "ArmFlapping", "Biting", "Finger", "Toe", "HeadBanging", and "Spinning"; 2.2 Random cropping, flipping and color jittering are used to enhance the data set to expand the diversity of the data set. The enhanced autism stereotyped behavior data set is divided into a training set and a validation set in a ratio of 7:
3. Model training is started and children appearing in the images are mosaicked. 3) Autism stereotyped behavior detection based on improved YOLOv11 includes the following steps: 3.1 In the Backbone network of YOLOv11, the original convolutional layers except the first layer are replaced by ADown convolutional layers. The input image first passes through the Backbone part, and then gradually extracts features and reduces the spatial resolution through the alternately connected ADown downsampling module and C3k2 feature enhancement module; the ADown module contains a 1×1 convolutional layer, a 3×3 deformable convolutional layer with a step size of 2, and finally enhances the nonlinear expression of features through batch normalization and SiLU activation function; 3.2 Replace the Neck network of YOLOv11 with the CCFM structure that integrates the SCDown module; The feature map enters the Neck part, which includes the SCDown joint downsampling module and the CCFM structure of cross-layer feature splicing; the SCDown module performs spatial downsampling through 3×3 convolution and 1×1 convolution channel downsampling; the CCFM structure enhances the correlation between features through cross-channel feature modulation; 3.3 Introduce the SCSA attention mechanism in the large object detection layer before the detection head; SCSA consists of a shared multi-semantic spatial attention SMSA module and a progressive channel self-attention PCSA module, which integrates multi-semantic information and effectively guides channel recalibration; SMSA module: First, multi-scale, depth-shared one-dimensional convolution is used to extract spatial information of different semantic levels from four independent sub-features, and group normalization is used to accelerate model convergence. Then, the feature map modulated by SMSA is input into PCSA. PCSA module: Combining progressive compression and channel-specific single-head self-attention mechanism, using input-aware single-head self-attention mechanism to alleviate the semantic differences between different sub-features in SMSA and promote information fusion; 3.4 Input the data into the model for testing and obtain the test results; The model can accurately and quickly identify which of the six types of stereotyped behaviors appear in the image or video, and mark the specific location where the behavior occurs in the image; 4) Model evaluation: Model evaluation indicators include: mAP@0.5, FPS, Params, FLOPs, confusion matrix, Precision, Recall, and F1 score. mAP@0.5 indicates the mAP value when the IOU threshold is 0.
5. mAP is used to evaluate the average value of the average accuracy AP of multiple categories. The larger the mAP value, the better the performance of the detection model. FPS indicates the number of frames that the model can process per second. Params and FLOPs are important indicators for measuring the computational complexity of the model. The fewer the number of parameters, the lighter the model is; the confusion matrix is a summary of the target recognition prediction results, counting the number of correct and incorrect predictions; precision refers to the proportion of samples predicted as positive examples that are actually positive examples, indicating the accuracy of the model in positive example prediction; recall refers to the proportion of samples that are actually positive examples that the model successfully predicts as positive examples, indicating the model's ability to recognize positive examples; the F1 score is the harmonic mean of precision and recall, with a value between 0 and 1, 1 indicating the best performance and 0 indicating the worst performance; the mathematical expressions of Precision, Recall and F1 are: Among them: TP is the number of positive classes predicted as positive classes, that is, the number of correctly predicted positive samples; TN is the number of negative classes predicted as negative classes, that is, the number of correctly predicted negative samples; FN is the number of positive classes predicted as negative classes, that is, the number of incorrectly predicted positive samples; FP is the number of negative classes predicted as positive classes, that is, the number of incorrectly predicted negative samples.
Citation Information
Cited By
Deep learning algorithm-based autistic child engraving behavior identification method and system
CN121260424A