A Method and System for Recognizing Stereotyped Behaviors in Children with Autism Based on Deep Learning Algorithms

This method for identifying stereotyped behaviors in children with autism using deep learning algorithms leverages multi-view data acquisition and frequency-Kolmogrove-Arnold networks for feature extraction. By combining algebraic normalization layers and relative entropy constraints, it addresses the accuracy and efficiency issues in identifying stereotyped behaviors in children with autism, achieving highly efficient stereotyped behavior recognition.

CN121260424BActive Publication Date: 2026-04-03JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-05
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies lack objective and low-cost methods for identifying stereotyped behaviors in children with autism, and traditional diagnostic methods have failed to effectively incorporate deep learning algorithms, resulting in low diagnostic accuracy and efficiency.

Method used

A method for identifying stereotyped behaviors in children with autism based on deep learning algorithms is adopted. Through multi-view data collection and processing, frequency features of stereotyped behaviors are extracted using a frequency-Kolmogrove-Arnold network, and compressed encoding is performed through algebraic normalization layer. Combined with relative entropy constraint, the output probability distribution is improved to enhance the model's generalization ability and robustness.

Benefits of technology

It significantly improved the accuracy of identifying stereotyped behaviors in children with autism, alleviated the problem of computational complexity, and enhanced the model's generalization ability and recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121260424B_ABST
    Figure CN121260424B_ABST
Patent Text Reader

Abstract

This invention relates to the field of behavior recognition technology and provides a method and system for recognizing stereotyped behaviors in children with autism based on deep learning algorithms. The method includes: collecting stereotyped behavior data from multiple perspectives and performing standardization processing; selecting representative stereotyped behaviors from the multi-perspective dataset to obtain video clips; processing the dataset using a frequency-Kolmogrove-Arnold network to extract frequency features of stereotyped behaviors; generating feature maps through algebraic normalization layer compression encoding; and constraining the difference between the output probability distribution and the prior distribution by introducing relative entropy as a regularization term in the cross-entropy loss, using the asymmetry of relative entropy to distinguish gradients of different categories. This invention introduces frequency information into the KAN network, extracting frequency components from autistic behaviors to improve prediction accuracy; uses an algebraic normalization layer to alleviate computational complexity in autistic scenarios; and addresses the difficulty in recognizing numerous and complex stereotyped behaviors in autistic scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of behavior recognition technology, and in particular relates to a method and system for recognizing stereotyped behaviors in children with autism based on deep learning algorithms. Background Technology

[0002] The population of children with mental illnesses, such as autism, is large, but there is a severe shortage of professional doctors. Doctors mainly rely on scales for subjective rating, lacking objective and low-cost accurate assessment methods.

[0003] Current research on the diagnosis of mental illnesses is basically limited to posterior analysis of small sample data. The model is relatively simple, the database is small and not open source, there is a lack of corresponding analysis algorithms, and it is not well integrated with traditional methods and foundations for the diagnosis of mental illnesses. This is insufficient to support a complete logical chain of diagnosis and seriously limits the medical application value of the solutions. Summary of the Invention

[0004] The purpose of this invention is to provide a method for identifying stereotyped behaviors in children with autism based on deep learning algorithms, aiming to solve the problems mentioned in the background art.

[0005] The present invention is implemented as follows: a method for identifying stereotyped behaviors in children with autism based on deep learning algorithms, comprising the following steps:

[0006] Collect multi-perspective data on children's stereotyped behaviors, construct a multi-perspective dataset, and perform standardization processing;

[0007] Video segments were obtained by selecting representative stereotyped behaviors from a multi-view dataset;

[0008] Based on the frequency-Kolmogrove-Arnold network, frequency features of children with stereotyped behaviors are extracted from the dataset.

[0009] Feature maps are generated through compression encoding using an algebraic normalization layer;

[0010] By introducing relative entropy as a regularization term into the cross-entropy loss, the difference between the output probability distribution and the prior distribution is constrained, and the gradients of different categories are distinguished by the asymmetry of relative entropy.

[0011] Another objective of this invention is to provide a system for recognizing stereotyped behaviors in children with autism based on deep learning algorithms, for implementing a method for recognizing stereotyped behaviors in children with autism based on deep learning algorithms, comprising:

[0012] The data acquisition and processing unit is used to collect multi-perspective data on children's stereotyped behaviors, construct multi-perspective datasets, and perform standardized processing.

[0013] The analysis unit is used to select representative stereotyped behaviors from the multi-view dataset to obtain video segments;

[0014] The feature extraction unit is used to process datasets based on frequency-Kolmogrove-Arnold networks to extract frequency features of children exhibiting stereotyped behaviors.

[0015] Normalization units are used to compress and encode features through algebraic normalization layers to generate feature maps.

[0016] The backpropagation unit is used to constrain the difference between the output probability distribution and the prior distribution by introducing relative entropy as a regularization term in the cross-entropy loss, and to distinguish the gradients of different categories through the asymmetry of relative entropy.

[0017] The stereotyped behavior recognition method for autistic children based on deep learning algorithms provided in this invention introduces frequency information into the KAN network, effectively extracting the frequency components of autistic behaviors, thus significantly improving prediction accuracy compared to traditional methods. It uses an algebraic normalization layer to replace the traditional normalization layer, effectively alleviating the computational complexity in autistic scenarios by addressing the underlying concept of normalization. Furthermore, it introduces relative entropy as a regularization term in the cross-entropy loss, which improves the model's generalization ability and robustness by constraining the difference between the model's output probability distribution and the prior distribution. The asymmetry of relative entropy distinguishes gradients for different categories, solving the problem of difficulty in recognizing numerous and complex stereotyped behaviors in autistic scenarios in existing technologies. Attached Figure Description

[0018] Figure 1 A flowchart of a method for identifying stereotyped behaviors in children with autism based on deep learning algorithms, provided in an embodiment of the present invention;

[0019] Figure 2 A schematic diagram of a frequency-Kolmogrove-Arnold network provided for an embodiment of the present invention;

[0020] Figure 3 A schematic diagram of a single frequency-Kolmogrove-Arnold network module component provided in an embodiment of the present invention;

[0021] Figure 4 A schematic diagram comparing the hyperbolic tangent function and the algebraic saturation function provided in an embodiment of the present invention;

[0022] Figure 5 This is a schematic diagram illustrating the gradient calculation for different categories using the balanced entropy loss function provided in this embodiment of the invention.

[0023] Figure 6 A comparison of the forward and backward propagation computation times of the hyperbolic tangent function and the algebraic saturation function provided in this embodiment of the invention;

[0024] Figure 7 A comparison of the calculated heatmap (a) and the original image (b) provided for an embodiment of the present invention. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0026] The specific implementation of the present invention will be described in detail below with reference to specific embodiments.

[0027] like Figures 1 to 3 The flowchart shown is a method for identifying stereotyped behaviors in children with autism based on deep learning algorithms, provided in an embodiment of the present invention, and includes the following steps:

[0028] Step 1: Multi-view data acquisition and preprocessing: A four-camera acquisition system is set up in the hospital with doctors and children sitting opposite each other. The four perspectives are mutually verified to collect real children's data from multiple perspectives. The most obvious stereotypical behaviors are selected to construct a multi-view dataset.

[0029] Step 2, Stereotype Behavior Segment Analysis: To avoid irrelevant parts interfering with the constructed dataset, representative stereotype behavior clips are selected to obtain the final video segments. The training and test sets are then split according to the amount of data within each category. Specifically:

[0030] Following the format of the SSV2 dataset, 50 medical video cases were selected, and a total of 2258 action video clips were captured from four perspectives. The action targets were segmented, and each clip lasted 1-8 seconds, containing the complete process of a single action performed by a child, doctor, and parent. There were 7 action categories: turning head, pointing, playing with a toy airplane, pushing a toy cart, jumping with a toy, drinking water, and looking at people. The training set contained 1806 video clips, and the test set contained 452 video clips.

[0031] Step 3: Extracting Frequency Features of Stereotyped Behaviors in Children: Since stereotyped behavior analysis contains higher-level semantic information, the repetitiveness and persistence of stereotyped behaviors essentially quantify the key features of the behavior in the time dimension. Stereotyped behavior is a repetitive behavior, and it shows a strong correlation with the frequency of children's actions. Therefore, stereotyped behaviors are strongly correlated in both time and frequency. To better capture this correlation in stereotyped behavior videos, the traditional KAN network is augmented with Fourier transform calculations to better compute this correlation. Specifically:

[0032] The processed dataset is fed into the Frequency-Kolmogorov-Arnold Networks, which can sequentially generate temporal prediction information in both frequency and video dimensions. This information is placed in the last layer of the KAN module, and the remaining layers do not modify it. Combined with the algebraic normalization layer, the feature map is finally generated.

[0033] The formula for calculating the frequency-Kolmogrove-Arnold network is as follows:

[0034] (1);

[0035] , (2);

[0036] (3);

[0037] (4);

[0038] (5);

[0039] (6);

[0040] In the formula, To construct the basic spline functions for the network; It refers to the different outputs of different function types, among which... It is a self-learning hyperparameter matrix, specifically, it is a dimensionless matrix. A three-dimensional real field matrix, Represents the real number field. It is 392. The value of 4 represents the fitting dimension of the spline function; The output of the overall Kolmogrove-Arnold network, where the output of a single-layer Kolmogrove-Arnold network is... What is extracted is the frequency component of the data. ; It is through Iterative computation and presequence The calculated new sequence components, where and These are the Fourier transform and the inverse Fourier transform, respectively, where the independent variables are the preceding sequence. The new sequence obtained after Fourier transform is then multiplied by one dimension, and the multiplied portion is set to zero. This multiplication is called the process of multiplying the dimension and setting the multiplied portion to zero. In the frequency domain, this is equivalent to copying the original spectrum to the low-frequency part of the target spectrum, while keeping the high-frequency part as zero. It is equivalent to adding a rectangular window of the original frequency domain to the overall function. Since Freq-KAN relies on Fourier transform to extract frequency features, it has limited adaptability to non-periodic behavior (such as random actions). Windowing is the core technical means to solve the global limitations of Fourier transform in analyzing non-stationary signals (such as random actions). It captures the frequency components in the signal that change with time through localization analysis. The final output is obtained through a linear coding layer. Encode the final total sequence Received;

[0041] First, the video clips The overall part obtained by feeding into the network is given by formula (3), and the required function calculation method has been given by formulas (1) and (2). It can be seen that the spline function is composed of cosine functions, which further confirms the previous frequency-related part. The frequency-time related part is calculated using the classic Fourier transform, which is to convert the data into a combination of sine and cosine functions. It can be seen that the previous It is an extension of the Fourier transform, and its mathematical expression is shown in formula (4). However, the sequence after the Fourier transform is not fixed. In order to ensure that the sequence has the same dimension as before the transform, it is achieved by windowing, i.e., setting the high frequencies to 0, corresponding to the formula. This enhances the adaptability to aperiodic behavior (such as random actions), which is then subjected to an inverse Fourier transform and the expanded portion is set to zero to obtain the final sequence. Finally, the required result is obtained through linear encoding. ;

[0042] Step 4: Algebraic Normalization Layer Compression Encoding: The calculated feature maps are initially fed into the algebraic normalization layer compression encoding part, which can prevent gradient problems and alleviate overfitting.

[0043] ;

[0044] In the formula, The algebraic saturation function is the functional representation of the algebraic normalization layer, and its independent variable represents the characteristic map obtained after calculation. , and This represents the hyperparameters in an algebraic saturation function, which can be learned autonomously by the network. The magnitude of the function. Deviation in the descriptive function;

[0045] Traditional normalization formula (hyperbolic tangent function):

[0046] ;

[0047] Its independent variable represents the feature map obtained after calculation. , and This represents the hyperparameters in the normalization layer, which can be learned autonomously by the network. The magnitude of the function. The deviation of the expression function This represents the sample mean of the feature maps in the normalization layer. That is, the sample variance of the feature maps in the normalization layer, and the denominator is... This is a constant, a value added to the denominator for numerical stability, to avoid the denominator being 0, and the default value is 0.00001;

[0048] However, traditional methods are relatively complex and computationally intensive, requiring the calculation of the mean and variance of all images before normalization. Based on this, the dynamic hyperbolic tangent method addresses the core idea of ​​the normalization layer—extreme value compression—by using a function-driven normalization layer. This method replaces the traditional normalization method with the hyperbolic tangent function. However, because its calculation requires the exponent e, it is not well-suited for low-resource devices, hindering its widespread adoption in assisted diagnostics. Therefore, this invention proposes a function-driven approach. Compared to the hyperbolic tangent function, it does not require the traditional exponent e, effectively alleviating the computational bottleneck through power function calculation. Furthermore, the added hyperparameters effectively mitigate the gradient vanishing and exploding problems caused by traditional saturated activation functions. A comparison between the hyperbolic tangent function and algebraic saturation functions is shown below. Figure 4 As shown;

[0049] Step 5, Backpropagation based on entropy loss: Add relative entropy as a regularization term to the cross-entropy loss to constrain the difference between the model's output probability distribution and the prior distribution. The asymmetry of relative entropy is used to distinguish gradients for different categories.

[0050] Backpropagation based on entropy loss is as follows:

[0051] ;

[0052] ;

[0053] ;

[0054] ;

[0055] In the formula, The overall loss function consists of two parts: the cross-entropy loss function. With relative entropy loss function , For allocation coefficients, and These represent the true label and the predicted probability in the data, respectively. and This refers to the actual label values ​​for different categories. These are real labels for different categories. Through uniform distribution Label values ​​after uniform distribution of different categories It is the real label Through uniform distribution The subsequent label value, This indicates the control over uniform distribution and label smoothness.

[0056] Where the allocation coefficient It can be set to 0.9, where the relative entropy uses uniformly distributed labels. By smoothing the noise distribution of the labels, the model is allowed to assign a small number of probabilities to non-target categories. This is because the correct probability obtained by taking the derivative of the balanced entropy loss function with respect to the labels is still not much different from the original loss function. However, due to the introduction of a uniform distribution, the probability of error is evenly distributed for the relative entropy part, thereby improving the robustness of the model and accelerating the convergence of the network.

[0057] Figure 5 This is a schematic diagram illustrating the gradient calculation of the balanced entropy loss function for different categories;

[0058] A comparison of the forward and backward propagation times for hyperbolic tangent functions and algebraic saturated functions, as follows: Figure 6 As shown;

[0059] A comparison between the image processed using the method provided in this embodiment and the original image is as follows: Figure 7 As shown.

[0060] The method provided in this embodiment of the invention is tested against other model methods using metrics. Here, True Positive (TP) represents a true positive instance, which is the number of samples that are actually positive and correctly predicted as positive by the model; False Negative (FN) represents a false negative instance, which is the number of samples that are actually positive but incorrectly predicted as negative by the model; False Positive (FP) represents a false positive instance, meaning the number of samples that are actually negative but incorrectly predicted as positive by the model; and True Negative (TN) represents a true negative instance, which is the number of samples that are actually negative and correctly predicted as negative by the model.

[0061] Accuracy is one of the most fundamental evaluation metrics for classification models. It measures the percentage of correctly predicted samples out of the total number of samples. The formula for calculating accuracy is:

[0062] ;

[0063] Precision, also known as accuracy, represents the proportion of truly correct predictions made by a model. It reflects the probability that a positive sample identified by the model is actually a positive sample. The formula for calculating precision is:

[0064] ;

[0065] Recall, also known as the detection rate, refers to the proportion of samples correctly identified by the model out of all actual positive samples. It reflects the model's coverage of positive samples. The calculation formula is:

[0066] ;

[0067] Precision and recall are a pair of mutually restrictive metrics. In practical applications, a trade-off usually needs to be struck between the two. For example, in the medical diagnosis scenario, for the diagnosis of rare diseases, doctors may prefer to improve the recall rate and try to find all possible patients to avoid omissions. Therefore, it is very important to measure the balance between the two. In order to comprehensively evaluate the performance of these two metrics and find the best balance between them, it is necessary to calculate the F1 score.

[0068] The F1 score is the harmonic mean of precision and recall, which takes into account the balance between precision and recall. Its calculation formula is as follows:

[0069] ;

[0070] The comparison results are shown in Table 1:

[0071] Table 1

[0072]

[0073] It can be seen that the method provided by the embodiments of the present invention is better than some model methods in the prior art in many aspects.

[0074] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for identifying stereotyped behaviors in children with autism based on deep learning algorithms, characterized in that, Includes the following steps: Collect multi-perspective data on children's stereotyped behaviors, construct a multi-perspective dataset, and perform standardization processing; Video segments were obtained by selecting representative stereotyped behaviors from a multi-view dataset; Based on frequency-Kolmogrove-Arnold network processing of video clips, frequency features of children exhibiting stereotyped behaviors are extracted. The frequency features of children with stereotyped behaviors are compressed and encoded through an algebraic normalization layer to generate feature maps. By introducing relative entropy as a regularization term into the cross-entropy loss, the difference between the output probability distribution and the prior distribution is constrained, and the gradient of different categories is distinguished by the asymmetry of relative entropy. The step of extracting frequency features of children exhibiting stereotyped behaviors from video clips using a frequency-Kolmogrove-Arnold network specifically includes: The formula for calculating the frequency-Kolmogrove-Arnold network is: ; , ; ; ; ; ; In the formula, To construct the basic spline functions of the network, It refers to the different outputs of different function types, among which... It is a hyperparameter matrix that can be learned autonomously. Represents the real number field. It is 392. The value of 4 represents the fitting dimension of the spline function; The output of the overall Kolmogrove-Arnold network, where the output of a single-layer Kolmogrove-Arnold network is... What is extracted is the frequency component of the data. , It is through Iterative computation and presequence The calculated new sequence components, This indicates that the dimension of the sequence used as the independent variable will be doubled. Indicates Fourier transform, Indicates the inverse Fourier transform. The final output is obtained through a linear coding layer. Encode the final total sequence The obtained data represents the frequency characteristics of children exhibiting stereotyped behaviors. The step of compressing and encoding the frequency features of stereotyped behaviors in children through an algebraic normalization layer to generate feature maps is as follows: The algebraic normalization layer is: ; In the formula, The algebraic saturation function is the functional representation of the algebraic normalization layer. Its independent variable represents the frequency characteristics of stereotyped behaviors in children obtained after calculation. , and This represents the hyperparameters in an algebraic saturation function, which can be learned autonomously by the network. The magnitude of the function. Deviation in the expression function.

2. The method for identifying stereotyped behaviors in autistic children based on deep learning algorithms according to claim 1, characterized in that, In the steps of collecting multi-view children's stereotyped behavior data, constructing a multi-view dataset, and performing standardization processing, multi-view children's stereotyped behavior data is collected based on cameras distributed in four directions of space.

3. The method for identifying stereotyped behaviors in autistic children based on deep learning algorithms according to claim 1, characterized in that, The step of introducing relative entropy as a regularization term into the cross-entropy loss to constrain the difference between the output probability distribution and the prior distribution, and to distinguish different categories of gradients through the asymmetry of relative entropy, specifically involves: Backpropagation based on entropy loss is as follows: ; ; ; ; In the formula, The overall loss function consists of two parts: the cross-entropy loss function. With relative entropy loss function , For allocation coefficients, and These represent the true label and the predicted probability in the data, respectively. and This refers to the actual label values ​​for different categories. These are real labels for different categories. Through uniform distribution Label values ​​after uniform distribution of different categories It is the real label Through uniform distribution The subsequent label value, This indicates the control over uniform distribution and label smoothness.

4. A system for recognizing stereotyped behaviors in autistic children based on deep learning algorithms, used to implement the method for recognizing stereotyped behaviors in autistic children based on deep learning algorithms as described in any one of claims 1-3, characterized in that, include: The data acquisition and processing unit is used to collect multi-perspective data on children's stereotyped behaviors, construct multi-perspective datasets, and perform standardized processing. The analysis unit is used to select representative stereotyped behaviors from the multi-view dataset to obtain video segments; The feature extraction unit is used to process video clips based on a frequency-Kolmogrove-Arnold network and extract frequency features of children exhibiting stereotyped behaviors. The normalization unit is used to compress and encode the frequency features of children with stereotyped behaviors through an algebraic normalization layer to generate a feature map. The backpropagation unit is used to constrain the difference between the output probability distribution and the prior distribution by introducing relative entropy as a regularization term into the cross-entropy loss, and to distinguish the gradients of different categories through the asymmetry of relative entropy.

Citation Information

Patent Citations

  • Multi-aspect segmentation network implementation method for visible light medical image segmentation

    CN119180963A

  • Satellite network data anomaly detection method based on CNN and RWav-KAN

    CN119854795A