Autism behavior data processing method and device, computer equipment and storage medium
By receiving interaction data between subjects and pre-set digital humans, performing multimodal behavior analysis and deep learning recognition, and combining machine learning classification, the problems of subjectivity and low efficiency in early ASD screening technology are solved, achieving efficient and accurate ASD risk screening.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-05
- Publication Date
- 2026-04-10
AI Technical Summary
Existing early screening technologies for ASD suffer from problems such as strong subjectivity, insufficient objectivity, low efficiency, high cost, and low levels of automation and intelligence, making it difficult to achieve efficient and accurate screening for young children.
By receiving interaction data between subjects and pre-built digital humans, multimodal behavior analysis is performed, deep learning models are used for identification and analysis, and machine learning models are combined for overall state classification. Finally, interpretable analysis content is output using an interpretability processing mechanism.
It enables objective, rapid, and interpretable screening of ASD risk, improves the accuracy and efficiency of intelligent analysis of behavioral data, and provides reliable diagnostic reference.
Smart Images

Figure CN121834471A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer software technology, and in particular to a method, apparatus, computer equipment, and storage medium for processing autism behavioral data. Background Technology
[0002] Autism Spectrum Disorder (ASD) is a neurodevelopmental disorder characterized by impairments in social interaction and communication, restricted interests, and repetitive and stereotyped behaviors. It typically manifests early in life and has long-term effects on a child's cognitive development, social functioning, and quality of life. Numerous studies have confirmed that early identification and intervention for ASD are crucial for improving the prognosis of affected children. Therefore, developing an objective and practical early screening technology for ASD suitable for young children has significant clinical application value and social implications.
[0003] In recent years, with the continuous rise in the incidence of ASD, traditional screening technologies relying on manual clinical diagnosis have become insufficient to meet the practical needs of early, widespread, and efficient screening. Especially in the preschool stage, young children have limited language expression abilities and cooperativeness, and are generally sensitive to contact-based testing instruments, leading to numerous technical challenges in traditional screening processes. The convenience, accuracy, and accessibility of screening are significantly limited. Against this backdrop, how to obtain objective behavioral indicators reflecting the core characteristics of ASD through technological means in a natural, low-burden setting without relying on contact instruments, and to achieve automated screening for ASD risk, has become a key technical challenge that urgently needs to be solved in this field.
[0004] Currently, early screening and assessment of ASD mainly rely on three types of technical methods: First, manual assessment techniques based on clinical experience, where professionals use clinical observation, structured interviews, and other manual methods to comprehensively judge children's social interactions, language behaviors, and stereotyped behaviors. This technique is highly dependent on the professional experience and subjective judgment of the assessors and lacks unified technical judgment standards, requiring a high degree of consistency among different assessors. Second, screening techniques based on parent reports and behavioral scales, which indirectly obtain data on children's daily behavioral performance by having parents fill out screening questionnaires or scales. This method is relatively simple to operate and can be used for large-scale initial screening, but the screening results are easily affected by non-technical factors such as parents' cognitive level, subjective understanding, and recall bias, resulting in insufficient objectivity. Third, manual assessment techniques based on standardized assessment tools, which use specially designed standardized tasks or assessment processes to manually assess children's language, cognitive, and social abilities. This technique usually requires a long assessment time and has high requirements for the testing environment and the professional training of assessors, making it difficult to achieve large-scale promotion and application.
[0005] In general, existing early screening technologies for ASD primarily rely on manual questionnaires, interviews, and observation as their core technical approaches. A standardized screening system based on multimodal objective behavioral data, capable of automated ASD risk analysis and assessment, has not yet been established. These technologies generally suffer from the following technical deficiencies: First, they exhibit strong subjectivity and insufficient objectivity. Most existing screening technologies depend on the experience of assessors or subjective reports from parents, lacking quantifiable and standardized objective behavioral data collection and analysis indicators, resulting in poor stability and repeatability of screening results. Second, screening efficiency is low and technology promotion costs are high. Most techniques require long-term, full-process professional involvement, making the screening process complex, time-consuming, and labor-intensive, hindering large-scale implementation in community or primary healthcare institutions. The challenges include: 1) limited application and difficulty in implementation; 2) low participation and cooperation among children; 3) traditional screening techniques are often simplistic, inefficient, and lack engaging elements, especially for younger children, which can lead to resistance or distraction, affecting the quality of screening data and the effectiveness of subsequent analysis; 4) low levels of automation and intelligence; existing screening technologies are not deeply integrated with artificial intelligence technologies such as machine learning and deep learning, lacking the ability to automatically collect, analyze, and intelligently screen children's behavioral data; even when some research or applications incorporate AI algorithms, their decision-making processes are often "black box" outputs, making it difficult to explain the basis for screening results and lacking clear explanations of specific behavioral manifestations and screening task contexts, resulting in limited clinical comprehensibility and acceptability. Therefore, existing technologies cannot simultaneously meet the clinical screening needs for objective, rapid, reproducible, and clearly causative screening, further highlighting the urgency of developing new early autism screening technologies that are both intelligent and interpretable. Summary of the Invention
[0006] This invention provides a method, apparatus, computer device, and storage medium for processing autism behavioral data, aiming to improve the intelligent analysis effect of behavioral data.
[0007] In a first aspect, embodiments of the present invention provide a method for processing autism behavioral data, including: The system receives raw behavioral signals input by the user and preprocesses the raw behavioral signals to obtain target behavioral data; wherein, the raw behavioral signals include interaction data between the subject and a preset digital human. Multimodal behavior analysis is performed on the target behavior data to obtain task fragments containing behavioral semantics; The task segments are identified and analyzed using a deep learning model, and the corresponding segment identification results are output. The overall state classification result is obtained by using a machine learning model to classify the segment recognition results. The overall state classification results are analyzed using an interpretability processing mechanism, and the corresponding interpretability analysis content is output.
[0008] Secondly, embodiments of the present invention provide an autism behavior data processing device, comprising: A signal processing unit is used to receive raw behavioral signals input by the user and preprocess the raw behavioral signals to obtain target behavioral data; wherein, the raw behavioral signals include interaction data between the subject and a preset digital human; The behavior analysis unit is used to perform multimodal behavior analysis on the target behavior data to obtain task fragments containing behavioral semantics; The segment recognition unit is used to identify and analyze the task segments using a deep learning model and output the corresponding segment recognition results. The state classification unit is used to perform an overall state classification judgment on the segment recognition result using a machine learning model, and obtain the overall state classification result. The result interpretation unit is used to perform interpretability analysis on the overall state classification result using an interpretability processing mechanism, and output the corresponding interpretability analysis content.
[0009] Thirdly, embodiments of the present invention provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the autism behavior data processing method as described in the first aspect.
[0010] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the autism behavior data processing method as described in the first aspect.
[0011] This invention provides a method, apparatus, computer device, and storage medium for processing autism behavioral data. The method includes: receiving raw behavioral signals input by a user and preprocessing the raw behavioral signals to obtain target behavioral data; wherein the raw behavioral signals include interaction data between the subject and a pre-set digital human; performing multimodal behavioral analysis on the target behavioral data to obtain task segments containing behavioral semantics; using a deep learning model to identify and analyze the task segments and output corresponding segment identification results; using a machine learning model to classify the segment identification results into an overall state to obtain an overall state classification result; and using an interpretability processing mechanism to perform interpretive analysis on the overall state classification result and output corresponding interpretive analysis content. This invention systematically integrates core technical aspects such as the collection, preprocessing, analysis, identification, and interpretation of multimodal interaction data through standardized end-to-end technical operations, effectively improving the accuracy, efficiency, and interpretability of intelligent behavioral data analysis, and significantly optimizing the intelligent analysis effect of behavioral data. Attached Figure Description
[0012] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 A flowchart illustrating a method for processing autism behavioral data according to an embodiment of the present invention; Figure 2 This is a schematic block diagram of an autism behavior data processing device provided in an embodiment of the present invention. Detailed Implementation
[0014] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0015] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0016] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0017] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0018] Please see below. Figure 1 This invention provides a method for processing autism behavioral data, specifically including steps S101 to S105.
[0019] Step S101: Receive the raw behavioral signal input by the user and preprocess the raw behavioral signal to obtain target behavioral data; wherein, the raw behavioral signal includes the interaction data between the subject and the preset digital human; Step S102: Perform multimodal behavior analysis on the target behavior data to obtain task fragments containing behavioral semantics; Step S103: Use a deep learning model to identify and analyze the task segments, and output the corresponding segment identification results; Step S104: Use a machine learning model to classify the overall state of the segment recognition results to obtain the overall state classification result; Step S105: Perform interpretability analysis on the overall state classification results using an interpretability processing mechanism, and output the corresponding interpretability analysis content.
[0020] In this embodiment, the raw behavioral signals generated by the interaction between the subject and the preset digital human are first received, and the raw behavioral signals are transformed into standardized target behavioral data through preprocessing. Then, multimodal behavioral analysis is performed on the target behavioral data to break it down into task segments with clear behavioral semantics. Next, a deep learning model is used to identify and analyze the task segments, and a machine learning model is used to classify the overall state of the segment identification results, outputting the segment identification results and the overall state classification results respectively. Finally, an interpretability processing mechanism is used to perform interpretive analysis on the overall state classification results, and the corresponding interpretive analysis content is output.
[0021] This embodiment systematically integrates core technologies such as the acquisition, preprocessing, analysis, identification, and interpretation of multimodal interaction data through standardized end-to-end technical operations. This effectively improves the accuracy, efficiency, and interpretability of intelligent behavioral data analysis, significantly optimizing the results. In practical applications, the autism behavioral data processing method provided in this embodiment can objectively and quantitatively uncover hidden features and correlation patterns in subject behavioral data, providing reliable technical references for users to conduct related diagnostic work. This helps improve the objectivity and efficiency of diagnostic work, while avoiding the problems of strong subjectivity and insufficient data utilization in traditional methods.
[0022] In practical applications, the original behavioral signals can be obtained by collecting the subject's reactions to the presented digital human image, the hide-and-seek game scenario, and the digital human's emotional expressions. During the data collection process, the digital human can be controlled to switch between different emotional states according to control commands, achieving multi-form, controllable scene and emotion display. Specifically, the interactive flow of the hide-and-seek game can be controlled, adjusting the digital human to present multiple preset emotional states (facial expressions) during the game, and setting a continuous display time of no less than 5 seconds for each emotional state. Through the above settings, the subject can obtain sufficient emotional stimulation duration during the hide-and-seek interaction, thereby stably inducing their visual attention, fixation patterns, and motor responses to different emotional information, providing reliable data support for subsequent analysis of emotional visual processing features.
[0023] In one embodiment, receiving the raw behavioral signal input by the user and preprocessing the raw behavioral signal to obtain target behavioral data includes: The original behavioral signals are separated and labeled based on a preset task event labeling strategy to obtain multiple modal signals; wherein, the task event labeling includes the time points of the appearance and disappearance of the digital human, the spatial location information of the digital human, the current emotion category of the digital human, and the start and end time windows of each emotion category; Based on the physical characteristics of different modal signals, corresponding filtering and denoising algorithms are used to filter and denoise each modal signal. Standardize the modal signal after filtering and denoising. Based on the start and end time windows of the emotion categories, the standardized modal signals are time-synchronized to obtain standardized behavioral data fragments corresponding to each emotion category. The normalized behavior data fragment is output as the target behavior data.
[0024] After obtaining the raw behavioral signals, this embodiment requires systematic preprocessing. The first step is the raw behavioral signal separation and labeling step. This step can separate and label the collected raw behavioral signals based on the task event labeling information output by the experimental control-related unit. The task event labeling includes at least the time points of the digital human's appearance and disappearance, the spatial location information of the digital human, the current emotion category of the digital human, and the start and end time windows of each emotion (the duration of a single emotion is not less than 5 seconds). Through the above event labeling, the continuously collected behavioral signals can be divided into multiple initial time segments that correspond one-to-one with specific task stages and emotional states. After signal separation and labeling, multimodal signal filtering and denoising are performed. This step employs appropriate filtering and denoising methods based on the physical characteristics of different modal signals: For eye-tracking signals, such as fixation coordinates and fixation duration in eye-tracking data, low-pass filtering or sliding window smoothing algorithms can be used to remove high-frequency noise introduced by momentary blinks and device jitter. Simultaneously, linear interpolation or spline interpolation can be used to repair short-term lost or abnormally abrupt data points. For head-movement signals, such as head posture angle, angular velocity, and angular acceleration signals, band-pass filtering can be used to remove high-frequency measurement noise and low-frequency drift components. For abnormal peak values caused by sudden large movements of the subject, a threshold detection mechanism is set to suppress or replace outliers. For body motion signals, such as displacement and velocity of key body points, smoothing is performed to eliminate random noise caused by non-task-related actions while retaining effective motion characteristics reflecting autonomous exploration and search behavior. After filtering and denoising, signal standardization and scaling are performed. To eliminate the impact of individual differences and inconsistent dimensions of different sampling devices on model training, the signals of each modality need to be standardized. For example, mean normalization or Z-score standardization can be used for continuous numerical features, time-related indicators (such as reaction time and head-turning delay) can be uniformly converted to millisecond or second scales, and spatial displacement or angle signals can be uniformly converted with reference to the same coordinate system. After standardization, behavioral data from different subjects, different modalities, and different emotional conditions can be compared and modeled in a unified feature space.
[0025] Subsequently, a multimodal time alignment step based on the emotion presentation time window is performed. After completing the single-modal preprocessing, the multimodal behavioral data is time-synchronized based on the start and end time windows of the emotion category. Specifically, the start time of the digital human's emotion presentation can be used as the time zero point. Eye movement, head movement, and body movement signals are resampled and the time axis is reconstructed. For signals with different sampling frequencies, interpolation or resampling methods are used to align them at the same time step. Multimodal data under the same emotional condition and within the same time window are combined into synchronized behavioral segments. In this way, multimodal, time-aligned, and standardized behavioral data segments that correspond one-to-one with different emotional states can be formed. In one embodiment, performing multimodal behavior analysis on the target behavior data to obtain a task fragment containing behavioral semantics includes: Based on the task event marking strategy, the target behavior data is processed by event-driven segmentation to divide the target behavior data into multiple initial task segments corresponding to task flow nodes. The initial task fragment is semantically segmented by behavioral semantic constraints to obtain an intermediate task fragment with behavioral semantics. Obtain the target modal signal corresponding to the intermediate task segment, and extract the corresponding behavioral features from the target modal signal; The intermediate task fragments and the behavioral features are combined into a task fragment containing behavioral semantics.
[0026] This embodiment uses a task event labeling strategy to divide the continuous behavioral data of the subject during the digital human hide-and-seek interaction into multiple task segments with clear behavioral semantics. Specifically, firstly, based on the input task structure information (i.e., the task event labeling strategy), the continuous behavioral data is processed in an event-driven segmentation manner. For example, starting with the digital human's appearance, behavioral data within the corresponding time window is extracted to form a behavioral segment for the digital human's appearance phase; starting with the digital human's disappearance, behavioral data within the corresponding time window is extracted to form a behavioral segment for the search phase after the digital human's disappearance; simultaneously, according to the task setting of the digital human appearing in different spatial locations, the behavioral data is further labeled as task segments for the corresponding spatial directions, and combined with the different emotional categories presented by the digital human, the behavioral segments are associated with emotional states. In this way, continuous behavioral signals can be divided into multiple initial task segments that correspond one-to-one with task flow nodes, completing the initial segmentation of behavioral data.
[0027] Secondly, based on the initial behavioral segmentation and acquisition of initial task fragments, behavioral semantic constraints are introduced to further subdivide the initial task fragments, ensuring that each fragment possesses clear behavioral semantics. For example, in the digital human appearance stage, behavioral data is divided into pre-reaction and post-reaction stages based on whether the subject exhibits gaze shift, head rotation, or body orientation behavior; in joint attention-related tasks, behavior before and after reaction delay is segmented based on the time point when gaze or head movement first points to the digital human target; in emotion expression processing-related tasks, viewing behavior fragments of the subject's facial expressions on the digital human are extracted using the time window for stable emotion presentation as the boundary; and in the digital human disappearance stage, behavioral data is divided into search initiation and search duration stages based on whether the subject exhibits active search behavior.
[0028] Then, for each intermediate task segment obtained after fine segmentation, corresponding behavioral features are constructed from signals of different modalities to meet the input requirements of subsequent models. This feature construction process covers at least three modalities: first, constructing feature indicators reflecting visual attention allocation and emotional processing from eye-tracking signals; second, constructing feature indicators reflecting head movement volume, head movement stereotypes, orientation response, and joint attention ability from head movement signals; and third, constructing feature indicators reflecting body movement volume, body movement stereotypes, exploratory behavior, and autonomous search ability from body movement signals. Furthermore, the constructed features can be directly retained in time series form or statistically summarized within the segments, thus flexibly adapting to the input requirements of different subsequent models.
[0029] After the above processing steps, the corresponding behavior analysis results can be obtained. These results mainly include multiple behavior task segments divided according to task structure, spatial orientation, and emotional state. Each task segment contains time-aligned multimodal behavior data and its corresponding behavioral semantic labels. These output task segments can be directly used as sequence inputs for deep learning models or as feature inputs for subsequent machine learning classification models, providing standardized and effective data support for subsequent recognition and analysis.
[0030] In one embodiment, the step of performing multimodal behavior analysis on the target behavior data to obtain a task fragment containing behavioral semantics further includes: Based on different modalities, time-series behavioral data are constructed for each task segment; wherein, the task segments include joint attention segments, emotion and expression processing segments, and target disappearance and active search segments; the modalities include eye-tracking modality, head-tracking modality, and body-movement modality.
[0031] In this embodiment, the task segments include at least three categories, and the specific definitions of each category of segments are clarified in conjunction with the digital human hide-and-seek task scenario as follows: The first is the shared attention segment, which specifically refers to the behavioral data segment corresponding to the subjects' gaze, head turning, and body reaction behavior when the digital human appears in different spatial locations or issues prompts; the second is the emotion and expression processing segment, which specifically refers to the behavioral data segment corresponding to the subjects' visual attention behavior towards the facial area and emotional information when the digital human presents different emotional expressions and maintains a preset display duration; the third is the target disappearance and active search segment, which specifically refers to the behavioral data segment corresponding to whether the subjects exhibit gaze, head movement, or body movement behavior in actively searching for the target after the digital human temporarily disappears. In each task segment, time-series behavioral data needs to be constructed from different modalities to provide standardized input for subsequent deep learning analysis. This includes, but is not limited to: eye-tracking data (the changes in eye fixation points along the x-axis, y-axis, and xy-plane); head-movement modal data (the time series of head movements along three axes); and body-movement modal data (the time series of movements at various joints of the body). This multimodal behavioral data will be input into subsequent deep learning analysis steps in time-series format to support the subsequent recognition and analysis work. It should be noted that the shared attention segments, emotion and expression processing segments, and target disappearance and active search segments are all automatically generated based on the preset task events and their time structure in the digital human hide-and-seek task, requiring no manual intervention and ensuring the objectivity and standardization of segment division. Correspondingly, in practical applications, after receiving the aforementioned task event information, for joint attention segments, when the digital human appears or issues a prompt at a preset spatial location, the time of the digital human's appearance or the prompt is taken as the segment start time, and the end time of the corresponding task phase is taken as the segment end time. The eye movement, head movement, and body movement data collected within this time interval are collectively divided into joint attention segments. This segment is mainly used to analyze the subject's behavioral performance under joint attention task conditions and to explore their joint attention-related behavioral characteristics. For emotion and expression processing segments, when the digital human enters the emotion and expression presentation phase, the time when the emotion and expression begin to be presented is taken as the segment start time. The segment ends at the point when the emotional expression ends. The multimodal behavioral data collected within this time interval is divided into emotional and facial expression processing segments, and corresponding emotional category labels are associated with them. This facilitates subsequent analysis of the subjects' visual attention and emotional processing characteristics under different emotional conditions. For the target disappearance and active search segments, when the digital human enters the disappearance phase, the segment starts at the point when the digital human disappears and ends at the end of the preset search time window. The multimodal behavioral data collected within this time interval is divided into target disappearance and active search segments. These segments are used to analyze the subjects' exploration and search behavior characteristics under the target disappearance condition.
[0032] After completing the above three-category task segment division, all task segments can be uniformly labeled to ensure that the attributes of each task segment are clear and traceable, which facilitates subsequent model analysis and data retrieval. For example, each task segment should at least include: task segment type label (clearly distinguishing the three types of segments), digital human spatial orientation information, digital human emotion category label, and segment start and end time information.
[0033] In one embodiment, the step of using a deep learning model to identify and analyze the task segments and outputting the corresponding segment identification results includes: The time-series behavioral data is input into a deep learning network structure; The time-series behavioral data is feature-encoded using the deep learning network structure. The deep learning network structure outputs a correlation prediction result reflecting the relationship between the corresponding task segment and the preset behavior pattern, and sets the correlation prediction result as the corresponding segment recognition result.
[0034] This embodiment provides a deep learning network structure for modeling and analyzing multimodal time-series behavioral data to achieve segment-level recognition of subject behavior. In practical applications, the specific workflow is divided into a model training phase and a model inference phase, specifically: During the model training phase, the deep learning network structure is systematically trained using pre-defined labeled samples. The aim is to enable the model to accurately learn the mapping relationship between subject behavior patterns and autism characteristics in different task segments, ensuring the accuracy of subsequent inferences. During model training, labeled multimodal time-series behavioral data is first input into the pre-defined deep learning network structure. Then, the deep learning network structure encodes the input time-series behavioral data, deeply mining and learning the inherent patterns of behavior changes over time in different task segments. Next, based on the error between the model's predicted results and the true labels of the samples, the corresponding loss function is calculated. Then, through backpropagation, the parameters of the deep learning model are iteratively updated based on the value of the loss function. This process of data input, feature encoding, loss calculation, and parameter updating is repeated until the model converges or reaches the preset training termination condition. Through this complete training process, the deep learning network structure gains the ability to automatically identify the correlation between subject behavior patterns and autism characteristics in different task segments. After model training is complete, the model inference phase begins. This phase primarily involves segment-level identification of the behavioral data of the subjects to be screened. During the inference phase, the deep learning network receives the input time-series data and performs inference operations on each task segment. For example, it first encodes and extracts targeted features from the multimodal time-series data corresponding to the task segment, uncovering hidden behavioral features within the segment. Then, based on the extracted behavioral features, it outputs a prediction result reflecting the correlation between the task segment and autism behavioral patterns. This prediction result can be at least one of probability values, risk scores, or binary classification results to flexibly adapt to subsequent analysis needs. Through these inference operations, accurate segment-level identification of the subjects' behavioral performance in different task segments is achieved. After the inference operation is completed, the corresponding recognition result can be output. In practical applications, the result includes at least: the autism-related recognition result for each task segment, the segment-level predicted probability or risk score for each segment, and the correlation information between the task segment type and the corresponding prediction result.
[0035] In one embodiment, the step of using a machine learning model to perform an overall state classification judgment on the segment recognition result to obtain an overall state classification result includes: Set the segment recognition results corresponding to the task segment as independent feature dimensions, and summarize the segment recognition results of the same type of task segments into segment type-level features; The features of each of the aforementioned fragment types are concatenated to construct an individual-level feature vector of uniform dimension; The individual-level feature vectors are input into a machine learning classification model for classification prediction, and the machine learning classification model outputs the corresponding overall state classification result.
[0036] This embodiment uses a machine learning classification model to integrate features and perform classification inference on the fragment recognition results, thereby achieving a classification judgment of the overall state of the subject. It completes a step-by-step decision-making process from "fragment-level recognition" to "individual-level classification," ultimately outputting the comprehensive screening results of the subject. It also consists of a model training phase and a model inference phase, specifically: The purpose of the model training phase is to enable the machine learning classification model to master the mapping relationship between individual-level feature vectors and different developmental state categories of subjects, providing support for accurate classification in the subsequent inference phase. In the model training phase, individual-level training samples are first constructed, i.e., individual-level sample data used to train the machine learning classification model. This individual-level sample data can be obtained by summarizing and integrating the segment-level recognition results of the same subject in all different task segments. Then, an individual-level feature vector reflecting the overall behavioral characteristics of the subject is further constructed. This involves using the segment recognition results corresponding to each type of task segment as independent feature dimensions to ensure that the recognition information of different types of segments is not confused. Statistical summarization of multiple segment recognition results for the same type of task segment is performed to form segment type-level features, integrating the overall performance of the same behavioral dimension. Finally, the segment type-level features corresponding to different task segment types are concatenated to construct a unified-dimensional individual-level feature vector. This feature vector comprehensively reflects the subject's overall behavioral performance under different task conditions, providing comprehensive feature support for subsequent classification inference.
[0037] The constructed individual-level feature vectors are input into a pre-defined machine learning classification model for systematic training. Specific steps include: pairing each individual-level feature vector with its corresponding true class label for the subject to ensure the accuracy of the training data; inputting the paired feature vectors and true labels into the machine learning classification model to start model training; calculating the error and iteratively updating the model parameters based on the difference between the model's output classification result and the sample's true class label. This process of data input, classification calculation, and parameter update is repeated until the model reaches the pre-defined convergence condition, completing the training. Through this complete training process, the machine learning classification model can accurately learn the mapping relationship between individual-level feature vectors and different developmental state categories of subjects, possessing individual-level classification judgment capabilities.
[0038] The model inference phase is primarily used to classify and judge the overall developmental status of the subjects to be screened. The specific steps correspond to those in the training phase, ensuring consistent processing logic and accurate results. For example, in the model inference phase, the subjects to be screened first undergo the same processing procedure as those in the training phase. This includes: acquiring multiple task segments generated by the subject during the screening process based on the same digital human hide-and-seek screening task and task event-driven segmentation method; calling the pre-trained deep learning segment-level recognition submodule to identify and analyze each task segment separately, outputting the corresponding segment-level recognition results; summarizing the segment-level recognition results of the subject across all task segments; and converting the summarized segment-level recognition results into an individual-level feature vector for the subject, following a construction method completely consistent with the model training phase, ensuring that the dimension and format of the feature vector are consistent with the training samples.
[0039] Then, the constructed individual-level feature vectors of the subjects to be screened are input into the trained machine learning classification model to perform individual-level classification inference operations. The specific process includes: the machine learning classification model comprehensively analyzes and classifies the input individual-level feature vectors; based on the mapping relationship learned during the model training phase, it outputs the specific classification result of the subject to be screened as belonging to autism spectrum disorder, developmental delay, or typical development; at the same time, it outputs at least one of the confidence level, probability value, or risk score corresponding to the classification result to help judge the reliability of the classification result and provide a reference for the subsequent screening conclusions.
[0040] In one embodiment, the step of using an interpretability processing mechanism to perform interpretable analysis on the overall state classification results and outputting corresponding interpretable analysis content includes: Based on the overall state classification results, the target task segment with abnormal behavior type is obtained; The target task fragment is used to generate interpretable analysis content with task semantics and behavioral descriptions through an interpretability processing mechanism based on SHAP.
[0041] This embodiment uses an interpretability processing mechanism based on SHAP to perform interpretive analysis on the overall state classification results (i.e., the comprehensive screening results of the subjects) output by the machine learning classification model. This means that the main behavioral basis leading to the judgment result is given, and the specific manifestations of abnormal behavior are clarified, thereby achieving an accurate explanation of the reasons for the screening conclusion and making the screening results more understandable and referable.
[0042] In practical applications, to achieve accurate detection and interpretation of behavioral deviations, behavioral expectation patterns are predefined for three different task segments in the digital human hide-and-seek task. These patterns serve as reference standards for subsequent behavioral comparison and deviation identification. For example, the behavioral expectation pattern for the joint attention segment is: whether the subject gazes at the digital human within a reasonable time window; whether a head turn occurs consistent with the digital human's position; and whether the gaze and turning behaviors are consistent in direction, ensuring that the behavioral response matches the digital human's appearance. The behavioral expectation pattern for the emotion and expression processing segment is: whether the subject allocates primary visual attention to the digital human's facial area; whether the facial area gaze ratio reaches a preset range; and whether the gaze behavior is continuous and stable, thereby judging the subject's ability to process emotions and expressions. The behavioral expectation pattern for the target disappearance and active search segment is: after the digital human disappears, whether the subject exhibits active gaze search behavior; and whether it is accompanied by exploratory movements of the head or body, reflecting the subject's ability to maintain and explore social targets.
[0043] During the model inference phase, after the machine learning classification model outputs the overall state classification result of the subject, targeted analysis is performed on the subject's actual behavioral data in each task segment. This may include: first, extracting the subject's actual behavioral indicators in each task segment and comparing them one by one with the predefined expected behavioral patterns for the corresponding segment; then, based on the comparison results, identifying behavioral deviations occurring in the current segment and clarifying the specific type of behavioral abnormality. For example, a significantly low proportion of attention to the facial area, a gaze or head reaction delay exceeding a reasonable range, or a lack of proactive search-related behaviors. This step ultimately outputs a set of associations for "task segment × behavioral abnormality type," and labels the specific behavioral abnormalities present in each type of task segment.
[0044] To ensure the accuracy and relevance of the explanations and to avoid irrelevant behavioral deviations interfering with the causal explanations, after obtaining individual-level screening results and behavioral deviation sets, this embodiment further identifies which task segments exhibit the most concentrated and severe behavioral deviations among subjects identified as high-risk or with autism spectrum disorder (ASD). Priority is given to selecting task segments with high discriminative power during model judgment (i.e., segments that have the greatest impact on the overall classification results) as the core explanatory basis. The overall state classification results output by the machine learning model are bound one-to-one with the selected behavioral deviation information to ensure that the explanations are not arbitrarily generated but are highly consistent with the model's actual judgment path and logic, thereby enhancing the credibility of the explanations.
[0045] After completing the detection and association confirmation of behavioral deviations, the above-mentioned behavioral deviation information is combined with the task semantics through the interpretability processing mechanism based on SHAP to generate interpretive analysis content with clear task semantics and specific behavioral descriptions, so as to realize the popularization and precision of the cause of the judgment. Its forms include, but are not limited to: "In the emotion and expression processing segment, the subject did not focus visual attention on the digital human's facial area, showing a low proportion of facial attention, and this behavioral pattern is consistent with the characteristics of autism"; "In the joint attention segment, the subject's gaze at the location of the digital human and head turning response were significantly delayed, and a stable joint attention behavior was not formed"; "In the target disappearance stage, no obvious active search-related behavior was detected, indicating insufficient social goal maintenance and exploration ability."
[0046] Ultimately, based on the above, the complete output can include three parts to form a closed-loop screening feedback: first, the overall screening classification results of the subject (i.e., the determination of autism spectrum disorder, developmental delay, or typical development); second, the description of the task segments that contributed significantly to the screening results, clarifying which behavioral abnormalities in those segments are the core factors leading to the current screening results; and third, the description of the main behavioral deviations corresponding to each key task segment, enabling relevant personnel to clearly understand the specific behavioral abnormalities of the subject and providing a clear direction for subsequent interventions.
[0047] Figure 2 This is a schematic block diagram of an autism behavior data processing device 200 provided in an embodiment of the present invention. The device 200 includes: The signal processing unit 201 is used to receive raw behavioral signals input by the user and preprocess the raw behavioral signals to obtain target behavioral data; wherein, the raw behavioral signals include interaction data between the subject and a preset digital human. Behavior analysis unit 202 is used to perform multimodal behavior analysis on the target behavior data to obtain task fragments containing behavioral semantics; The segment recognition unit 203 is used to identify and analyze the task segments using a deep learning model and output the corresponding segment recognition results; The state classification unit 204 is used to perform an overall state classification judgment on the segment recognition result using a machine learning model to obtain an overall state classification result. The result interpretation unit 205 is used to perform interpretability analysis on the overall state classification result using an interpretability processing mechanism and output the corresponding interpretability analysis content.
[0048] In one embodiment, the signal processing unit 201 includes: The separation and labeling unit is used to separate and label the original behavioral signals based on a preset task event labeling strategy to obtain multiple modal signals; wherein, the task event labeling includes the time points of the appearance and disappearance of the digital human, the spatial location information of the digital human, the current emotion category of the digital human, and the start and end time windows of each emotion category; The filtering and denoising unit is used to perform filtering and denoising processing on each modal signal according to the physical characteristics of different modal signals and the corresponding filtering and denoising algorithms. The standardization unit is used to standardize the modal signal after filtering and denoising. The time synchronization unit is used to perform time synchronization processing on the standardized modal signal based on the start and end time window of the emotion category to obtain the standardized behavioral data fragments corresponding to each emotion category. The data output unit is used to output the normalized behavior data fragment as the target behavior data.
[0049] In one embodiment, the behavior analysis unit 202 includes: The segmentation processing unit is used to perform event-driven segmentation processing on the target behavior data based on the task event marking strategy, so as to divide the target behavior data into multiple initial task segments corresponding to task flow nodes. A semantic segmentation unit is used to semantically segment the initial task fragments through behavioral semantic constraints to obtain intermediate task fragments with behavioral semantics. The feature extraction unit is used to acquire the target modal signal corresponding to the intermediate task segment and extract the corresponding behavioral features from the target modal signal; The feature combining unit is used to combine the intermediate task fragment and the behavioral features into a task fragment containing behavioral semantics.
[0050] In one embodiment, the behavior analysis unit 202 further includes: The data construction unit is used to construct time-series behavioral data for each task segment based on different modalities; wherein the task segments include joint attention segments, emotion and expression processing segments, and target disappearance and active search segments; and the modalities include eye-tracking modality, head-tracking modality, and body-movement modality.
[0051] In one embodiment, the fragment recognition unit 203 includes: The data input unit is used to input the time series behavioral data into the deep learning network structure; A feature encoding unit is used to encode features into the time-series behavioral data using the deep learning network structure. The result setting unit is used to output a correlation prediction result reflecting the relationship between the corresponding task segment and the preset behavior pattern through the deep learning network structure, and set the correlation prediction result as the corresponding segment recognition result.
[0052] In one embodiment, the state classification unit 204 includes: The feature aggregation unit is used to set the segment recognition results corresponding to the task segment as independent feature dimensions, and to statistically aggregate the segment recognition results of the same type of task segments into segment type-level features. The feature construction unit is used to concatenate the features of each of the aforementioned fragment types to construct an individual-level feature vector of a unified dimension. The classification prediction unit is used to input the individual-level feature vector into the machine learning classification model for classification prediction, and the machine learning classification model outputs the corresponding overall state classification result.
[0053] In one embodiment, the result interpretation unit 205 includes: An anomaly acquisition unit is used to acquire target task segments with abnormal behavior types based on the overall state classification results. The content generation unit is used to generate interpretable analysis content with task semantics and behavioral descriptions from the target task fragment through an interpretability processing mechanism based on SHAP.
[0054] Since the embodiments of the apparatus and the embodiments of the method correspond to each other, please refer to the description of the embodiments of the method for the embodiments of the apparatus, which will not be repeated here.
[0055] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed, can perform the steps provided in the above embodiments. The storage medium may include various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0056] This invention also provides a computer device, which may include a memory and a processor. The memory stores a computer program, and when the processor calls the computer program in the memory, it can implement the steps provided in the above embodiments. Of course, the computer device may also include various network interfaces, power supplies, and other components.
[0057] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to in the method section. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
[0058] It should also be noted that, in this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
Claims
1. An autism behavior data processing method, characterized by, The method comprises the following steps: Receiving raw behavior signals input by a user, and preprocessing the raw behavior signals to obtain target behavior data; wherein the raw behavior signals include interaction data between a subject and a pre-set digital human; Performing multi-modal behavior analysis on the target behavior data to obtain a task segment containing behavior semantics; Using a deep learning model to perform recognition analysis on the task segment, and outputting a corresponding segment recognition result; Using a machine learning model to perform overall state classification judgment on the segment recognition result, and obtaining an overall state classification result; Using an explainability processing mechanism to perform explainability analysis on the overall state classification result, and outputting corresponding explainability analysis content.
2. The autism behavior data processing method of claim 1, wherein, The step of receiving raw behavior signals input by a user, and preprocessing the raw behavior signals to obtain target behavior data comprises the following steps: Based on a pre-set task event marking strategy, the raw behavior signals are separated and marked to obtain multiple modal signals; wherein the task event marking includes the time points of appearance and disappearance of the digital human, the spatial orientation information of the digital human, the current emotion category of the digital human, and the start and end time windows of each emotion category; According to the physical characteristics of different modal signals, a corresponding filtering and denoising algorithm is used to filter and denoise each modal signal; The modal signals after filtering and denoising are standardized; Based on the emotion category start and end time window, the standardized modal signals are time-synchronized to obtain standardized behavior data segments corresponding to each emotion category; The standardized behavior data segments are output as the target behavior data.
3. The autism behavior data processing method of claim 2, wherein, The step of performing multi-modal behavior analysis on the target behavior data to obtain a task segment containing behavior semantics comprises the following steps: Based on the task event marking strategy, the target behavior data is subjected to event-driven segmentation processing to divide the target behavior data into multiple initial task segments corresponding to task flow nodes; The initial task segments are semantically divided by behavior semantic constraints to obtain intermediate task segments with behavior semantics; The target modal signals corresponding to the intermediate task segments are obtained, and corresponding behavior features are extracted from the target modal signals; The intermediate task segments and the behavior features are combined into a task segment containing behavior semantics.
4. The autism behavior data processing method of claim 3, wherein, The step of performing multi-modal behavior analysis on the target behavior data to obtain a task segment containing behavior semantics further comprises the following steps: Based on different modalities, a time series behavior data is constructed for each task segment; wherein the task segments include a common attention segment, an emotion and expression processing segment, a target disappearance and active search segment; and the modalities include an eye movement modality, a head movement modality, and a body movement modality.
5. The autism behavior data processing method of claim 4, wherein, The step of using a deep learning model to perform recognition analysis on the task segment, and outputting a corresponding segment recognition result comprises the following steps: The time series behavior data is input into a deep learning network structure; The time series behavior data is subjected to feature coding by the deep learning network structure; The deep learning network structure outputs a correlation prediction result between the corresponding task segment and the preset behavior mode, and sets the correlation prediction result as a corresponding segment recognition result.
6. An autism behavior data processing apparatus characterized by comprising: The method comprises the following steps: A signal processing unit is configured to receive a raw behavior signal input by a user and pre-process the raw behavior signal to obtain target behavior data. An interaction data of the user and a preset digital person is included in the raw behavior signal. A behavior analysis unit is configured to perform multi-modal behavior analysis on the target behavior data to obtain a task segment containing behavior semantics. A segment recognition unit is configured to perform recognition analysis on the task segment by using a deep learning model and output a corresponding segment recognition result. A state classification unit is configured to perform overall state classification and judgment on the segment recognition result by using a machine learning model to obtain an overall state classification result.
7. The autism behavior data processing apparatus according to claim 6, characterized by A result interpretation unit is configured to perform interpretive analysis on the overall state classification result by using an interpretability processing mechanism and output corresponding interpretive analysis content. The signal processing unit comprises: A separation and labeling unit is configured to separate and label the raw behavior signal based on a preset task event marking strategy to obtain multiple modal signals. The task event marking includes time points of appearance and disappearance of the digital person, spatial orientation information of the digital person, current emotion category of the digital person, and start and end time windows of each emotion category. A filtering and denoising unit is configured to filter and denoise each modal signal by using corresponding filtering and denoising algorithms according to physical characteristics of different modal signals. A standardization unit is configured to standardize the modal signals after filtering and denoising.
8. The autism behavior data processing apparatus according to claim 7, characterized by, A time synchronization unit is configured to perform time synchronization processing on the standardized modal signals based on the start and end time windows of each emotion category to obtain standardized behavior data segments corresponding to each emotion category. A data output unit is configured to output the standardized behavior data segments as the target behavior data. The behavior analysis unit comprises: A segmentation processing unit is configured to perform event-driven segmentation processing on the target behavior data based on the task event marking strategy to divide the target behavior data into multiple initial task segments corresponding to task flow nodes. A semantic division unit is configured to perform semantic division on the initial task segments by behavior semantic constraints to obtain intermediate task segments with behavior semantics.
9. A computer device, comprising: A feature extraction unit is configured to obtain target modal signals corresponding to the intermediate task segments and extract corresponding behavior features from the target modal signals. A feature combination unit is configured to combine the intermediate task segments and the behavior features into task segments containing behavior semantics. The method comprises a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the autism behavior data processing method according to any one of claims 1 to 5.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the autism behavior data processing method in any one of claims 1 to 5.