A multimodal intelligent screening method, system, device and medium for childhood autism
By employing multimodal data fusion and phased assessment methods, this approach addresses the subjectivity and lag issues in existing autism screening technologies for children, enabling early and accurate autism screening that adapts to specific symptoms at different developmental stages and provides a fully intelligent screening solution.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2026-04-03
AI Technical Summary
Existing methods for predicting childhood autism are highly subjective, have significant time lags, and poor adaptability across different stages, making it difficult to meet the needs of clinical practice and large-scale rapid screening.
By establishing a mapping relationship between children's age and three attention periods—social interaction, play performance, and school adaptation—multimodal data is integrated for phased assessment. This includes analyzing audio, video, and behavioral modal data from family interaction videos to identify social characteristics such as name-calling responses and joint attention initiation; extracting text, video, and behavioral modal data from play videos to capture abnormalities such as performance deficits and stereotyped behaviors; and integrating homework texts, group discussion audio, and teacher interaction data from classroom scenes to assess characteristics of group activities such as the breadth of social relationships and instruction compliance. Finally, gender parameters are used to determine the level of loneliness.
It enables accurate and early autism screening, avoids the limitations of a single standard, improves the objectivity and accuracy of the assessment, shortens the screening cycle, adapts to the specific symptoms of children at different developmental stages, and provides a stable and reliable intelligent screening method throughout the entire process.
Smart Images

Figure QLYQS_8 
Figure QLYQS_23 
Figure QLYQS_24
Abstract
Description
Technical Field
[0001] This invention relates to the field of autism screening technology, specifically to a multimodal intelligent screening method, system, device, and medium for childhood autism. Background Technology
[0002] Autism significantly impacts children's social interaction, language communication, behavioral patterns, and cognitive development, and comorbidities can exacerbate developmental delays. Therefore, precise and early-intervention-enabled predictive technologies are urgently needed to achieve early detection and intervention for childhood autism.
[0003] Current methods for predicting childhood autism include behavioral observation scales, biomarker testing, imaging analysis, and genetic testing, but each method has significant limitations. Behavioral scales are highly subjective and exhibit significant time lag; biomarkers lack specificity; imaging methods are costly and have poor accessibility; and genetic testing struggles to cover sporadic cases. Existing technologies lack a unified analytical framework and accurate predictive capabilities, making it difficult to meet the needs of clinical practice and large-scale rapid screening. Summary of the Invention
[0004] The purpose of this invention is to provide a multimodal intelligent screening method, system, device, and medium for childhood autism.
[0005] The technical solution of this invention is as follows:
[0006] A multimodal intelligent screening method for childhood autism includes the following operations:
[0007] S1. Based on the mapping relationship between children's age information and attention periods, obtain the type of children's attention period; if the type of children's attention period is social interaction attention period, execute S2; if the type of children's attention period is play execution attention period, execute S3; if the type of children's attention period is school adaptation attention period, execute S4.
[0008] S2. Extract the first audio modality data, first video modality data, and first behavioral modality data from the children's family interaction video samples, perform fusion analysis to obtain the children's name-calling response information, joint attention initiation information, and social approach-avoidance information, calculate the children's autism prediction value, and execute S5;
[0009] S3. Extract the second text modality data, second video modality data, and second behavioral modality data from the children's play performance video samples. After modality aggregation processing, obtain performance defect information, action stuttering information, and stereotyped behavior information, calculate the autism prediction value of the children, and then proceed to S5.
[0010] S4. Based on the homework text modal data, group discussion audio modal data, and teacher interaction behavior modal data in children's classrooms, information on the breadth of social relationships, information on dialogue turn-taking delay, and information on compliance with collective instructions are obtained. After embedding processing, information interaction processing is performed to obtain a social activity feature vector. The social activity feature vector is then processed nonlinearly to obtain a predicted value for childhood autism, and S5 is executed.
[0011] S5. Based on the child's autism prediction value and gender, obtain the autism level and match the corresponding treatment plan.
[0012] In S2, the name-calling response information is the name-calling response rate, which is obtained based on the ratio of the total number of name-calling responses to the total number of names called in the text modal data, video modal data, and behavioral modal data; the joint attention initiation information is the number of joint attention initiations, which is obtained based on the number of children's proactive social behaviors outside the response scenario in the first audio modal data, and / or the first video modal data, and / or the first behavioral modal data; and the social approach-avoidance information is the number of social approach-avoidance events, which is obtained based on the number of times the child stops when approaching and the number of times the child turns when approaching in the first behavioral modal data.
[0013] In S3, execution defect information refers to the degree of incompleteness of actions during game execution; action stuttering information refers to the number of times actions are stuttered; and stereotyped behavior information refers to the number of times stereotyped behaviors are performed.
[0014] In S4, the social relationship breadth information is obtained from the entropy values of different number of names and title types in the children's classroom homework text modal data; the dialogue turn-off delay information is obtained from the children's group discussion audio modal data where the children's response time delay is greater than the time delay threshold; and the collective instruction compliance information is obtained from the number of times the children perform actions within the first time range after the teacher issues an instruction in the teacher's interaction behavior modal data.
[0015] In S2, the predictive value for childhood autism is calculated using the following formula: Q1=100·(1-(w c ·C+w a ·f a -w s ·f s Q1 is the predictor of childhood autism, C is the name-calling response rate, and f a To generate a contribution value for shared attention, f s Contribution value to social avoidance, w c For the call-to-name response weight, w a To initiate weights for shared attention, w s Social avoidance weight.
[0016] Common attention initiates contribution value f aIt is obtained through the following formula: when the number of times joint attention is initiated, a ≤ a base hour, When the number of times joint attention is initiated, a > a base hour, 'a' represents the number of times joint attention is initiated. ref As a reference value for the number of times common attention is initiated, a base To establish a baseline for the number of times common attention is initiated, w 1,a w 2,a These are the first initiation weights and the second initiation weights, respectively; the social avoidance contribution value f. s It is obtained through the following formula: when the number of social approach / avoidance attempts s≤s ref At that time, f s =0; when the number of social approach / avoidance attempts s>s ref hour, s b This serves as a baseline for the number of social avoidance events.
[0017] In S3, the operation to obtain motion lag information is as follows: the second row of modal data is divided into discrete motion segments, and the duration of each motion segment is obtained; if the duration of the current motion exceeds the duration threshold, and the current motion is different from the two adjacent motions, then the current motion is a lag motion; all lag motions are counted to obtain the number of motion lags, which is used as motion lag information.
[0018] In S4, the operation to obtain the social activity feature vector is as follows: the social relationship breadth information, the dialogue turn-taking delay information, and the collective instruction compliance information are embedded into vectors to obtain the social activity embedding vector; the social activity embedding vector is processed by a multilayer perceptron and an attention mechanism to obtain the social activity feature vector.
[0019] A multimodal intelligent screening system for childhood autism, used to implement the aforementioned multimodal intelligent screening method for childhood autism, includes:
[0020] The Child Attention Period Type Generation Module is used to obtain the child's attention period type based on the mapping relationship between the child's age information and attention periods. If the child's attention period type belongs to the social interaction attention period, the social interaction attention period prediction module is executed; if the child's attention period type belongs to the game execution attention period, the game execution attention period prediction module is executed; if the child's attention period type belongs to the school adaptation attention period, the school adaptation attention period prediction module is executed.
[0021] The social interaction attention period prediction module is used to extract first audio modality data, first video modality data, and first behavioral modality data from children's family interaction video samples, perform fusion analysis, obtain children's name-calling response information, joint attention initiation information, and social approach and avoidance information, calculate the prediction value of autism in children, and execute the autism level and treatment plan generation module;
[0022] The game execution attention period prediction module is used to extract second text modal data, second video modal data, and second behavioral modal data from children's game execution video samples. After modal aggregation processing, it obtains execution defect information, action stuttering information, and stereotyped behavior information, calculates the prediction value of childhood autism, and generates the autism level and treatment plan.
[0023] The school adaptation attention period prediction module is used to obtain information on the breadth of social relationships, the delay in dialogue turn-taking, and compliance with collective instructions based on homework text modal data, group discussion audio modal data, and teacher interaction behavior modal data in children's classrooms. After embedding and processing, information interaction processing is performed to obtain a social activity feature vector. The social activity feature vector is then processed nonlinearly to obtain a predicted value for childhood autism, and an autism level and treatment plan generation module is executed.
[0024] The loneliness level and treatment plan generation module is used to obtain the loneliness level based on the child's autism prediction value and gender, and match the corresponding treatment plan.
[0025] A multimodal intelligent screening device for childhood autism includes a processor and a memory, wherein the processor executes a computer program stored in the memory to implement the aforementioned multimodal intelligent screening method for childhood autism.
[0026] A computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the above-described multimodal intelligent screening method for childhood autism.
[0027] The beneficial effects of this invention are as follows:
[0028] This invention presents a multimodal intelligent screening method for childhood autism. By establishing a mapping relationship between a child's age and three attention periods—social interaction, play execution, and school adaptation—it achieves phased scenario data acquisition and accurate autism assessment. During the social interaction attention period, it integrates audio, video, and behavioral modal data from family interaction recordings to analyze social characteristics such as name-calling response and joint attention initiation. During the play execution attention period, it extracts text, video, and behavioral modal data from play recordings to capture abnormalities such as execution deficits and stereotyped behaviors. During the school adaptation attention period, it integrates homework text, group discussion audio, and teacher interaction data from classroom scenarios to assess the breadth of social relationships. This method identifies characteristics of autism, such as compliance with instructions, in group activities. Data from each stage is processed through feature extraction, modality aggregation, and nonlinear processing to generate autism prediction values, which are then combined with gender parameters to determine the degree of autism. By dynamically adapting to the core characteristics of children's developmental stages, this method avoids the limitations of single-standard screening. It enhances the objectivity of assessment by using multi-source modality data fusion and shortens the screening cycle with automated algorithms. This achieves full-process intelligentization from data collection to intelligent screening, effectively solving the problems of strong subjectivity, significant lag, and poor cross-stage adaptability of existing methods. It provides a stable, reliable, accurate, efficient, and comprehensive early screening method for autism that covers all age groups in clinical autism screening. Detailed Implementation
[0029] This embodiment provides a multimodal intelligent screening method for childhood autism, including the following operations:
[0030] S1. Based on the mapping relationship between children's age information and attention periods, obtain the type of children's attention period; if the type of children's attention period is social interaction attention period, execute S2; if the type of children's attention period is play execution attention period, execute S3; if the type of children's attention period is school adaptation attention period, execute S4.
[0031] S2. Extract the first audio modality data, first video modality data, and first behavioral modality data from the children's family interaction video samples, perform fusion analysis to obtain the children's name-calling response information, joint attention initiation information, and social approach-avoidance information, calculate the children's autism prediction value, and execute S5;
[0032] S3. Extract the second text modality data, second video modality data, and second behavioral modality data from the children's play performance video samples. After modality aggregation processing, obtain performance defect information, action stuttering information, and stereotyped behavior information, calculate the autism prediction value of the children, and then proceed to S5.
[0033] S4. Based on the homework text modal data, group discussion audio modal data, and teacher interaction behavior modal data in children's classrooms, information on the breadth of social relationships, information on dialogue turn-taking delay, and information on compliance with collective instructions are obtained. After embedding processing, information interaction processing is performed to obtain a social activity feature vector. The social activity feature vector is then processed nonlinearly to obtain a predicted value for childhood autism, and S5 is executed.
[0034] S5. Based on the child's autism prediction value and gender, obtain the autism level and match the corresponding treatment plan.
[0035] The specific steps are detailed below.
[0036] S1. Based on the mapping relationship between children's age information and attention periods, obtain the type of children's attention period; if the type of children's attention period is social motivation attention period, execute S2; if the type of children's attention period is play execution attention period, execute S3; if the type of children's attention period is school adaptation attention period, execute S4.
[0037] Children of different ages have different specific symptoms. By obtaining children's age information and mapping the relationship between children's age information and attention periods, we can obtain the types of children's attention periods, which will facilitate the subsequent execution of different specific analyses based on children's age characteristics.
[0038] If a child is between 2 and 4 years old, this period is the sensitive period for pre-language social interaction. Normal children are more receptive to interactive signals. Therefore, this period of childhood is defined as the social motivation attention period. The connection between children and autism can be analyzed by performing the social motivation information analysis operation in S2.
[0039] If a child is between 4 and 6 years old, they are in the preoperational stage, during which symbolic thinking begins to develop rapidly. During this stage, normal children can use play to stimulate their imagination, social understanding, and executive functions. Therefore, this period of childhood is defined as the play-executive attention period. The closeness of the relationship between the child and autism can be analyzed by performing the operations based on play-executive information analysis in S2.
[0040] If the child is between 7 and 8 years old, then the child is entering school age and the social environment shifts from the family or kindergarten to the school. During this stage, normal children can understand abstract social rules (such as taking turns speaking, teacher-student interaction, and simple social relationship expansion). Therefore, this period of childhood is defined as the school adaptation attention period. The closeness of the relationship between the child and autism can be analyzed by performing the operation based on school adaptation information analysis in S4.
[0041] This embodiment uses a mapping relationship between children's age information and attention period to screen for autism. This method can accurately match the laws of children's development, cover specific symptoms at different stages, and reduce interference from differences in developmental levels and subjective bias by matching targeted assessment indicators and tools. This reduces the risk of misjudgment and omission, and can also improve credibility by integrating multi-dimensional evidence. It can not only detect abnormalities in the golden window of early identification and intervention, but also provide guidance for stratified intervention, adapt to the heterogeneous characteristics of the autism spectrum, and achieve an upgrade from "extensive screening" to "precise positioning", which can significantly improve the accuracy of screening.
[0042] S2. Extract the first audio modality data, the first video modality data, and the first behavioral modality data from the children's family interaction video samples. After preprocessing, perform fusion analysis to obtain the children's name-calling response information, joint attention initiation information, and social approach-avoidance information. Calculate the children's autism prediction value and execute S5.
[0043] By extracting the first audio modality data, first video modality data, and first behavioral modality data from children's family interaction video samples, and performing preprocessing such as noise reduction and feature standardization, multimodal fusion analysis is conducted to accurately capture children's name-calling response information, joint attention initiation information, and social approach-avoidance information. This allows for the assessment of children's social abilities from multiple dimensions, including auditory, visual, and behavioral interaction, facilitating early screening and accurate analysis of autism in children aged 2-4.
[0044] First, when the childhood period falls within the social motivation focus period, we obtain family interaction video samples that more easily demonstrate children's social initiation and response patterns as analysis data, and extract the first audio modality data, the first video modality data, and the first behavioral modality data.
[0045] Specifically, the family interaction video samples are processed using audio-video separation techniques (including but not limited to recurrent neural network separation or generative adversarial network separation) to obtain a first audio modality sequence and a first video modality sequence. Simultaneously, children's behaviors in the family interaction video samples are labeled (action tag classification) to obtain a first behavioral modality sequence. The first audio modality sequence is then processed for noise reduction and normalization to obtain first audio modality data. The first video modality sequence undergoes image enhancement and frame sampling to improve video clarity and remove redundant information, resulting in first video modality data. The first behavioral modality sequence undergoes format unification and outlier removal processing to maintain a consistent data format, reduce noise interference with behavioral pattern recognition, and make the extracted behavioral features closer to real social behavior patterns, resulting in first behavioral modality data.
[0046] Then, the first audio modal data, the first video modal data, and the first behavioral modal data are fused and analyzed to obtain the child's name-calling response information, joint attention initiation information, and social approach-avoidance information.
[0047] The call response information is the call response rate, which is obtained based on the ratio of the total number of call responses to the total number of calls in the first text modal data, the first video modal data, and the first behavioral modal data.
[0048] The operation to obtain the name-calling response information is as follows: The time frame in the first audio modality data where the name-calling keyword (e.g., "baby" or the child's corresponding name) appears is taken as a keyframe. The subsequent neighboring time ranges of the keyframes in the first audio modality data, first video modality data, and first behavioral modality sequence are respectively taken as the audio keyframe neighborhood range, video keyframe neighborhood range, and behavioral keyframe neighborhood range. If a child's response sound (e.g., "hey," "ah," "um," etc.) appears in the audio keyframe neighborhood of a keyframe, or a child's corresponding name appears in the corresponding video keyframe neighborhood... Head turning or eye tracking (e.g., head turning towards the location where the keyword was emitted, eye movement towards the location where the keyword was emitted; head turning or eye tracking behaviors can be obtained by detecting video modal data using the YOLOv8 model), or the appearance of a child's corresponding body movement label within the neighborhood of the corresponding behavior keyframe (e.g., the child's position moves towards the location where the keyword was emitted), is recorded as a name call response; the number of name call responses corresponding to all keyframes is counted to obtain the total number of name call responses; the ratio of the total number of name call responses to the total number of name calls (total number of keyframes) is used as the name call response rate to obtain the name call response information.
[0049] The joint attention initiation information is the number of joint attention initiations, which is obtained based on the number of children's proactive social behaviors outside the response scene (a time set formed by several audio keyframe neighborhoods, video keyframe neighborhoods, and behavioral keyframe neighborhoods) in the first audio modality data, and / or the first video modality data, and / or the first behavioral modality data. It is the sum of the number of children's proactive social behaviors. Children's proactive social behaviors include: proactively calling out keywords in the first audio modality data (such as "Mommy," "Look," "This," etc.), proactive eye contact in the first video modality data (such as the child's gaze being directly facing the parent's gaze, and the duration being greater than a first duration threshold), and proactive trajectory approach in the first behavioral modality data (such as the child's position moving closer to the parent in the child's trajectory, and the closest distance being less than a first distance threshold).
[0050] Social approach-avoidance information is the number of social approach-avoidance occurrences, derived from the number of times the child stopped (action-labeled) and turned (action-labeled) upon approaching in the first behavioral modality data. It is the sum of these two counts. Stopping upon approaching and turning upon approaching are the child's stopping and turning behaviors in the behavioral modality data when the child is in a circular position within the parent's neighborhood (a circular position centered on the parent, formed by the first and second distances).
[0051] Finally, based on the child's name-calling response information, joint attention initiation information, and social approach-avoidance information, the predictive value for childhood autism was calculated.
[0052] The above predictive values for childhood autism are calculated using the following formula:
[0053] Q1 = 100·(1-(w) c ·c+w a ·f a -w s ·f s )),
[0054] Q1 is the predictor of childhood autism, C is the name-calling response rate, and f a To generate a contribution value for shared attention, f s Contribution value to social avoidance, w c For the call-to-name response weight, w a To initiate weights for shared attention, w s Social avoidance weight.
[0055] Common attention initiates contribution value f a It is obtained through the following formula: when the number of times joint attention is initiated, a ≤ a base At that time, a reward contribution value is given to those who pay normal, shared attention. When the number of times joint attention is initiated, a > a base At that time, additional shared attention is given to initiate a reward contribution value. 'a' represents the number of times joint attention is initiated. ref The reference value for the number of times joint attention is initiated (belonging to the average level of normal children), a base The baseline value for the number of times joint attention is initiated (belonging to the lowest level of normal children), w 1,a w 2,a These are the first initiation weight and the second initiation weight, respectively.
[0056] Social approach / avoidance contribution value f s It is obtained through the following formula: when the number of social approach / avoidance attempts s≤s ref At that time, no penalty of points will be imposed, f s =0; when the number of social approach / avoidance attempts s>s ref At that time, a penalty of points will be imposed. s b This is the baseline value for the number of times a child exhibits social avoidance (which falls within the average level for normal children).
[0057] S3. Extract the second text modality data, second video modality data, and second behavioral modality data from the children's game performance video samples. After modality aggregation processing, obtain performance defect information, action stuttering information, and stereotyped behavior information, calculate the predicted value of autism in children, and then execute S5.
[0058] Text, video, and behavioral multimodal data are extracted from video samples of children playing games to accurately capture abnormal behaviors related to autism in children. After aggregation processing, information on executive deficits, motion pauses, and stereotyped behaviors is obtained. By integrating multidimensional data, the accuracy and depth of capturing behavioral characteristics of the autism identification system for children aged 4-6 are improved, and more accurate autism prediction values are obtained, providing a more comprehensive and accurate basis for early and accurate identification.
[0059] First, when the childhood period falls within the play performance attention period, second text modal data, second video modal data, and second behavioral modal data are extracted from the children's play performance video samples. Multimodal data can record children's behavior from multiple dimensions such as language expression, body movements, and game operations, accurately capturing abnormal manifestations related to autism in children.
[0060] Specifically, the game execution video samples undergo speech recognition processing to extract the child's speech during the game, converting it into text to obtain the second text modality data. The game execution video samples then undergo keyframe extraction and region of interest enhancement processing to focus on the child's movement areas, improving video quality and facilitating behavior analysis, resulting in the second video modality data. Finally, the game execution video samples undergo pose estimation and motion capture processing (which can be achieved using tools such as MediaPipe) to extract the coordinate sequences of the child's key body points (time-series data), forming action sequence information, thus obtaining the second behavioral modality data.
[0061] Then, the second text modal data, the second video modal data, and the second behavioral modal data are processed by modal aggregation to obtain execution defect information, motion stuttering information, and stereotyped behavior information.
[0062] The execution defect information refers to the incompleteness of actions during game execution. The process for obtaining this information involves: multimodal feature alignment and fusion processing of the second text modality data, second video modality data, and second behavioral modality data to obtain a fused feature vector; action segmentation and recognition processing based on the second behavioral modality data and the fused feature vector (achieved through training a segmentation model, where action labels can be customized according to actual needs); utilizing the fused feature vector improves the accuracy of action segmentation and recognition, outputting action labels for each time segment to obtain action sequence labels; and then performing action label matching between these action sequence labels and corresponding standard action sequence labels. The ratio of the total number of action sequence labels that do not match the standard action sequence labels to the total number of labels in the corresponding standard action sequence labels is used as the incompleteness, thus obtaining the execution defect information. A higher incompleteness indicates a poorer willingness and ability to understand and execute instructions, a greater likelihood of executive function deficits or attention problems, and a stronger tendency towards autism.
[0063] Action stuttering information refers to the number of action stutters. The process for obtaining this information is as follows: the second-row modal data undergoes action temporal analysis to obtain action stuttering information. Specifically, the action temporal analysis process involves using action segmentation methods to divide the continuous actions in the second-row modal data into discrete action segments, obtaining the duration of each action segment. If the duration of the current action exceeds a duration threshold, and the current action differs in type (label) from the two adjacent actions (i.e., the current action is at an action transition point), then the current action is considered a stuttering action. All stuttering actions are counted to obtain the number of action stutters, which serves as the action stuttering information. A higher number of action stutters indicates greater difficulty in connecting action sequences and a stronger tendency towards autism. The action boundaries in the above action segmentation method can be defined and obtained through action classification or speed change detection methods.
[0064] The above-mentioned operation of obtaining motion stuttering information (number of motion stutters) can also be achieved by processing the fused feature vector obtained by multimodal feature alignment and fusion processing of the second text modal data, the second video modal data, and the second line modal data with the second line modal data through a trained stuttering detection model, which can further improve the accuracy of information acquisition.
[0065] Stereotyped behavior information is the number of stereotyped behaviors. The process for obtaining this information (number of stereotyped behaviors) is as follows: Second-behavioral modality data undergoes stereotyped action detection processing to obtain stereotyped behavior information. Specifically, the stereotyped action detection processing involves segmenting continuous behaviors in the second-behavioral modality data into discrete action fragments, classifying each action fragment using an action recognition model to obtain action labels, acquiring consecutive action fragments with the same action label to form repetitive action fragments, and identifying repetitive action fragments whose action labels are not in the standard action label set (indicating stereotyped behavior that existed before the child was freed from game instruction control), or whose total duration exceeds the repetition duration threshold (indicating insufficient behavioral flexibility in the child). In these cases, the repetitive action fragment is considered a stereotyped behavior. All stereotyped behaviors are counted to obtain the number of stereotyped behaviors, which is then used as the stereotyped behavior information.
[0066] The above-mentioned operation of obtaining action stuttering information (number of stereotyped behaviors) can also be achieved by processing the fused feature vector obtained by multimodal feature alignment and fusion processing of the second text modal data, the second video modal data, and the second action modal data with the second action modal data through a trained stereotyped action detection model, which can further improve the accuracy of information acquisition.
[0067] The above predictive values for childhood autism are calculated using the following formula:
[0068]
[0069] Q2 is the predictor of childhood autism, q is the number of executive deficits, and q max The maximum number of execution defects is represented by k, where k is the number of execution pauses, and μ is the maximum number of execution defects. k σ k These represent the mean and standard deviation of the number of stutters in the normal child group, respectively, where x represents the number of stereotyped behaviors. min x max These represent the minimum and maximum values of the stereotyped behavior, w. q w k w x The weights for executive deficits, motion pauses, and stereotyped behaviors are respectively used. The formula for calculating the predictive value of childhood autism in this embodiment integrates three core indicators: executive deficits, motion pauses, and stereotyped behaviors. By weighting and aggregating the number of executive deficits, motion pauses, and stereotyped behaviors, complex behavioral manifestations are transformed into quantifiable predictive values, providing more comprehensive coverage, reducing missed diagnoses and misdiagnoses, and improving the reliability of the analysis of potential relationships in autism.
[0070] S4. Based on the homework text modal data, group discussion audio modal data, and teacher interaction behavior modal data in children's classrooms, obtain information on the breadth of social relationships, the delay in dialogue turn-taking, and the compliance with collective instructions. After embedding and information interaction processing, obtain the social activity feature vector. After nonlinear processing, obtain the predicted value of autism in children, and execute S5.
[0071] Based on classroom multimodal data (homework text, group discussion audio, teacher interaction behavior), multidimensional information such as social relationship breadth, dialogue turn-taking delay, and collective instruction compliance is extracted. Social activity feature vectors are constructed through embedding and interactive processing, and autism prediction values are obtained through nonlinear transformation. This approach comprehensively captures the social behavior characteristics of 7-8 year old children from multiple dimensions, integrating text expression, verbal interaction, instruction response, and other aspects. It also uncovers the potential complex relationships between features, avoiding the one-sidedness of a single modality or indicator, thereby significantly improving the accuracy of autism prediction analysis.
[0072] First, based on children's classroom homework text modal data, group discussion audio modal data, and teacher interaction behavior modal data, information on the breadth of social relationships, dialogue turn-taking delay, collective instruction compliance, and peer interaction initiation is obtained.
[0073] Social relationship breadth information, also known as social relationship span, is derived from the number of different names (e.g., real or fictional names appearing in the essays) and the entropy value of address types in children's classroom homework text modal data (e.g., children's essay data). It reflects the breadth of children's social relationship cognition. The more singular the social relationship breadth information, the narrower the child's social network, and the greater the tendency towards autism. The number of different names can be obtained by performing name entity recognition and peer filtering on the homework text modal data (which can be achieved by using a NER model to identify all names) to obtain a list of peer names and then counting the results.
[0074] The breadth of social relationships is calculated using the following formula:
[0075]
[0076] G represents the breadth of social relationships, λ is the scaling factor, λ = 0.25, N name For different numbers of names, λ E C is the entropy adjustment factor for the address type, where c is the entropy value of the address type. min C max Let p be the minimum and maximum entropy values of the name type. i Let C be the frequency of occurrence of the i-th type of address, and I be the total number of address types (such as calling someone by their first name, nickname, etc.). When all types of addresses are evenly distributed, C max =log2I, The name diversity coefficient uses an exponential decay model to accurately reflect the marginal effect, avoiding the scale distortion caused by simple counting. This formula for calculating the breadth of social relationships uses name diversity... With the entropy adjustment factor λ of the naming type E The integration of these technologies quantifies children's social characteristics using mathematical models. It uses the number of names to reflect the breadth of social contacts and the entropy value of names and moderating factors to characterize the complexity of interaction patterns, achieving a precise mapping from text data to the breadth of social relationships.
[0077] Dialogue turn-off delay information (dialogue turn-off delay) is obtained from children's group discussion audio modal data, specifically dialogues where the children's response time delay is greater than a time delay threshold. It represents the proportion of dialogues where the children's response time delay is greater than the time delay threshold. It can directly quantify the fluency of interaction in children's social dialogues. The greater the dialogue turn-off delay, the worse the interaction fluency and the greater the tendency towards autism.
[0078] The method for obtaining dialogue turn-taking delay information (dialogue turn-taking delay) is as follows: using deep learning methods, based on the voiceprint features in the audio, the speech of different children is distinguished, and the speaker label (e.g., child A, child B) and the corresponding time interval (start time t) of each speech segment are output. start End time t endThe speech segments are sorted out in chronological order. When the speakers of adjacent speech segments are different, it is determined to be a dialogue turn. The time difference before and after the turn is extracted as the response delay. All dialogue turns are traversed, and the number of dialogue turns that meet the condition that the child's response delay is greater than the delay threshold is counted. The ratio of the number of dialogue turns to the total number of dialogue turns is used to obtain the dialogue turn delay degree as the dialogue turn delay degree information.
[0079] Collective instruction compliance information is obtained from teacher interaction behavior modal data (video modal) based on the number of times children perform actions (which can be raising hands or verbal responses) within the immediate time frame after the teacher issues an instruction. It is the ratio of the number of times children perform actions within the immediate time frame after the teacher issues an instruction to the total number of times the teacher issues instructions.
[0080] Then, the information on the breadth of social relationships, the delay in dialogue turn-taking, and compliance with collective instructions are each processed through embedding (mapping from scalar to vector space) to form vectors, resulting in social activity embedding vectors. These social activity embedding vectors are then processed through information interaction to obtain social activity feature vectors. During the embedding process, the information on the breadth of social relationships is inversely proportional to the embedding value (the greater the breadth of social relationships, the smaller the embedding value); the information on the delay in dialogue turn-taking is directly proportional to the embedding value (the greater the breadth of social relationships, the larger the embedding value); and the information on compliance with collective instructions is inversely proportional to the embedding value (the greater the breadth of social relationships, the smaller the embedding value).
[0081] The information interaction processing operation is as follows: the social activity embedding vector is processed by a multilayer perceptron and an attention mechanism to uncover the implicit correlations between features, realize dynamic information interaction, and obtain the social activity feature vector. The dimensions of the multilayer perceptron can be defined according to actual needs.
[0082] Finally, the social activity feature vector is processed nonlinearly (using the ReLU activation function and the sigmoid function) to obtain the predicted value of childhood autism.
[0083]
[0084] H = ReLU(W²·v) s +b2),
[0085] Q3 is the predicted value for childhood autism, σ() is the sigmoid function, H is the hidden layer output, W1 and W2 are the first and second weights respectively, b1 and b2 are the first and second biases respectively, ReLU() is the ReLU activation function, and v s This represents the feature vector of social activities.
[0086] S5. Based on the child's autism prediction value and gender, obtain the autism level and match the corresponding treatment plan.
[0087] The predicted autism level for a child is multiplied by a gender factor to obtain an updated predicted autism level. If the updated predicted autism level is not less than a first predicted threshold, the autism level is Level 1; if the updated predicted autism level is between the first and second predicted thresholds, the autism level is Level 2; and if the updated predicted autism level is not greater than the first predicted threshold, the autism level is Level 3. Boys generally have a higher risk of autism than girls; therefore, in this embodiment, the gender factor for boys is set to be greater than that for girls. By fusing the predicted autism level with gender to obtain the autism level, and by introducing gender-specific risk weights and differentiated thresholds, the risk of misdiagnosis or missed diagnosis due to group differences can be effectively reduced, significantly improving prediction accuracy.
[0088] This embodiment also provides a multimodal intelligent screening system for childhood autism, used to implement the above-mentioned multimodal intelligent screening method for childhood autism, including:
[0089] The Child Attention Period Type Generation Module is used to obtain the child's attention period type based on the mapping relationship between the child's age information and attention periods. If the child's attention period type belongs to the social interaction attention period, the social interaction attention period prediction module is executed; if the child's attention period type belongs to the game execution attention period, the game execution attention period prediction module is executed; if the child's attention period type belongs to the school adaptation attention period, the school adaptation attention period prediction module is executed.
[0090] The social interaction attention period prediction module is used to extract first audio modality data, first video modality data, and first behavioral modality data from children's family interaction video samples, perform fusion analysis, obtain children's name-calling response information, joint attention initiation information, and social approach and avoidance information, calculate the prediction value of autism in children, and execute the autism level and treatment plan generation module;
[0091] The game execution attention period prediction module is used to extract second text modal data, second video modal data, and second behavioral modal data from children's game execution video samples. After modal aggregation processing, it obtains execution defect information, action stuttering information, and stereotyped behavior information, calculates the prediction value of childhood autism, and generates the autism level and treatment plan.
[0092] The school adaptation attention period prediction module is used to obtain information on the breadth of social relationships, the delay in dialogue turn-taking, and compliance with collective instructions based on homework text modal data, group discussion audio modal data, and teacher interaction behavior modal data in children's classrooms. After embedding and processing, information interaction processing is performed to obtain a social activity feature vector. The social activity feature vector is then processed nonlinearly to obtain a predicted value for childhood autism, and an autism level and treatment plan generation module is executed.
[0093] The loneliness level and treatment plan generation module is used to obtain the loneliness level based on the child's autism prediction value and gender, and match the corresponding treatment plan.
[0094] This embodiment also provides a multimodal intelligent screening device for childhood autism, including a processor and a memory, wherein the processor executes a computer program stored in the memory to implement the above-described multimodal intelligent screening method for childhood autism.
[0095] This embodiment also provides a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the above-described multimodal intelligent screening method for childhood autism.
[0096] This embodiment provides a multimodal intelligent screening method for childhood autism. By establishing a mapping relationship between a child's age and three attention periods—social interaction, play execution, and school adaptation—it achieves phased scenario data acquisition and accurate autism assessment. During the social interaction attention period, it integrates audio, video, and behavioral modal data from family interaction recordings to analyze social characteristics such as name-calling response and joint attention initiation. During the play execution attention period, it extracts text, video, and behavioral modal data from play recordings to capture abnormalities such as execution deficits and stereotyped behaviors. During the school adaptation attention period, it integrates homework text, group discussion audio, and teacher interaction data from classroom scenarios to assess the breadth of social relationships. This method identifies key characteristics of autism, such as compliance with instructions and other group activities. Data from each stage is processed through feature extraction, modal aggregation, and nonlinear processing to generate autism prediction values, which are then combined with gender parameters to determine the degree of autism. By dynamically adapting to the core characteristics of each child's developmental stage, this method avoids the limitations of single-standard screening. It enhances the objectivity of assessment through multi-source modal data fusion and shortens the screening cycle with automated algorithms, achieving intelligent processing throughout the entire process from data collection to intelligent screening. This effectively addresses the problems of existing methods, such as strong subjectivity, significant lag, and poor cross-stage adaptability, providing a stable, reliable, accurate, efficient, and comprehensive early autism screening method covering all age groups for clinical autism screening.
Claims
1. A multimodal intelligent screening method for childhood autism, characterized in that, This includes the following operations: S1. Based on the mapping relationship between children's age information and attention periods, the types of children's attention periods are obtained; If the childhood period falls under the social interaction and attention period, execute S2; If the childhood period falls under the play execution attention period, execute S3; If the childhood period falls under the school adjustment and attention period, proceed with S4; S2. Extract the first audio modality data, first video modality data, and first behavioral modality data from the children's family interaction video samples, perform fusion analysis to obtain the children's name-calling response information, joint attention initiation information, and social approach-avoidance information, calculate the children's autism prediction value, and execute S5; The predictive value for childhood autism is calculated using the following formula: , For predicting childhood autism, For the name-calling response rate, To generate contribution values for shared attention, Contribution value to social avoidance Weighting for the call-to-name response. To initiate weights based on common attention, Social avoidance weight; Joint attention initiates contribution value It is obtained through the following formula: when the number of times joint attention is initiated hour, ; When the number of times joint attention is initiated hour, ; To ensure that the number of times the initiative is initiated is shared, For reference values of the number of times common attention is initiated, To establish a baseline for the number of times common attention is initiated, , These are the first initiation weight and the second initiation weight, respectively. Social avoidance contribution value It is obtained through the following formula: when the number of social avoidance events... hour, ; When social avoidance frequency hour, ; This serves as a baseline for the number of social avoidance / avoidance occurrences. S3. Extract the second text modality data, second video modality data, and second behavioral modality data from the children's play performance video samples. After modality aggregation processing, obtain performance defect information, action stuttering information, and stereotyped behavior information, calculate the autism prediction value of the children, and then proceed to S5. S4. Based on the homework text modal data, group discussion audio modal data, and teacher interaction behavior modal data in children's classrooms, information on the breadth of social relationships, information on dialogue turn-taking delay, and information on compliance with collective instructions are obtained. After embedding processing, information interaction processing is performed to obtain a social activity feature vector. The social activity feature vector is then processed nonlinearly to obtain a predicted value for childhood autism, and S5 is executed. S5. Based on the child's autism prediction value and gender, obtain the autism level and match the corresponding treatment plan.
2. The multimodal intelligent screening method for childhood autism according to claim 1, characterized in that, In S2, The call response information is the call response rate, which is obtained based on the ratio of the total number of call responses to the total number of names called in text modal data, video modal data, and behavioral modal data. The joint attention initiation information is the number of joint attention initiations, which is obtained based on the number of children's proactive social behaviors outside the response scenario in the first audio modality data, and / or the first video modality data, and / or the first behavioral modality data; The social approach-avoidance information is the number of social approach-avoidance events, which is obtained based on the number of times the approach was stopped and the number of times the approach was turned in the first behavioral modality data.
3. The multimodal intelligent screening method for childhood autism according to claim 1, characterized in that, In S3, execution defect information refers to the degree of incompleteness of actions during game execution; action stuttering information refers to the number of times actions are stuttered; and stereotyped behavior information refers to the number of times stereotyped behaviors are performed.
4. The multimodal intelligent screening method for childhood autism according to claim 1, characterized in that, In S4, The social relationship breadth information is obtained from the entropy values of different name counts and address types in the modal data of children's classroom homework texts; The dialogue turn-off delay information is obtained from dialogues in children's group discussion audio modal data where the children's response time delay is greater than a time delay threshold; Collective instruction compliance information is obtained from the number of times children perform actions within a short period of time after the teacher issues an instruction, based on teacher interaction behavior modal data.
5. The multimodal intelligent screening method for childhood autism according to claim 1, characterized in that, In S3, the operation to obtain motion lag information is as follows: The second line of modal data is segmented into discrete action segments, and the duration of each action segment is obtained. If the duration of the current action exceeds the duration threshold, and the current action is of a different type than the two adjacent actions, then the current action is a stuttering action. All stuttering actions are counted to obtain the number of stuttering actions, which is used as the stuttering information.
6. The multimodal intelligent screening method for childhood autism according to claim 1, characterized in that, In S4, the operation to obtain the feature vector of social activities is as follows: Information on the breadth of social relationships, the delay in dialogue turn-taking, and compliance with collective instructions are embedded into vectors to obtain social activity embedding vectors. These social activity embedding vectors are then processed by a multilayer perceptron and an attention mechanism to obtain social activity feature vectors.
7. A multimodal intelligent screening system for childhood autism, used to implement the multimodal intelligent screening method for childhood autism according to claim 1, characterized in that, include: The child attention period type generation module is used to obtain the child attention period type based on the mapping relationship between the child's age information and the attention period. If the childhood period falls under the social interaction attention period, execute the social interaction attention period prediction module; If the childhood period falls under the game execution attention period, execute the game execution attention period prediction module; If the childhood period falls under the school adaptation and attention period prediction module, execute the school adaptation and attention period prediction module. The social interaction attention period prediction module is used to extract first audio modality data, first video modality data, and first behavioral modality data from children's family interaction video samples, perform fusion analysis, obtain children's name-calling response information, joint attention initiation information, and social approach and avoidance information, calculate the prediction value of autism in children, and execute the autism level and treatment plan generation module; The game execution attention period prediction module is used to extract second text modal data, second video modal data, and second behavioral modal data from children's game execution video samples. After modal aggregation processing, it obtains execution defect information, action stuttering information, and stereotyped behavior information, calculates the prediction value of childhood autism, and generates the autism level and treatment plan. The school adaptation attention period prediction module is used to obtain information on the breadth of social relationships, the delay in dialogue turn-taking, and compliance with collective instructions based on homework text modal data, group discussion audio modal data, and teacher interaction behavior modal data in children's classrooms. After embedding and processing, information interaction processing is performed to obtain a social activity feature vector. The social activity feature vector is then processed nonlinearly to obtain a predicted value for childhood autism, and an autism level and treatment plan generation module is executed. The loneliness level and treatment plan generation module is used to obtain the loneliness level based on the child's autism prediction value and gender, and match the corresponding treatment plan.
8. A multimodal intelligent screening device for childhood autism, characterized in that, It includes a processor and a memory, wherein the processor executes a computer program stored in the memory to implement the multimodal intelligent screening method for childhood autism as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, Used to store a computer program, wherein the computer program, when executed by a processor, implements the multimodal intelligent screening method for childhood autism as described in any one of claims 1-6.
Citation Information
Patent Citations
Human-machine interaction multi-mode early intervention system for improving social interaction capacity of autistic children
CN102354349A
APP-based autism high-risk infant screening system
CN109545293A
Data acquisition and management system for autism spectrum disorder child information
CN119181494A