Early diagnosis and treatment system for children with autism based on facial representation multi-modal fusion
By adopting a multimodal fusion strategy based on facial representation and a classification network of multi-valve memory mechanism in the autism diagnosis and treatment system, combining expression recognition and line of sight estimation calculation methods, the limitations of the existing system in data acquisition and multimodal fusion are solved, efficient diagnostic and personalized intervention are achieved, and the applicability and diagnostic accuracy of the system are improved.
Patent Information
- Application Number
- CN202510076510.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-06-24
AI Technical Summary
The existing autism diagnosis and treatment system has limitations in data collection and multimodal fusion, which is difficult to promote in daily life scenarios, and the evaluation system is rough and lacks personalized intervention guidance.
A multimodal fusion strategy based on facial representation is adopted, and a multimodal fusion classification network with face detection algorithm and multi-valve memory mechanism is used to combine expression recognition and line of sight estimation algorithm to learn and classify video features, output diagnostic results, and record user training data through intervention games to generate a personalized intervention plan.
It improves the robustness and diagnostic accuracy of the diagnosis and treatment system, lowers the threshold for use of the system, realizes promotion in daily life scenarios, and provides personalized and dynamically adjusted intervention plans, improving the rehabilitation effect.
Smart Images

Figure CN120189115A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross - field of artificial intelligence and medical health, and particularly relates to an early diagnosis and treatment system for autistic children based on multi - modal fusion of facial representations. Background Art
[0002] Existing autism diagnosis and treatment systems have several limitations and urgently need to be further optimized to improve the diagnosis and treatment efficiency and applicability. First, the current technical means of data collection highly rely on special devices such as electroencephalographs, eye trackers, high - definition cameras, etc. These devices have strict requirements for the collection environment and usually need to be completed in laboratories or professional medical institutions, involving complex preparation work and high operation costs, making it difficult to popularize these systems in daily life scenarios and resulting in low popularity. At the same time, the professionalism of the devices also poses high requirements for the technical level of operators, further increasing the usage threshold of the system.
[0003] Secondly, existing multi - modal fusion methods based on facial videos also have significant limitations in data processing and analysis. These methods usually adopt simple fusion strategies, directly splicing or weighted - fusing the hidden features or classification results of two or more modalities. This shallow - level fusion method cannot fully capture the deep complementary information between different modalities, resulting in the model showing low robustness and diagnostic accuracy when facing complex and diverse actual scenarios. In addition, this simple fusion ignores the potential dynamic relationships and time - series dependencies between different modalities, restricting the model's comprehensive understanding of complex behavior patterns. Therefore, designing more efficient and intelligent multi - modal fusion strategies to deeply explore the correlations between modalities has become a research hotspot in the field of autism diagnosis and treatment.
[0004] Finally, existing intervention training methods are too single, and the evaluation system is also relatively rough. Many systems judge the intervention effect by analyzing the overall performance of users, and this evaluation method is difficult to accurately quantify the improvement of users in a specific ability, resulting in the system lacking sufficient pertinence when providing personalized intervention guidance for patients. Therefore, developing refined and multi - dimensional evaluation models, combining performance data in different fields, and formulating personalized and dynamically adjusted intervention plans for patients are the keys to achieving precise rehabilitation of autistic children. Summary of the Invention
[0005] The present invention proposes an early diagnosis and treatment system for autistic children based on multi - modal fusion of facial representations to solve the problems mentioned in the above background art. The technical solution steps of the present invention are as follows:
[0006] Step 1: Data collection and pre - processing;
[0007] Step 1.1: The user sits directly in front of the computer screen according to the schematic diagram and voice prompts, watches the video materials. The video materials consist of three different types of segments, namely, the comparison between social and non-social pictures, popular science explanations, and animation stimuli. The total duration is 7 minutes and 8 seconds. The system calls the camera to record the user's facial video during this process;
[0008] Step 1.2: Use the face detection algorithm to adaptively extract the first 2000 frame images containing faces in the facial video as the input of the diagnostic model;
[0009] Step 2: Use the multi-modal fusion model based on facial expression and gaze estimation algorithm trained on the existing dataset to perform video feature learning and classification, and output the diagnostic result;
[0010] Step 2.1: Extract the expression information of each frame of face image through the cross-domain expression recognition model based on Contrastive Language-Image Pretraining. This expression recognition model has been pre-adaptively trained on the self-collected autism dataset;
[0011] Step 2.2: Use GazeMobileNet, which has been adaptively trained on the self-collected autism dataset, as the gaze estimation network to extract the gaze directions of the left and right eyes of each frame of face image;
[0012] Step 2.3: Use the multi-modal fusion classification network based on the multi-valve memory mechanism to perform multi-modal fusion and classification. This network fuses the facial expression and gaze estimation data of each frame, extracts multi-modal temporal features through multiple groups of gated units and outputs the classification result. This multi-valve memory network consists of three layers of gated combination units, namely, the preliminary memory module, the key extraction module, and the hidden learning module;
[0013] Step 2.3.1: When the multi-modal temporal data passes through the preliminary memory module, it performs preliminary forgetting and retention of historical data and current data. The preliminary memory module consists of a forgetting gate, an input gate, a cell gate, and an output gate. The calculation formula is as follows:
[0014]
[0015] f t =σ(W if X t +b if +W hf h t-1 +b hf )
[0016] i t =σ(W ii X t+b ii +W hi h t-1 +b hi )
[0017] g t = tanh(W ig X t +b ig +W hg h t-1 +b hg )
[0018] o t = σ(W io X t +b io +W ho h t-1 +b ho )
[0019] where and are the three outputs of the initial module respectively, is the output of the key extraction module, f t is the forget gate, g t is the cell gate, i t is the input gate, o t is the output gate, h t-1 is the hidden state at time t-1, c t-1 is the cell state at time t-1, X t is the input at the current time;
[0020] Step 2.3.2: Use the historical data retained by the forget gate and the current data obtained through the input gate in the preliminary memory module as the historical data and current data of the current module respectively for further forgetting and memorization, ensuring that the network can focus on the information most important for the task at the current time. The key extraction module consists of a forget gate, an input gate, and a cell gate, and its calculation formula is as follows:
[0021]
[0022] Step 2.3.3: It uses the latent variable obtained through the output gate of the preliminary memory module as the current data of this module, and the output of the key extraction module as the historical data of this module. Using the fine control of the input gate and the output gate, it effectively transmits and fuses long-term dependence information, thereby generating a richer and more accurate latent variable representation. The hidden learning module consists of an input gate, a cell gate, and an output gate, and its calculation formula is as follows:
[0023] h t = o t ⊙ tanh(c t )
[0024]
[0025] where h t and c t are the outputs of the hidden learning module, h t represents the hidden state at the current moment, and c t represents the cell state at the current moment;
[0026] Step 2.3.4: Obtain the final classification result by passing the latent variable representation learned by the multi-valve memory network through a linear layer;
[0027] Step 3: The system sequentially launches intervention games, guides the user to complete the intervention training through voice prompts, and records the completion status and scores during the games;
[0028] Step 3.1: Launch the whack-a-mole game to train the user's reaction ability;
[0029] Step 3.1.1: Practice stage: A total of 20 trials are set, the duration of the mole is 2 seconds, and there is sound effect feedback. When hitting the mole, it prompts "You did it right!", and when not hitting the mole or hitting the mushroom sprite, it prompts "You did it wrong!";
[0030] Step 3.1.2: Test stage: A total of 120 trials are set, with 30 trials in each group, and there is a 10-second break between groups. In the test stage, only the sound effect of setting a strike is made for the user's actions, without voice feedback. The ratio of the number of moles to mushroom sprites appearing in this game is 4:1, and the scoring standard is: getting 1 point for hitting the mole, getting no points for not hitting the mole, and deducting 1 point for hitting the mushroom sprite;
[0031] Step 3.1.3: Record the user's reaction time (ms), hitting position (x, y), the positions of the mole and mushroom sprite (x, y), hitting type (mole, mushroom sprite, whether hitting or not), and game score during the game;
[0032] Step 3.2: Launch the shape-color recognition game to train the user's cognitive learning ability;
[0033] Step 3.2.1: Color selection stage: The user clicks on the color block that needs to be cognized in advance according to the voice prompt. This color will be used as the target color for the next step;
[0034] Step 3.2.2: Test stage: Correctly identify the "**colored** shape" given in the question from different options, with a time limit of 7 seconds. A total of 5 levels are set, and several shapes and colors are randomly generated in each level, with the difficulty increasing gradually. Getting 1 point for clicking correctly within the specified time, and getting no points for exceeding the time limit or clicking wrongly;
[0035] Step 3.2.3: Record the reaction time, score for each level, and total score of the user during the game.
[0036] Step 3.3: Start the gesture control game to train the user's endurance and hand-eye coordination ability.
[0037] Step 3.3.1: Practice stage: The user completes two rounds of drag-and-drop tasks according to the voice prompts. The user needs to drag all the graphics on the screen into the corresponding graphic frames. If dragged too fast, the shape will fall. The time limit is 5 minutes.
[0038] Step 3.3.2: Test stage: Randomly generate four graphics and corresponding graphic frames in four areas of the screen, including rectangles, triangles, circles, rhombuses, etc. The graphics will show color changes according to the user's actions to prompt whether the drag is successful. One point is scored for each successful match within the specified time.
[0039] Step 3.3.3: Record the user's completion time and total score.
[0040] Step 3.4: Start the language training game to train the user's language expression ability.
[0041] Step 3.4.1: Imitation stage: A picture of an item, such as a cup, is displayed on the screen. Voice prompt: "Kid, this is a cup. Please repeat after me: 'cup'." Wait for the user to say "cup" (using the speech-to-text model with fuzzy matching. If the user's speech contains the word "cup", it is considered a match). The time limit is 15 seconds. If the correct speech is not recognized within 15 seconds, it starts over. The total duration of this stage is 4 minutes. If it is still not completed within 4 minutes, it directly enters the next stage.
[0042] Step 3.4.2: Reinforcement stage: The picture of the item from the imitation stage is displayed on the screen. Voice question: "What is this?" The user needs to answer within 15 seconds.
[0043] Step 3.4.3: Cognition stage: Several pictures are displayed on the screen, including the picture of the item shown in the previous two stages. Voice question: "Which one here is the {cup}?" The user needs to click on the correct option within 15 seconds.
[0044] Step 3.4.4: Learning stage: Unlearned items are added in this stage. Several pictures are displayed on the screen. These pictures are of the same item with different colors. The item is the one used in the previous two stages, such as "cup". Voice question: "Please help me select the '** color' cup." The user needs to click on the correct option within 15 seconds.
[0045] Step 3.4.5: Record the percentage of completion of the user. This game consists of four stages, and 25% is added for each completed stage.
[0046] Step 4: Analyze the training data for multiple times to generate an evaluation report that changes over time;
[0047] Step 4.1: The user completes the intervention training according to the intervention plan, which is customized according to the user's personalization. The system saves the data records (completion status, scores, etc.) of each training;
[0048] Step 4.2: Analyze the recorded training data, and generate a curve graph that changes over time for each training item. This curve graph intuitively reflects the training effect;
[0049] Step 4.3: Generate an evaluation report by synthesizing the training effects of the four types of intervention trainings, providing a basis for formulating subsequent intervention plans. Brief Description of the Drawings
[0050] Figure 1 It is a flowchart of the data collection, diagnosis and intervention process of the present invention;
[0051] Figure 2 It is a framework diagram of the multi-modal fusion algorithm based on the multi-valve control mechanism of the present invention; Detailed Embodiment
[0052] The following further describes the present invention in detail with reference to the preferred embodiments. More details are elaborated in the following description to facilitate a full understanding of the present invention. However, the present invention can obviously be implemented in many other ways different from this description. Those skilled in the art can make similar generalizations and deductions according to the actual application situation without departing from the connotation of the present invention. Therefore, the protection scope of the present invention should not be limited by the content of this specific embodiment.
[0053] An early diagnosis and treatment system for autistic children based on multi-modal fusion of facial representations, as Figure 1 shown, includes the following steps:
[0054] S1: Use an ordinary camera to collect the facial video of the user;
[0055] S2: Use a face detection algorithm to preprocess the video and extract the images of the detected faces in the first 2000 frames;
[0056] S3: Input the preprocessed data into the diagnosis algorithm based on the fusion of expression and gaze;
[0057] S3.1: Use an expression recognition algorithm to obtain the expression information of the face in each frame;
[0058] S3.2: Use a gaze estimation algorithm to obtain the gaze information of the face in each frame;
[0059] S3.3: Combine the expression information with the gaze information to obtain multi-modal time-series data, and input it into the multi-modal fusion classification algorithm to get the classification result;
[0060] S3.4: Output the detection report;
[0061] S4: The user completes the intervention game according to the intervention plan, and the system records the completion status and score of each game each time;
[0062] S5: Analyze the multiple training data, generate a curve graph, and evaluate the intervention effects of the game on the user's concentration, cognitive learning ability, coordination ability, and language ability;
[0063] S6: Generate a comprehensive evaluation report.
[0064] A multi-modal fusion algorithm framework diagram based on a multi-valve control mechanism, as Figure 2 shown, includes the following:
[0065] S1: The multi-modal time-series data passes through the preliminary memory module to perform preliminary forgetting and retention on historical data and current data. The preliminary memory module consists of a forgetting gate, an input gate, a cell gate, and an output gate. The calculation formulas are as follows:
[0066]
[0067] f t = σ(W if X t + b if + W hf h t-1 + b hf )
[0068] i t = σ(W ii X t + b ii + W hi h t-1 + b hi )
[0069] g t = tanh(W ig X t + b ig + W hg h t-1 + b hg )
[0070] o t = σ(W io X t + b io + W ho h t-1 + bho )
[0071] Among them and are respectively the three outputs of the initial module is the output of the key extraction module, f t is the forget gate, g t is the cell gate, i t is the input gate, o t is the output gate, h t-1 is the hidden state at time t-1, c t-1 is the cell state at time t-1, X t is the input at the current time;
[0072] Step 2.3.2: Use the historical data retained by the forget gate and the current data obtained through the input gate in the preliminary memory module as the historical data and current data of the current module respectively for further forgetting and memorizing, ensuring that the network can focus on the information most important for the task at the current time. The key extraction module consists of a forget gate, an input gate, and a cell gate, and its calculation formula is as follows:
[0073]
[0074] Step 2.3.3: It uses the latent variable obtained by the preliminary memory module through the output gate as the current data of this module, and the output of the key extraction module as the historical data of this module. By using the fine control of the input gate and the output gate, it effectively transmits and fuses long-term dependence information, thereby generating a richer and more accurate representation of the latent variable. The hidden learning module consists of an input gate, a cell gate, and an output gate, and its calculation formula is as follows:
[0075] h t = o t ⊙tanh(c t )
[0076]
[0077] Among them, h t , c t are the outputs of the hidden learning module, h t represents the hidden state at the current time, c t represents the cell state at the current time.
[0078] The above is only the preferred solution of the present invention, and is not a further limitation of the present invention. All equivalent changes made by using the content of the specification and drawings of the present invention are within the protection scope of the present invention.
Claims
1. An early diagnosis and treatment system for autistic children based on multimodal fusion of facial representation, characterized by: Data collection module, diagnosis module based on multimodal fusion classification model, game intervention module; the diagnosis and treatment process is as follows: Step 1: Data collection and preprocessing: the pre-set video data is played on the computer screen, and the camera is called to collect the user's facial video and preprocess the video; Step 2: Use the multimodal fusion model based on facial expression and gaze estimation algorithms trained on the existing dataset to learn and classify video features and output the diagnosis results; Step 3: The system starts the intervention games in sequence, guides the children to complete the intervention training through voice prompts, and records the completion status and scores during the game; Step 4: Analyze multiple training data and generate evaluation reports that change over time.
2. The early diagnosis and treatment system for autistic children based on multimodal fusion of facial representations as described in claim 1, characterized in that: The specific contents of data collection in step 1 are: 1.1: The user sits in front of the computer screen according to the diagram and voice prompts to watch the video material, and the system calls the camera to record the user's facial video during the process; 1.2: The video material consists of three different content clips, namely, social and non-social picture comparison, popular science explanation, and animation stimulation, with a total length of 7 minutes and 8 seconds; 1.3: Use the face detection algorithm to adaptively extract the first 2000 frames of images containing faces in the facial video as the input of the diagnosis model.
3. The early diagnosis and treatment system for autistic children based on multimodal fusion of facial representation according to claim 1 is characterized in that: The specific architecture of the multimodal fusion model based on facial expression and sight estimation algorithm in step 2 is: 2.1: Extract the expression information of each frame of facial image through a cross-domain expression recognition model based on contrastive language-image pretraining. The expression recognition model has been adaptively trained on a self-collected autism dataset in advance; 2.2: Use GazeMobileNet, which has been adaptively trained on a self-collected autism dataset, as the gaze estimation network to extract the gaze direction of the left and right eyes in each frame of the face image; 2.3: Use a multimodal fusion classification network based on a multi-valve memory mechanism to perform multimodal fusion and classification. The network fuses the facial expression and gaze estimation data of each frame, extracts multimodal temporal features through multiple groups of gating units, and outputs classification results.
4. The early diagnosis and treatment system for autistic children based on multimodal fusion of facial representation according to claim 3 is characterized in that: In step 2.3, the multimodal fusion classification network architecture based on the multi-valve memory mechanism is: 2.3.1: The multi-valve memory network consists of three layers of gated combination units, namely the preliminary memory module, the key extraction module and the hidden learning module; 2.3.2: The preliminary memory module consists of a forget gate, an input gate, and an output gate. Its purpose is to perform preliminary forgetting and retention of historical data and current data. The calculation formula is as follows: f t =σ(W if X t +b if +W hf h t-1 +b hf ) i t =σ(W ii X t +b ii +W hi h t-1 +b hi ) g t =tanh(W ig X t +b ig +W hg h t-1 +b hg ) o t =σ(W io X t +b io +W ho h t-1 +b ho ) in and are the three outputs of the initial module, is the output of the key extraction module, f t is the forget gate, g t It is the unit door, i t is the input gate, o t is the output gate, h t-1 is the hidden state at time t-1, c t-1 is the cell state at time t-1, X t is the input at the current moment; 2.3.3: The key extraction module consists of a forget gate and an input gate. It uses the historical data retained by the forget gate in the preliminary memory module and the current data obtained by the input gate as the historical data and current data of the current module for further forgetting and memorization, respectively, to ensure that the network can focus on the most important information for the task at the current moment. The calculation formula is as follows: 2.3.4: The hidden learning module consists of an input gate and an output gate. It uses the latent variables obtained by the preliminary memory module through the output gate as the current data of the module, and the output of the key extraction module as the historical data of the module. It uses the fine control of the input gate and the output gate to effectively transmit and fuse long-term dependency information, thereby generating a richer and more accurate latent variable representation. The calculation formula is as follows: h t =o t ⊙tanh(c t ) where h t 、c t is the output of the hidden learning module, h t represents the hidden state at the current moment, c t Indicates the cell state at the current moment; 2.3.5: The latent variable representation learned by the multi-valve memory network passes through a linear layer to obtain the final classification result.
5. The early diagnosis and treatment system for autistic children based on multimodal fusion of facial representation according to claim 1, characterized in that: The intervention games in step 3 include: 3.1: Play whack-a-mole to train users’ reaction ability; 3.2: Graphic-color recognition, training users’ cognitive learning ability; 3.3: Gesture control, training the user's endurance and hand-eye coordination; 3.4: Language training, training users’ language expression ability.
6. The early diagnosis and treatment system for autistic children based on multimodal fusion of facial representation according to claim 5, characterized in that: The game settings for whack-a-mole in 3.1 are: 3.1.1: Game task: hit the moles quickly and correctly, and successfully avoid the interference items (mushroom elves), where the ratio of moles to mushroom elves is 4:1; 3.1.2: Score setting: Hit the mole to get one point, miss the mole to get no points, hit the mushroom elf to deduct one point: 3.1.3: Practice phase: 20 trials in total, the duration of the hamster is 2 seconds, and there is sound feedback. When you hit the hamster, you will be prompted "You did it right!", and when you miss the hamster or hit the mushroom elf, you will be prompted "You did it wrong!"; 3.1.4: Test phase: A total of 120 trials were set, with 30 trials in each group and a 10-second rest period between groups. During the test phase, only the percussion sound effects were set for the user's actions, and no voice feedback was given. 3.1.5: Data recording: reaction time (ms), hitting position (x, y), the position where the mole and mushroom elf appear (x, y), the type of hit (mole, mushroom elf, whether it hits), and the game score.
7. The early diagnosis and treatment system for children with autism based on multimodal fusion of facial representation according to claim 5, characterized in that: The game settings for playing graphics-color recognition in 3.2 are: 3.2.1: Game task: correctly identify the "**color**figure" given in the question from different options, within 7 seconds; 3.2.2: Score setting: one point is awarded for clicking correctly within the specified time, no points are awarded for exceeding the time limit or clicking incorrectly; 3.2.3: Color selection stage: The user clicks on the color block that needs to be recognized in advance according to the voice prompt, and this color will be used as the next target color; 3.2.4: Test phase: There are 5 levels in total, each level randomly generates a number of shapes and colors, and the difficulty gradually increases; 3.2.5: Data recording: reaction time for each level, score for each level, total score.
8. The early diagnosis and treatment system for autistic children based on multimodal fusion of facial representation according to claim 5, characterized in that: In 3.3, the game settings for gesture control are: 3.3.1: Game task: Drag all the shapes on the screen into the corresponding shape boxes. If you drag too fast, the shapes will fall off. The time limit is 5 minutes; 3.3.2: Score setting: one point is awarded for each match completed within the specified time; 3.3.3: Practice phase: Users complete two rounds of dragging tasks according to voice prompts; 3.3.4: Testing phase: Four graphics and corresponding graphic frames are randomly generated in four areas of the screen, including rectangles, triangles, circles, diamonds, etc. The graphics will change color according to the user's actions to prompt the user whether the drag is successful; 3.3.5: Data recording: completion time, total score.
9. The early diagnosis and treatment system for autistic children based on multimodal fusion of facial representation according to claim 5, characterized in that: The game settings for language training in 3.4 are: 3.4.1: Game task: Users express themselves according to the requirements of the questions, which are arranged from easy to difficult; 3.4.2: Scoring setting: one point will be awarded for each task completed within the specified time; 3.4.3: Imitation stage: The screen displays a picture of an object, such as a cup. Voice prompt, "Children, this is a cup, please read with me: "cup". Wait for the user to say "cup" (using the speech-to-text model, using fuzzy matching, if the user's voice contains the word "cup", the match is successful), the time limit is 15 seconds. If the correct voice is not recognized within 15 seconds, restart, the total duration of this stage is 4 minutes, if it is still not completed within 4 minutes, go directly to the next stage. 3.4.4: Reinforcement stage: The screen displays the picture of the object in the imitation stage, and the voice asks "What is this?" The user needs to give a correct answer within 15 seconds; 3.4.5: Cognitive stage: Several pictures are displayed on the screen, including the pictures of objects shown in the first two stages, and a voice question is asked, "Which one here is {cup}?" The user needs to click the correct option within 15 seconds; 3.4.6: Learning stage: In this stage, unlearned items are added, and the screen displays several pictures of the same object in different colors. The object is the object used in the first two stages, such as "cup". The voice question "Please help me choose the cup of '** color'", and the user needs to click the correct option within 15 seconds; 3.4.7: Data record: The percentage of users completing. The game contains four tasks in total, and each completed task increases by 25%.
10. The early diagnosis and treatment system for autistic children based on multimodal fusion of facial representations as claimed in claim 1, characterized in that: In step 4, the steps for generating the evaluation report are: 4.1: The user completes the intervention training according to the intervention plan. The intervention plan is customized according to the user, and the system saves the data records of each training (completion status, score, etc.); 4.2: Analyze the recorded training data and generate a curve chart over time for each training. The curve chart intuitively reflects the training effect: 4.3: Generate an evaluation report based on the training effects of the four types of intervention training to provide a basis for the subsequent formulation of intervention plans.
Citation Information
Cited By
Autism risk assessment method and system based on specific social animation stimulation
CN121331469A