Autism Children Intervention and Treatment Assistance System and Method Based on Intelligent Interaction

Through the intelligent interaction system combined with AI technology, analyzing children's facial expressions and audio data, the problem that traditional intervention methods are difficult to capture subtle behaviors in real time is solved, and accurate judgment and personalized intervention of the behavioral status of children with autism are achieved.

CN119541839BActive Publication Date: 2025-06-24JILIN UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510096375.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-24
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

Traditional autistic children’s intervention methods rely on human observers and are difficult to capture children’s subtle behaviors and expressions in real time, resulting in inaccurate assessments and insufficient intervention.

Method used

Using an intelligent interactive system based on AI, children's facial expression images and audio data are obtained through cameras and microphones, and features are extracted using the audio-image multimodal joint analysis module to judge children's behavioral status and generate intervention feedback.

Benefits of technology

Accurate judgment of the behavioral status of children with autism is achieved, ensuring that no instantaneous behavior is missed, and more targeted personalized intervention feedback is generated, improving the accuracy and comprehensiveness of the intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119541839B_ABST
    Figure CN119541839B_ABST
Patent Text Reader

Abstract

The present application provides an intelligent interaction-based auxiliary system and method for the intervention and treatment of autistic children, which relates to the field of intelligent data analysis. It uses AI-based data processing technology to extract audio features from the audio data of autistic children, and at the same time extracts facial features from the facial expression images of autistic children. Based on this, it can intelligently judge whether the behavior state of the autistic children is an attempt to communicate based on the multi-modal fine-grained joint representation between the audio feature encoding features and the facial image feature encoding features. In this way, through continuous data collection, it is possible to more accurately capture the extremely subtle changes of autistic children and ensure that no behavior occurring at any moment is missed. Moreover, by combining facial expression analysis and audio feature extraction, the judgment of the behavior state of autistic children can be made more accurate, and then more targeted personalized intervention feedback can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data intelligent analysis, and more specifically, to an intervention treatment assistance system and method for autistic children based on intelligent interaction. Background Art

[0002] Autism, also known as autism spectrum disorder, is a neurodevelopmental disorder that typically appears in early childhood and has a wide range of effects on an individual's social interaction, communication skills, interests, and behavior patterns. Therefore, timely intervention is crucial for supporting the development of children with autism and can significantly promote their improvement in social interaction, language expression, and behavior management.

[0003] Traditional intervention methods for autistic children mainly rely on manual assessments by human observers and the experience guidance of therapists. However, due to manpower limitations, it is difficult for professionals to be present at all times for real-time observation, resulting in the easy neglect of some extremely important behaviors that occur instantaneously. In particular, some subtle expressions or brief vocalizations of autistic children, which often only appear in specific situations but are crucial for judging their behavioral states. In addition, traditional methods highly rely on the intuition and personal experience of observers, and the experience level of therapists directly affects the understanding and response quality of behavioral changes. Even in the case of face-to-face observation, human observation may miss key details, such as slight eye movements, subtle contractions of facial muscles, etc., due to distraction or other reasons, thus affecting the accuracy of assessment and the comprehensiveness of intervention treatment.

[0004] Therefore, an intervention treatment assistance solution for autistic children based on intelligent interaction is desired. Summary of the Invention

[0005] To solve the above technical problems, this application is proposed. This application provides an intervention treatment assistance system and method for autistic children based on intelligent interaction.

[0006] According to one aspect of this application, an intervention treatment assistance system for autistic children based on intelligent interaction is provided, which includes:

[0007] A child object face and audio acquisition module for acquiring facial expression images of an autistic child object collected by a camera and audio data of the autistic child object collected by a microphone;

[0008] An audio feature extraction module for extracting audio features from the audio data to obtain an audio feature encoding vector;

[0009] A facial feature extraction module for extracting facial features from the facial expression images to obtain a facial image feature encoding feature map;

[0010] An audio-image multimodal joint module is used to perform fine-grained joint of audio-image multimodal data on the audio feature encoding vector and the facial image feature encoding feature map to obtain an autism child object state multimodal joint encoding vector. Among them, the audio-image multimodal joint module includes: a facial image decomposition unit, which is used to perform feature decoupling and feature flattening on the facial image feature encoding feature map to obtain a set of facial image local feature encoding vectors; a clustering fusion unit, which is used to perform kernel feature calculation and fine-grained clustering fusion on the audio feature encoding vector and the set of facial image local feature encoding vectors to obtain the autism child object state multimodal joint encoding vector;

[0011] A behavior state judgment module is used to determine the behavior state of the autism child object based on the autism child object state multimodal joint encoding vector, and generate an intervention feedback based on the behavior state.

[0012] Further, the audio feature extraction module includes:

[0013] An audio data wavelet transform unit, which is used to perform discrete wavelet transform on the audio data to obtain an audio data time-frequency diagram;

[0014] An audio feature generation unit, which is used to perform audio feature extraction on the audio data time-frequency diagram using a feature extractor based on a CNN-GRU hybrid model to obtain the audio feature encoding vector.

[0015] Further, the facial feature extraction module is used to: extract facial features from the facial expression image using a facial feature extractor based on a dilated pyramid model to obtain the facial image feature encoding feature map.

[0016] Further, the facial image decomposition unit is used to:

[0017] Perform feature decoupling on the facial image feature encoding feature map along the channel dimension of the facial image feature encoding feature map to obtain a set of facial image feature encoding local feature matrices;

[0018] Flatten the features of each facial image feature encoding local feature matrix in the set of facial image feature encoding local feature matrices to obtain the set of facial image local feature encoding vectors.

[0019] Further, the clustering fusion unit includes:

[0020] A facial image kernel feature extraction subunit, which is used to perform facial image kernel feature extraction on the set of facial image local feature encoding vectors to obtain a facial image kernel feature encoding vector;

[0021] An audio-facial feature prior clustering center encoding vector determination subunit, configured to determine an audio-facial feature prior clustering center encoding vector based on the audio feature encoding vector and the facial image kernel feature encoding vector;

[0022] A multi-modal fine-grained joint subunit, configured to perform multi-modal fine-grained joint encoding on the set of the audio-facial feature prior clustering center encoding vector and the facial image local feature encoding vectors to obtain the multi-modal joint encoding vector of the autistic child object state.

[0023] Further, the multi-modal fine-grained joint subunit is configured to:

[0024] Calculate the Poincaré distance between the audio-facial feature prior clustering center encoding vector and each facial image local feature encoding vector in the set of the facial image local feature encoding vectors respectively to obtain a set of audio-facial feature semantic metric values;

[0025] Use a binary function to perform clustering judgment on the set of the audio-facial feature semantic metric values to obtain a set of audio-facial feature clustering coefficients;

[0026] Based on the set of the audio-facial feature clustering coefficients, perform fine-grained cross-domain fusion on the set of the audio-facial feature prior clustering center encoding vector and the facial image local feature encoding vectors to obtain the multi-modal joint encoding vector of the autistic child object state.

[0027] Further, the behavior state judgment module includes:

[0028] A behavior state result generation unit, configured to input the multi-modal joint encoding vector of the autistic child object state into a behavior state judge to obtain the behavior state of the autistic child object, where the behavior state of the autistic child object is whether to attempt to communicate;

[0029] An intervention feedback generation unit, configured to generate the intervention feedback in response to the behavior state of the autistic child object being an attempt to communicate.

[0030] Further, the behavior state judge is a behavior state judge based on a classifier.

[0031] According to another aspect of the present application, there is provided an intelligent interaction-based autistic child intervention and treatment assistance method, which includes:

[0032] Obtain a facial expression image of an autistic child object collected by a camera and audio data of the autistic child object collected by a microphone;

[0033] Extract audio features from the audio data to obtain an audio feature encoding vector;

[0034] Extract facial features from the facial expression image to obtain a facial image feature encoding feature map;

[0035] Perform audio-image multimodal data fine-grained joint on the audio feature encoding vector and the facial image feature encoding feature map to obtain an autism child object state multimodal joint encoding vector, including: performing feature decoupling and feature flattening on the facial image feature encoding feature map to obtain a set of facial image local feature encoding vectors; performing kernel feature calculation and fine-grained clustering fusion on the audio feature encoding vector and the set of facial image local feature encoding vectors to obtain the autism child object state multimodal joint encoding vector;

[0036] Based on the autism child object state multimodal joint encoding vector, determine the behavior state of the autism child object, and generate an intervention feedback based on the behavior state.

[0037] 10. The method for assisting in the intervention treatment of autism children based on intelligent interaction according to claim 9, wherein extracting audio features from the audio data to obtain an audio feature encoding vector includes:

[0038] Perform discrete wavelet transform on the audio data to obtain an audio data time-frequency diagram;

[0039] Use a feature extractor based on a CNN-GRU hybrid model to extract audio features from the audio data time-frequency diagram to obtain the audio feature encoding vector.

[0040] Compared with the prior art, the system and method for assisting in the intervention treatment of autism children based on intelligent interaction provided by the present application adopt an AI-based data processing technology to extract audio features from the audio data of an autism child object, and at the same time extract facial features from the facial expression image of the autism child object, so as to intelligently judge whether the behavior state of the autism child object is attempting to communicate based on the multimodal fine-grained joint representation between the audio feature encoding features and the facial image feature encoding features. In this way, through continuous data collection, extremely subtle changes in autism children can be captured more accurately, ensuring that no behavior occurring at any moment is missed. And by combining facial expression analysis and audio feature extraction, the judgment of the behavior state of autism children can be made more accurate, and then a more targeted personalized intervention feedback can be generated. Description of the Drawings

[0041] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings. In the accompanying drawings:

[0042] Figure 1 It is a system block diagram of an intelligent interaction-based autism children intervention and treatment assistance system according to an embodiment of the present application.

[0043] Figure 2 It is a block diagram of an audio feature extraction module in an intelligent interaction-based autism children intervention and treatment assistance system according to an embodiment of the present application.

[0044] Figure 3 It is a block diagram of an audio-image multimodal joint module in an intelligent interaction-based autism children intervention and treatment assistance system according to an embodiment of the present application.

[0045] Figure 4 It is a block diagram of a behavior state judgment module in an intelligent interaction-based autism children intervention and treatment assistance system according to an embodiment of the present application.

[0046] Figure 5 It is a flowchart of an intelligent interaction-based autism children intervention and treatment assistance method according to an embodiment of the present application. Detailed implementation manners

[0047] Next, example embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the example embodiments described herein.

[0048] Autism, scientifically known as autism spectrum disorder, belongs to the category of neurodevelopmental disorders. This disorder generally appears in early childhood and has a wide and profound impact on an individual's social interaction, communication ability, interests and hobbies, and behavior patterns. Therefore, timely intervention is crucial for helping the growth and development of autistic children, and can significantly promote their ability improvement in aspects such as social interaction, language expression, and behavior management.

[0049] Traditional intervention means for autistic children mainly rely on manual evaluation by human observers and guidance based on the experience of therapists. However, limited by human factors, professionals cannot always be present for real-time observation, which leads to the easy omission of some behaviors that occur instantaneously but are extremely important. Especially some subtle expressions or short voice expressions of autistic children, which often only appear in specific situations but play a key role in judging their behavior states.

[0050] In addition, traditional methods rely to a large extent on the intuition and personal experience of observers. The experience level of therapists is directly related to the depth of understanding of behavioral changes and the quality of responses. Even in the case of face-to-face observation, human observation may still miss key details such as slight eye movements and subtle contractions of facial muscles due to distraction or other factors, thus affecting the accuracy of assessment and the comprehensiveness of intervention treatment.

[0051] To address the above technical problems, the technical concept of this application is to obtain the facial expression images of an autistic child object collected by a camera and the audio data of the autistic child object collected by a microphone, and use AI-based image analysis technology and data processing algorithms to extract audio features from the audio data, while extracting facial features from the facial expression images, so as to intelligently judge whether the behavioral state of the autistic child object is attempting to communicate based on the multi-modal fine-grained joint representation between the audio feature encoding features and the facial image feature encoding features. This application realizes continuous data collection with the help of camera and microphone devices, can accurately capture extremely subtle changes, breaks through the limitations of human power, and ensures that no instantaneous behavior is missed. Moreover, by combining facial expression analysis and audio feature extraction, the behavioral performance of autistic children is comprehensively interpreted from multiple perspectives, making the judgment of the behavioral state more accurate, and then generating more targeted personalized intervention feedback.

[0052] Figure 1 System block diagram of an autistic child intervention treatment assistance system based on intelligent interaction according to an embodiment of the present application. As Figure 1 shown, in the autistic child intervention treatment assistance system 100 based on intelligent interaction, it includes: a child object face audio acquisition module 110, configured to acquire the facial expression images of an autistic child object collected by a camera and the audio data of the autistic child object collected by a microphone; an audio feature extraction module 120, configured to extract audio features from the audio data to obtain an audio feature encoding vector; a facial feature extraction module 130, configured to extract facial features from the facial expression images to obtain a facial image feature encoding feature map; an audio-image multi-modal joint module 140, configured to perform audio-image multi-modal data fine-grained joint on the audio feature encoding vector and the facial image feature encoding feature map to obtain an autistic child object state multi-modal joint encoding vector; a behavioral state judgment module 150, configured to determine the behavioral state of the autistic child object based on the autistic child object state multi-modal joint encoding vector, and generate an intervention feedback based on the behavioral state.

[0053] In the embodiment of the present application, the child object face audio acquisition module 110 is configured to acquire the facial expression images of the autistic child object collected by the camera and the audio data of the autistic child object collected by the microphone. It should be understood that the facial expression images of the autistic child object specifically include the basic facial features of the object, expression dynamic information, eye-related information, etc.; the audio data of the autistic child object specifically includes the language content information, voice pitch information, voice volume information, voice factor information, and some non-verbal information of the child object when speaking. Specifically, positive facial expressions in the facial images, such as smiling, bright eyes and looking at the communication object, are usually signals of an attempt to communicate. These expressions indicate that the child is interested in the communication object and is willing to initiate or maintain communication and interaction. On the contrary, a blank expression, frowning or avoiding eye contact may imply a lack of willingness to communicate or a communication barrier; the speech content in the audio data clearly expresses the child's needs, thoughts or emotions and is direct evidence for judging an attempt to communicate. For example, initiating a conversation actively, such as "What are you doing", indicates that the child has a strong willingness to communicate. Even if the language expression is incomplete or there are errors, as long as there is an intention to express, it shows an attempt to communicate; and appropriate pitch, volume and speech rate are also key information for judging whether a child is attempting to communicate. Generally speaking, the facial expression images and the audio data complement each other and jointly provide more comprehensive information. Through the multi-modal fusion analysis of the two types of data, it is possible to more accurately judge whether an autistic child is attempting to communicate, avoid misjudgments that may occur based on a single data source, and improve the accuracy and reliability of the judgment.

[0054] The following is a detailed elaboration of a specific implementation process of "acquiring the facial expression images of the autistic child object collected by the camera and the audio data of the autistic child object collected by the microphone":

[0055] First, the preparation of the environment and equipment is crucial. A dedicated intervention and treatment space should be created, which should be quiet, have soft and sufficient lighting, and at the same time create a warm, comfortable and attractive atmosphere in the layout, so as to minimize the interference of external factors on data collection. In terms of the room layout, it should be reasonably planned to ensure that children have enough space to move around while ensuring that the cameras and microphones can be in the best working positions. The installation position of the camera needs to be carefully debugged. Usually, it is installed at a suitable position above or near the screen. The adjustment of its angle and focal length should be precise to ensure that every subtle facial expression of the child during the interaction with the virtual character can be captured completely and clearly, such as the blinking of eyes, the opening and closing of the mouth, the raising or wrinkling of the eyebrows and other dynamic changes of key parts, without omission. The placement of the microphone is also particular. It should be placed close to the child, such as on the desktop or near the child's seat, so as to ensure that all kinds of sounds of the child can be accurately recorded, whether it is clear speech, cheerful laughter, aggrieved crying, or some subtle breathing sounds, muttering sounds, etc., can be accurately captured and stably connected to the data processing system. Before formal use, the equipment also needs to be comprehensively calibrated and strictly tested to ensure the stability and accuracy of data collection and lay a solid foundation for subsequent analysis work.

[0056] Secondly, the design of the virtual character and the setting of the interaction program are one of the core links in the whole implementation process. The created virtual character should be extremely amiable and attractive, with vivid and distinctive images, such as designed as cute cartoon animal images or friendly children's images. These virtual characters are presented on the screen in the form of exquisite animations, with rich and diverse expressions and natural and smooth movements, which can quickly attract the attention of autistic children and stimulate their interest. In terms of the interaction program, a series of creative and targeted interaction scripts need to be carefully compiled, covering various forms such as simple and interesting games, fascinating story-telling, and daily communication topics close to life. For example, design a game of "Animal Paradise Adventure", where the virtual character acts as a tour guide to lead the child to explore in the virtual animal paradise, encounter various animals along the way and introduce their characteristics and habits, and at the same time ask questions to the child to guide them to participate in the interaction. During this process, the voice intonation of the virtual character should be natural and gentle, with a moderate speed, and can be flexibly adjusted according to the child's reaction and ability. If the child shows signs of difficulty in understanding, the voice can be slowed down and the words used can be simpler; if the child responds actively, the voice can be more lively and vivid to enhance the fun of the interaction. At the same time, the interaction program should skillfully set the trigger points for data collection to ensure that the facial expression images and audio data of the child can be accurately obtained during the key interaction links, without missing any valuable information.

[0057] Furthermore, the data collection process is a crucial operational link in the entire implementation process. When a child enters the intervention and treatment area and is ready to interact with the virtual character, the system is officially launched. At this time, the virtual character will first give a short and attractive opening introduction and guidance. Meanwhile, the camera and microphone start continuously collecting data synchronously. The system will record the child's first reaction expression when initially seeing the virtual character, as well as any possible sounds made, which are used as important basic data. Throughout the interaction process, the camera and microphone always remain in a real-time working state, continuously capturing the child's behavior and sound information. For example, in the interactive session of "animal sound imitation show", the virtual character plays the sound of an animal and invites the child to imitate. At this time, the system will focus on collecting the child's facial expression changes and voice responses during listening to the sound, preparing to imitate, and the imitation process. The child may show expressions such as a concentrated look, curious eyes, confused frown, or excited smile, and the audio characteristics such as intonation, speech rate, and volume during imitating the sound will all be recorded in detail. For some non-verbal behaviors of the child, such as gestures like nodding, shaking the head, and waving the hand, the camera will also keenly capture the relevant image information and perform preliminary analysis and marking with the help of advanced image recognition technology, providing comprehensive data support for subsequent in-depth research.

[0058] The collected facial expression images and audio data will be immediately transmitted to a dedicated storage device and stored in a detailed classification according to the time sequence and interactive sessions. At the same time, the system will perform preliminary processing on the data, using professional algorithms and technologies to remove some obvious noise interferences, and performing simple clarity enhancement and optimization processing on the images to ensure the quality of the data for more in-depth and accurate analysis in the future.

[0059] In the embodiment of the present application, the audio feature extraction module 120 is used to extract audio features from the audio data to obtain an audio feature coding vector. Specifically, Figure 2 It is a block diagram of the audio feature extraction module in the autism child intervention and treatment assistance system based on intelligent interaction according to the embodiment of the present application. As Figure 2 shown, the audio feature extraction module 120 includes: an audio data wavelet transform unit 121, which is used to perform discrete wavelet transform on the audio data to obtain an audio data time-frequency diagram; an audio feature generation unit 122, which is used to perform audio feature extraction on the audio data time-frequency diagram using a feature extractor based on a CNN-GRU hybrid model to obtain the audio feature coding vector.

[0060] In the embodiment of the present application, the audio data wavelet transform unit 121 is configured to perform discrete wavelet transform on the audio data to obtain a time-frequency diagram of the audio data. Accordingly, considering that not all information in the audio data is equally important for judging the behavioral state of autistic children, which includes the voice, intonation, speech rate of the child's speech, and some background sounds, etc. Therefore, in order to focus on the key information related to the behavioral state, such as features like the fundamental frequency and formants of speech, the present application extracts audio features from the audio data to extract features closely related to emotion and intention expression, obtaining an audio feature coding vector, which helps to more accurately analyze the behavior of autistic children. In particular, in a specific embodiment of the present application, discrete wavelet transform is performed on the audio data to decompose the audio signal into sub-signals of different frequency bands and time scales, obtaining a time-frequency diagram of the audio data, presenting the change of the frequency of the audio signal over time in an intuitive visual way. It should be understood that an important characteristic of discrete wavelet transform is its ability to provide multi-resolution analysis. Audio signals usually contain multiple frequency components, and the performance of these components at different time scales is crucial for understanding the audio behavior of autistic children.

[0061] In the embodiment of the present application, the audio feature generation unit 122 is configured to use a feature extractor based on a CNN-GRU hybrid model to perform audio feature extraction on the time-frequency diagram of the audio data to obtain the audio feature coding vector. It should be understood that in order to effectively capture and mine the audio feature information in the time-frequency diagram, further, a feature extractor based on a CNN-GRU hybrid model is used to perform audio feature extraction on the time-frequency diagram of the audio data to obtain the audio feature coding vector. That is to say, the time-frequency diagram of the audio data is essentially a two-dimensional image representation, with the abscissa being time and the ordinate being frequency, and the values at different positions represent the signal strength at that time-frequency point. CNN (Convolutional Neural Networks) has natural advantages in processing image data. The convolutional kernels in its convolutional layer can slide on the time-frequency diagram, capturing local feature patterns in the time-frequency diagram through local receptive fields, such as key audio features representing the start of speech, intonation changes, etc. And audio signals have time series characteristics, and there is a correlation between the front and back audio information. GRU (Gated Recurrent Unit), as a variant of the recurrent neural network, can effectively process time series data and capture the dynamic changes of audio signals in the time dimension, such as the change patterns of speech prosody and speech rate in the time series.

[0062] In an embodiment of the present application, the facial feature extraction module 130 is used to extract facial features from the facial expression image to obtain a facial image feature encoding feature map. Specifically, in an embodiment of the present application, the facial feature extraction module 130 is used to: use a facial feature extractor based on a hollow pyramid model to extract facial features from the facial expression image to obtain the facial image feature encoding feature map. Accordingly, considering that the subtle expression changes in the facial expression image contain rich psychological and behavioral information, such as the corners of the mouth, eye contact, etc. Based on this, in order to dig out the characteristic information hidden in the facial expression, so as to help judge the child's communication willingness and emotional state, the present application extracts facial features from the facial expression image to dig out hidden information such as micro-expressions, and obtains a facial image feature encoding feature map. In particular, in a specific embodiment of the present application, a facial feature extractor based on a hollow pyramid model is used to extract facial features from the facial expression image to obtain the facial image feature encoding feature map. It should be understood that the rich information of facial expressions is reflected at different scales. For example, subtle muscle movements around the eyes (such as blinking and pupil changes) are small-scale features that can convey emotions such as surprise and concentration; while muscle stretching of the entire face, upturned or downturned corners of the mouth, etc. are larger-scale features that reflect more obvious emotional states, such as happiness or sadness. The dilated pyramid model can extract features from images at multiple scales through convolution operations with different dilation rates to capture important information at different scales in facial expression images, that is, it can effectively perceive changes from local details to overall facial structure, thereby more comprehensively portraying facial expression features, so as to more accurately reflect the information conveyed by the facial expressions of children with autism. For example, accurately captured facial muscle movement features can more accurately judge the emotional state of children with autism, and thus provide a more reliable basis for judging whether they are trying to communicate.

[0063] In the embodiment of the present application, the audio-image multimodal joint module 140 is used to perform a fine-grained audio-image multimodal data joint on the audio feature coding vector and the facial image feature coding feature map to obtain a multimodal joint coding vector of the autistic child object state. Specifically, Figure 3 FIG. 1 is a block diagram of an audio-image multimodal joint module in an intelligent interaction-based autism children intervention and treatment auxiliary system according to an embodiment of the present application. Figure 3As shown, the audio-image multimodal joint module 140 includes: a facial image decomposition unit 141, used to perform feature decoupling and feature flattening on the facial image feature coding feature map to obtain a set of facial image local feature coding vectors; a clustering fusion unit 142, used to perform kernel feature calculation and fine-grained clustering fusion on the set of the audio feature coding vector and the facial image local feature coding vector to obtain the autistic child object state multimodal joint coding vector.

[0064] It should be understood that the audio feature coding vector reflects the key information in the audio signal of autistic children, such as voice intonation, sound frequency, etc., while the facial image feature coding feature map captures the information conveyed by facial expressions, such as emotional expressions such as joy, anger, sorrow, and happiness. The data of these two modalities describe the behavioral state of autistic children from different angles. Therefore, in order to integrate these multi-source information, give full play to their respective advantages, overcome the limitations of single modality data, and understand their behavioral state more comprehensively, in the technical solution of the present application, the audio feature coding vector and the facial image feature coding feature map are subjected to audio-image multimodal data fine-grained joint to obtain autistic children's object state multimodal joint coding vector. In particular, audio and facial images belong to different information domains, and there is a potential fine-grained relationship between them. For example, a specific voice intonation may be accompanied by a specific facial expression, such as an angry tone may be accompanied by facial expressions such as frowning and staring. This method can dig out these cross-domain fine-grained relationships and deeply analyze the correlation between different modal data to more accurately interpret their behavioral states.

[0065] Specifically, in the embodiment of the present application, the facial image decomposition unit 141 is used to: perform feature decoupling on the facial image feature coding feature map along the channel dimension of the facial image feature coding feature map to obtain a set of facial image feature coding local feature matrices; and perform feature flattening on each facial image feature coding local feature matrix in the set of facial image feature coding local feature matrices to obtain a set of facial image local feature coding vectors. The above process can be expressed as:

[0066] ;

[0067] in, is the facial image feature encoding feature map, It is a feature decoupling and feature flattening operation. is the set of local feature encoding vectors of facial images, , , , and are respectively the first, second, and third in the set of local feature encoding vectors of the facial image individual, the individual and the facial image local feature encoding vectors.

[0068] It should be understood that the facial image feature encoding feature map contains information such as expressions, facial muscle movements, and facial feature shapes. To separate this information and provide more detailed information for subsequent analysis, in this application, a feature decoupling operation needs to be performed on the facial image feature encoding feature map along the channel dimension. Each facial image feature encoding local feature matrix in the set of facial image feature encoding local feature matrices obtained by decoupling only focuses on and carries a single type of semantic information, which can avoid interference between different semantic information, thereby making subsequent processing more targeted and accurate. Correspondingly, considering that there may be some redundant information in the multi-dimensional feature matrix, which may be caused by the spatial structure of the features or the correlation between dimensions. To be able to remove some of the redundancy brought by the spatial structure and enable the model to more effectively capture key information during the learning process without being interfered by redundant spatial structure information, in this application, each facial image feature encoding local feature matrix in the set of facial image feature encoding local feature matrices needs to be flattened to obtain the set of facial image local feature encoding vectors. Moreover, through the flattening operation of the features, it can be ensured that the local features can be uniformly processed while maintaining their respective characteristics.

[0069] Specifically, in the embodiment of this application, the clustering fusion unit 142 includes: a facial image kernel feature extraction subunit, configured to perform facial image kernel feature extraction on the set of facial image local feature encoding vectors to obtain facial image kernel feature encoding vectors; an audio-facial feature prior clustering center encoding vector determination subunit, configured to determine an audio-facial feature prior clustering center encoding vector based on the audio feature encoding vector and the facial image kernel feature encoding vector; and a multi-modal fine-grained joint subunit, configured to perform multi-modal fine-grained joint encoding on the audio-facial feature prior clustering center encoding vector and the set of facial image local feature encoding vectors to obtain the multi-modal joint encoding vector of the autistic child object state.

[0070] Next, facial image kernel feature extraction is performed on the set of facial image local feature encoding vectors to obtain facial image kernel feature encoding vectors. The above process can be expressed as:

[0071] ;

[0072] wherein, is the set of facial image local feature encoding vectors, , are respectively the -th and the -th facial image local feature encoding vectors in the set of the facial image local feature encoding vectors, is to perform kernel feature extraction on represents the first norm of the feature vector, represents the number of vectors in the set of the facial image local feature encoding vectors minus one, is the corresponding facial image semantic difference value, represents the exponential function with the natural constant as the base, is the number of vectors in the set of the facial image local feature encoding vectors, is the facial image kernel feature encoding vector.

[0073] It should be understood that when dealing with the problem of judging the behavioral state of autistic children, the audio feature encoding vector and the set of facial image local feature encoding vectors represent different modalities of data. In order to effectively fuse these two modalities of data, it is necessary to find the correlation points between them, that is, the cross-modal data interaction anchoring bridge. By performing kernel analysis on the set of facial image local feature encoding vectors, the most core and representative features in the facial image data can be mined, and these features can be used as the key bridge for interacting with the audio modality data to help establish an effective connection between the two modalities of data.

[0074] Then, based on the audio feature encoding vector and the facial image kernel feature encoding vector, an audio-facial feature prior clustering center encoding vector is determined. The above process can be expressed as:

[0075] ;

[0076] where, is the facial image kernel feature encoding vector, is the audio feature encoding vector, is the concatenation operation, is the weight transformation matrix, is the bias vector, is the audio-facial feature prior clustering center encoding vector.

[0077] It should be understood that the audio feature encoding vector and the facial image kernel feature encoding vector belong to data of different modalities, each containing unique information about the behavioral state of autistic children. To achieve effective fusion and interaction of the two-modal data, a clear correlation point is needed. The audio-facial feature prior clustering center encoding vector serves such a role. As an anchor bridge for cross-modal data interaction, it connects the data of the two modalities, namely audio and facial images, enabling orderly cross-modal fine-grained interaction. Moreover, by determining the audio-facial feature prior clustering center encoding vector, the advantages of the two modalities can be combined to more comprehensively understand the children's behavioral performance during communication.

[0078] More specifically, in the embodiment of the present application, the multi-modal fine-grained joint subunit is used to: respectively calculate the Poincaré distance between the audio-facial feature prior clustering center encoding vector and each facial image local feature encoding vector in the set of facial image local feature encoding vectors to obtain a set of audio-facial feature semantic metric values. This process can be expressed as:

[0079] ;

[0080] Wherein, is the th facial image local feature encoding vector in the set of facial image local feature encoding vectors, is the audio-facial feature prior clustering center encoding vector, is the square of the Euclidean norm of the vector, is the inverse hyperbolic cosine function, is and is the audio-facial feature semantic metric value between them;

[0081] Use a binary function to perform clustering judgment on the set of audio-facial feature semantic metric values to obtain a set of audio-facial feature clustering coefficients. This process can be expressed as:

[0082] ;

[0083] Wherein, is and is the audio-facial feature semantic metric value between them, is corresponding audio-facial feature clustering coefficient, is a preset threshold;

[0084] Based on the set of audio-facial feature clustering coefficients, perform fine-grained cross-domain fusion on the set of audio-facial feature prior clustering center encoding vectors and the set of facial image local feature encoding vectors to obtain the multi-modal joint encoding vector of the autistic child object state, and this process can be expressed as:

[0085] ;

[0086] where, is the number of vectors in the set of facial image local feature encoding vectors, is the th facial image local feature encoding vector in the set of facial image local feature encoding vectors, is the audio-facial feature prior clustering center encoding vector, is the corresponding audio-facial feature clustering coefficient, is the multi-modal joint encoding vector of the autistic child object state.

[0087] It should be understood that the audio-facial feature prior clustering center encoding vector integrates the core correlation information of both audio and facial image modal data, but the set of facial image local feature encoding vectors also contains rich facial detail information. Through multi-modal fine-grained joint encoding, using the audio-facial feature prior clustering center encoding vector as a bridge, these facial detail information can be fused with the existing core correlation information, and a new feature representation that can more comprehensively and comprehensively reflect various modal information can be generated, that is, the multi-modal joint encoding vector of the autistic child object state, so that all clues related to the autistic child's attempt to communicate can be captured more completely. It is also worth mentioning that when actually collecting the audio and facial image data of autistic children, due to environmental factors, equipment errors, or abnormal behaviors of children themselves, there are still noises or outliers in the feature information encoded in the original data. Through multi-modal fine-grained joint encoding, those feature information that may be noises or outliers can be filtered out, focusing on the part that can best reflect the essential characteristics of children's behaviors, thereby improving the reliability of the data and the accuracy of the analysis results.

[0088] In the embodiment of the present application, the behavior state judgment module 150 is used to determine the behavior state of the autistic child object based on the multi-modal joint encoding vector of the autistic child object state, and generate an intervention feedback based on the behavior state. Specifically, Figure 4 is the block diagram of the behavior state judgment module in the autistic child intervention treatment assistance system based on intelligent interaction according to the embodiment of the present application. As Figure 4As shown, the behavior state judgment module 150 includes: a behavior state result generation unit 151, configured to input the multi-modal joint encoding vector of the autistic child object state into a behavior state judge to obtain the behavior state of the autistic child object, where the behavior state of the autistic child object is whether to attempt communication; an intervention feedback generation unit 152, configured to generate the intervention feedback in response to the behavior state of the autistic child object being an attempt to communicate.

[0089] In an embodiment of the present application, the behavior state result generation unit 151 is configured to input the multi-modal joint encoding vector of the autistic child object state into a behavior state judge to obtain the behavior state of the autistic child object, where the behavior state of the autistic child object is whether to attempt communication. Specifically, the behavior state judge in the present application is a behavior state judge based on a classifier. That is, the multi-modal joint encoding vector of the autistic child object state obtained by performing fine-grained joint on the audio feature encoding vector and the facial image feature encoding feature map is classified, so as to utilize the excellent ability of the classifier in the category task to learn the feature patterns in the multi-modal joint encoding vector and accurately classify it into the corresponding behavior state category, thereby determining whether it belongs to the category of attempting communication or not attempting communication. Particularly, in a specific embodiment of the present application, inputting the multi-modal joint encoding vector of the autistic child object state into a behavior state judge to obtain the behavior state of the autistic child object, where the behavior state of the autistic child object is whether to attempt communication, includes: performing fully connected encoding on the multi-modal joint encoding vector of the autistic child object state using the fully connected layer of the classifier to obtain a multi-modal joint fully connected encoding feature vector of the autistic child object state; inputting the multi-modal joint fully connected encoding feature vector of the autistic child object state into the Softmax classification function of the classifier to obtain the probability values of the multi-modal joint encoding vector of the autistic child object state belonging to each classification label, where the classification labels include those for indicating that the behavior state of the autistic child object is an attempt to communicate and those for indicating that the behavior state of the autistic child object is not an attempt to communicate; and determining the classification label corresponding to the maximum of the probability values as the behavior state of the autistic child object.

[0090] In a preferred example, since the audio feature encoding vector and the facial image feature encoding feature map respectively represent the audio features of the audio data of the autistic child object and the facial image semantic features of its facial expression image, when performing cross-domain fine-grained aggregation analysis based on modal priors, the feature population attributes of different modalities will have cross-domain fine-grained aggregation fairness differences based on different modal prior levels, thereby affecting the aggregation feature distribution inclusiveness of the multi-modal joint encoding vector of the autistic child object state and reducing the accuracy of the determined behavior state of the autistic child object.

[0091] Therefore, when determining the behavior state of the autistic child object based on the multi-modal joint encoding vector of the autistic child object state, this application considers first optimizing the multi-modal joint encoding vector of the autistic child object state, including the steps of:

[0092] Determine the mean value of the eigenvalues corresponding to the multi-modal joint encoding vector of the autistic child object state and the standard deviation of the eigenvalues ;

[0093] Subtract the dot product vector of the multi-modal joint encoding vector of the autistic child object state and the mean value of the eigenvalues from the product of the standard deviation of the eigenvalues to obtain the first intermediate multi-modal joint encoding vector of the autistic child object state. This process can be expressed as:

[0094] ;

[0095] Wherein, represents the multi-modal joint encoding vector of the autistic child object state, represents element-wise multiplication, represents element-wise subtraction, represents the standard deviation of the eigenvalues corresponding to, represents the mean value of the eigenvalues corresponding to, represents the first intermediate multi-modal joint encoding vector of the autistic child object state;

[0096] Subtract the dot product vector of the multi-modal joint encoding vector of the autistic child object state and the standard deviation of the eigenvalues from the product of the mean value of the eigenvalues to obtain the second intermediate multi-modal joint encoding vector of the autistic child object state. This process can be expressed as:

[0097] ;

[0098] Wherein, represents the multi-modal joint encoding vector of the autistic child object state, represents element-wise multiplication, Indicates subtraction by position point, Indicates The standard deviation of the corresponding eigenvalue, Indicates The corresponding feature mean, Indicates the intermediate vector of the multimodal joint coding of the second autistic child object state;

[0099] After multiplying the reciprocal of each bit of the intermediate vector of the multimodal joint coding of the second autistic child object state by the intermediate vector of the multimodal joint coding of the first autistic child object state, take the logarithm to the base 2 of each bit to obtain the correction vector of the multimodal joint coding of the autistic child object state. This process can be expressed as:

[0100] ;

[0101] Wherein, Indicates the intermediate vector of the multimodal joint coding of the second autistic child object state, Indicates calculation The reciprocal of each eigenvalue in, Indicates multiplication by position point, Indicates the intermediate vector of the multimodal joint coding of the first autistic child object state, Indicates the correction vector of the multimodal joint coding of the autistic child object state;

[0102] Divide the eigenvalue mean By the standard deviation of the eigenvalue Take the square root of the quotient and multiply by the weight hyperparameter, and then dot-multiply with the correction vector of the multimodal joint coding of the autistic child object state to obtain the optimized multimodal joint coding vector of the autistic child object state:

[0103] ;

[0104] Wherein, Indicates the weight hyperparameter, Indicates The standard deviation of the corresponding eigenvalue, Indicates The corresponding feature mean, Indicates addition by position point, Indicates the correction vector of the multimodal joint coding of the autistic child object state, Indicates the optimized multimodal joint coding vector of the autistic child object state;

[0105] Determine the behavior state of the autistic child object based on the optimized multimodal joint coding vector of the autistic child object state.

[0106] That is, based on the fairness difference in the attribute levels of the data population corresponding to the sequence fusion features of the multi-modal joint coding vectors of the autistic child object states, in order to enhance the aggregation inclusiveness under the feature distribution diversity of the multi-modal joint coding vectors of the autistic child object states, by using the cross-probability value constraint based on the multi-modal joint coding vectors of the autistic child object states as the representation of the interactive fairness goal, the interactive propagation of the group feature information of the multi-modal joint coding vectors of the autistic child object states is corrected, and the unified statistical feature response interaction based on the multi-modal joint coding vectors of the autistic child object states is used as the bias of the multi-level fairness goal of the feature distribution, to achieve the robust unified representation of the distribution fairness of the multi-modal joint coding vectors of the autistic child object states, form the fairness collaboration paradigm under the feature distribution framework of the multi-modal joint coding vectors of the autistic child object states, and improve the accuracy of the determined behavior states of the autistic child object.

[0107] In the embodiment of the present application, the intervention feedback generation unit 152 is configured to generate the intervention feedback in response to the behavior state of the autistic child object being an attempt to communicate. It should be understood that when an autistic child attempts to communicate, this is a positive behavior manifestation. Generating intervention feedback can timely reinforce this positive behavior, enabling the child to realize that their communication attempts are recognized and encouraged, so that autistic children can gradually establish a connection between communication behaviors and positive experiences, and thus be more willing to actively communicate. The intervention feedback here can be simple voice prompts, virtual character emotion matching responses on the screen, etc.

[0108] The following is a detailed elaboration of a specific implementation process of "generating the intervention feedback in response to the behavior state of the autistic child object being an attempt to communicate":

[0109] First, from the perspective of voice feedback, the system will be designed specifically according to the child's language ability and expression characteristics. If the child's language expression is relatively simple and basic, the voice feedback will use concise, clear, and easy-to-understand vocabulary and sentence structures. For example, when the child makes some vague syllables or simple words, the system may respond with "You did a great job. Can you tell me more in detail?" The voice tone is gentle and encouraging, and the speaking speed will also be slowed down to ensure that the child can clearly receive the information and feel positive feedback. For children with slightly stronger language abilities, the voice feedback will be more rich and guiding, such as "Your idea is very interesting. We can explore it in depth together. What do you want to do next?" In this way, it can stimulate the child to further express their thoughts and feelings and expand the depth of communication.

[0110] Secondly, the visual feedback on the screen is also an important part. The animation effects of the virtual character will be dynamically adjusted according to the child's behavior. If the child shows focused eyes and positive expressions when trying to communicate, the virtual character will respond with more lively actions, such as jumping cheerfully, waving its arms, etc. At the same time, the facial expression will show a happy and encouraging look, with wide-open eyes and an upturned mouth, sending positive emotional signals to the child. If the child shows a bit of nervousness or uneasiness, the virtual character will make soothing actions, such as approaching slowly, nodding gently, and accompanied by a gentle expression, with soft and concerned eyes, making the child feel safe and supported visually, thus relieving the nervousness and being more willing to continue the communication.

[0111] Furthermore, the adjustment of interactive content is a key part of the intervention feedback. Once it is determined that the child is trying to communicate, the system will quickly analyze the theme or interest points that the child is concerned about. For example, if the child frequently mentions a certain toy or animal during the communication, the system will immediately retrieve relevant pictures, videos, or stories related to that toy or animal from its rich material library and display them on the screen in a vivid form, accompanied by voice explanations, such as "Look, this is the little rabbit you mentioned. It has many interesting stories. Do you want to hear them?" This content adjustment based on the child's interests can greatly enhance the child's participation and willingness to communicate, making them more involved in the interaction process.

[0112] In addition, the intervention feedback also takes into account factors such as time and frequency. The system will not provide too much information or feedback at one time, but will reasonably arrange the rhythm of the feedback according to the child's attention span and reaction speed. If the child can respond actively in a short time, the system will appropriately increase the frequency of the feedback, but always keep it within the range that the child can accept; if the child needs more time to digest and respond, the system will wait patiently and give a gentle reminder at the right time, such as "Don't worry, take your time to think".

[0113] In the actual implementation process, professionals will also continuously observe and evaluate the effectiveness of the intervention feedback. They will record the child's reactions to each feedback, including aspects such as expressions, actions, and verbal responses. Through the analysis of these data, the feedback strategy of the system will be continuously optimized. If it is found that a certain feedback method fails to achieve the expected effect, professionals will promptly adjust and improve the voice, vision, or interactive content to ensure that the intervention feedback can always effectively promote the child's communication and development.

[0114] In summary, the intelligent interaction-based autism children intervention and treatment assistance system 100 according to the embodiments of the present application is elucidated. It uses AI-based data processing technology to extract audio features from the audio data of an autism child object, and at the same time extracts facial features from the facial expression images of the autism child object, so as to intelligently judge whether the behavior state of the autism child object is attempting to communicate based on the multi-modal fine-grained joint representation between the audio feature encoding features and the facial image feature encoding features. In this way, through continuous data collection, extremely subtle changes in autism children can be captured more accurately, ensuring that no behavior occurring in any instant is missed. Moreover, by combining facial expression analysis and audio feature extraction, the judgment of the behavior state of autism children can be made more precise, and then more targeted personalized intervention feedback can be generated.

[0115] Figure 5 FIG. is a flowchart of an intelligent interaction-based autism children intervention and treatment assistance method according to an embodiment of the present application. As Figure 5 shown, in the intelligent interaction-based autism children intervention and treatment assistance method, it includes: S110, obtaining the facial expression images of an autism child object collected by a camera and the audio data of the autism child object collected by a microphone; S120, extracting audio features from the audio data to obtain an audio feature encoding vector; S130, extracting facial features from the facial expression images to obtain a facial image feature encoding feature map; S140, performing audio-image multi-modal data fine-grained joint on the audio feature encoding vector and the facial image feature encoding feature map to obtain an autism child object state multi-modal joint encoding vector; S150, determining the behavior state of the autism child object based on the autism child object state multi-modal joint encoding vector, and generating intervention feedback based on the behavior state.

[0116] Here, those skilled in the art can understand that the specific operations of each step in the above intelligent interaction-based autism children intervention and treatment assistance method have been introduced in detail in the description of the intelligent interaction-based autism children intervention and treatment assistance system above with reference to Figures 1 to 4 and thus, the repeated description thereof will be omitted.

[0117] In summary, the autism child intervention and treatment assistance method based on intelligent interaction according to the embodiments of the present application is elucidated. It uses AI-based data processing technology to extract audio features from the audio data of the autism child object, and at the same time extracts facial features from the facial expression images of the autism child object. Based on this, it intelligently judges whether the behavior state of the autism child object is attempting to communicate based on the multi-modal fine-grained joint representation between the audio feature encoding features and the facial image feature encoding features. In this way, through continuous data collection, extremely subtle changes in autism children can be captured more accurately, ensuring that no behavior occurring at any moment is missed. Moreover, by combining facial expression analysis and audio feature extraction, the judgment of the behavior state of autism children can be made more precise, and then more targeted personalized intervention feedback can be generated.

Claims

1. An intelligent interactive intervention and treatment auxiliary system for children with autism, characterized by: include: A child object facial audio acquisition module, used to acquire the facial expression image of the autistic child object acquired by the camera and the audio data of the autistic child object acquired by the microphone; An audio feature extraction module, used for extracting audio features from the audio data to obtain an audio feature encoding vector; A facial feature extraction module, used for extracting facial features from the facial expression image to obtain a facial image feature encoding feature map; An audio-image multimodal joint module is used to perform an audio-image multimodal data fine-grained joint of the audio feature coding vector and the facial image feature coding feature map to obtain a multimodal joint coding vector of the object state of an autistic child, wherein the audio-image multimodal joint module includes: a facial image decomposition unit, used to perform feature decoupling and feature flattening on the facial image feature coding feature map to obtain a set of facial image local feature coding vectors; a clustering fusion unit, used to perform kernel feature calculation and fine-grained clustering fusion on the set of the audio feature coding vector and the facial image local feature coding vector to obtain the multimodal joint coding vector of the object state of an autistic child; A behavior state judgment module, used to determine the behavior state of the autistic child object based on the multimodal joint coding vector of the state of the autistic child object, and generate intervention feedback based on the behavior state; Wherein, the clustering fusion unit comprises: The facial image kernel feature extraction subunit is used to extract the facial image kernel feature from the set of facial image local feature coding vectors to obtain a facial image kernel feature coding vector, which is expressed as: ; in, is the set of local feature encoding vectors of facial images, , are respectively the first , facial image local feature encoding vector, For Perform kernel feature extraction, represents the one-norm of the eigenvector, represents the number of vectors in the set of local feature encoding vectors of the facial image minus one, yes The corresponding facial image semantic difference value, Expressed as a natural constant The exponential function with base , is the number of vectors in the set of local feature encoding vectors of the facial image, is the facial image kernel feature encoding vector; The audio-facial feature prior cluster center coding vector determining subunit is used to determine the audio-facial feature prior cluster center coding vector based on the audio feature coding vector and the facial image kernel feature coding vector, expressed as: ; in, is the facial image kernel feature encoding vector, is the audio feature encoding vector, For cascade operation, is the weight transformation matrix, is the bias vector, is the audio-facial feature prior cluster center encoding vector; The multimodal fine-grained joint subunit is used to perform multimodal fine-grained joint encoding on the set of the audio-facial feature prior clustering center encoding vector and the facial image local feature encoding vector to obtain the multimodal joint encoding vector of the autistic child object state.

2. The intelligent interaction-based autism children intervention and treatment auxiliary system according to claim 1, characterized in that: The audio feature extraction module comprises: An audio data wavelet transform unit, used for performing discrete wavelet transform on the audio data to obtain a time-frequency diagram of the audio data; An audio feature generating unit is used to extract audio features from the audio data time-frequency diagram using a feature extractor based on a CNN-GRU hybrid model to obtain the audio feature encoding vector.

3. The intelligent interaction-based autism children intervention and treatment auxiliary system according to claim 2, characterized in that: The facial feature extraction module is used to extract facial features from the facial expression image using a facial feature extractor based on a dilated pyramid model to obtain the facial image feature encoding feature map.

4. The intelligent interaction-based autism children intervention and treatment auxiliary system according to claim 3 is characterized in that: The facial image decomposition unit is used to: Performing feature decoupling on the facial image feature coding feature map along the channel dimension of the facial image feature coding feature map to obtain a set of facial image feature coding local feature matrices; Each facial image feature encoding local feature matrix in the set of facial image feature encoding local feature matrices is feature flattened to obtain the set of facial image local feature encoding vectors.

5. The intelligent interaction-based autism children intervention and treatment auxiliary system according to claim 4, characterized in that: The multimodal fine-grained joint subunit is used to: The Poincare distances between the audio-facial feature prior cluster center encoding vector and each facial image local feature encoding vector in the set of facial image local feature encoding vectors are calculated respectively to obtain a set of audio-facial feature semantic metric values. The process can be expressed as: ; in, is the first in the set of local feature encoding vectors of the facial image facial image local feature encoding vector, is the audio-facial feature prior cluster center encoding vector, is the square of the Euclidean norm of the vector, is the inverse hyperbolic cosine function, for and The audio-facial feature semantic metric between them; Using a binary function to perform clustering judgment on the set of audio-facial feature semantic measurement values ​​to obtain a set of audio-facial feature clustering coefficients; Based on the set of audio-facial feature clustering coefficients, the set of audio-facial feature prior clustering center encoding vectors and the set of facial image local feature encoding vectors are fine-grained cross-domain fused to obtain the multimodal joint encoding vector of the autistic child object state.

6. The intelligent interaction-based autism children intervention and treatment auxiliary system according to claim 5, characterized in that: The behavior state judgment module includes: A behavior state result generating unit, used for inputting the multimodal joint coding vector of the state of the autistic child object into a behavior state judgement unit to obtain the behavior state of the autistic child object, wherein the behavior state of the autistic child object is whether to attempt to communicate; The intervention feedback generating unit is used to generate the intervention feedback in response to the behavior state of the autistic child subject being an attempt to communicate.

7. According to the intelligent interaction-based autism children intervention treatment auxiliary system according to claim 6, the behavior state determiner is a classifier-based behavior state determiner.

8. An intelligent interactive intervention and treatment aid method for children with autism, characterized in that: include: Acquire a facial expression image of an autistic child object captured by a camera and audio data of the autistic child object captured by a microphone; Extracting audio features from the audio data to obtain an audio feature encoding vector; Extracting facial features from the facial expression image to obtain a facial image feature encoding feature map; The audio feature coding vector and the facial image feature coding feature map are subjected to an audio-image multimodal data fine-grained union to obtain a multimodal joint coding vector of the object state of an autistic child, including: performing feature decoupling and feature flattening on the facial image feature coding feature map to obtain a set of facial image local feature coding vectors; performing kernel feature calculation and fine-grained clustering fusion on the set of the audio feature coding vector and the facial image local feature coding vector to obtain the multimodal joint coding vector of the object state of an autistic child; Determining the behavioral state of the autistic child subject based on the multimodal joint encoding vector of the state of the autistic child subject, and generating intervention feedback based on the behavioral state; The method of performing kernel feature calculation and fine-grained clustering fusion on the set of the audio feature coding vector and the facial image local feature coding vector to obtain the multimodal joint coding vector of the autistic child object state includes: Perform facial image kernel feature extraction on the set of facial image local feature coding vectors to obtain a facial image kernel feature coding vector, which is expressed as: ; in, is the set of local feature encoding vectors of facial images, , are respectively the first , facial image local feature encoding vector, For Perform kernel feature extraction, represents the one-norm of the eigenvector, represents the number of vectors in the set of local feature encoding vectors of the facial image minus one, yes The corresponding facial image semantic difference value, Expressed as a natural constant The exponential function with base , is the number of vectors in the set of local feature encoding vectors of the facial image, is the facial image kernel feature encoding vector; Based on the audio feature coding vector and the facial image kernel feature coding vector, an audio-facial feature prior clustering center coding vector is determined, which is expressed as: ; in, is the facial image kernel feature encoding vector, is the audio feature encoding vector, For cascade operation, is the weight transformation matrix, is the bias vector, is the audio-facial feature prior cluster center encoding vector; The set of the audio-facial feature prior cluster center coding vector and the facial image local feature coding vector is multimodally fine-grained jointly encoded to obtain the multimodal joint coding vector of the autistic child object state.

9. The intelligent interaction-based intervention and treatment auxiliary method for autistic children according to claim 8, characterized in that: Extracting audio features from the audio data to obtain an audio feature encoding vector includes: Performing discrete wavelet transform on the audio data to obtain a time-frequency graph of the audio data; A feature extractor based on a CNN-GRU hybrid model is used to extract audio features from the time-frequency graph of the audio data to obtain the audio feature encoding vector.

Citation Information

Patent Citations

  • Emotion recognition method and device, equipment and storage medium

    CN116013369A

  • Language evaluation method and device for autistic children and medium

    CN119028594A

  • Unmanned aerial vehicle public safety data management platform

    CN119131634A