Multi-modal monitoring method for autistic children

Through the multimodal monitoring method, combined with the fusion analysis of movement, voice and facial features, real-time monitoring and interactive guidance of behavior and emotions of autistic children is achieved, solving the problem of inability to prevent abnormal behaviors in a timely manner in the existing technology, and improving the accuracy and reliability of monitoring.

CN120234749APending Publication Date: 2025-07-01ZHEJIANG NORMAL UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510714147.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The prior art cannot achieve real-time observation and prompt interaction and guidance during abnormal behavior in the monitoring of children with autism, resulting in difficult prevention of abnormal behaviors.

Method used

Multimodal monitoring method is adopted to obtain video and voice data through image acquisition devices and microphones, action features are extracted using OpenPose algorithm, voice features are extracted by VGGish network, facial features are extracted by CNN network, and comprehensive features are generated through multimodal attention fusion mechanism, combined with LSTM network to analyze behavior abnormalities, and large language models are used to perform voice interaction and alarms.

Benefits of technology

It realizes comprehensive and meticulous monitoring of behaviors and emotions of autistic children, improves the accuracy and reliability of detection of abnormal behaviors, provides real-time interactive comfort, reduces the harm of abnormal behaviors, and improves monitoring efficiency and scientificity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234749A_ABST
    Figure CN120234749A_ABST
Patent Text Reader

Abstract

The invention aims to solve the problem that existing autistic children are difficult to effectively pacify in time in monitoring. According to the multi-mode monitoring method for the autistic children, the action features, the facial features and the voice features are fused, the comprehensive features are finally obtained, and compared with a single-mode data processing mode in the prior art, behavior information and emotion information of the autistic children can be captured more comprehensively and meticulously; the accuracy and reliability of behavior analysis are remarkably improved, the reliability of voice interaction of the large language model is also improved, the voice interaction of the large language model is combined, the blank that only warning can be given out and real-time interactive pacifying is lacked when the abnormal behavior occurs is effectively filled, and the voice interaction of the large language model is improved. Primary emotion pacification and behavior guidance can be carried out on children before guardians arrive, and harm possibly brought by abnormal behaviors is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of guardianship, and particularly to a multi-modal guardianship method for autistic children. Background Art

[0002] The behavioral characteristics of autistic children are relatively special and often require continuous attention and guardianship. Among them, real-time monitoring and early warning are a major direction in the field of autistic children's guardianship. Existing technologies include methods of monitoring physiological indicators such as heart rate, body temperature, sweat, and arm movement through smart bracelets, such as the content shown in the patent with the application publication number CN119498843A; there are also methods of predicting the emotional outbursts of autistic children through machine learning models, such as the content shown in the patent with the application publication number CN119830223A; there are also methods of monitoring the changes in brain waves during different tasks through devices such as wireless dry electrode electroencephalogram recorders similar to head-mounted headphones to achieve the auxiliary diagnosis of autism, such as the content shown in the patent with the application publication number CN118506988A. However, existing technologies usually mostly adopt the method of guardianship of personnel and alarm to notify personnel for handling. Although to a certain extent, they can reflect and predict some abnormal situations, in most cases, it is only after the behavior occurs that personnel are notified to handle it or an alarm is given, and it is impossible to timely and effectively dredge the abnormal behavior of the guarded personnel to prevent the occurrence of some non-permitted behaviors; on the other hand, the existing methods lack automatic and timely interaction with the guarded personnel, resulting in the guarded personnel being more likely to have abnormal behaviors. Therefore, a guardianship method that can observe the guarded personnel in real time and conduct interactive guidance during abnormal behaviors is needed. Summary of the Invention

[0003] The purpose of the present invention is to solve the deficiencies of the prior art and provide a multi-modal guardianship method for autistic children.

[0004] To solve the above problems, the present invention adopts the following technical solutions: A multi-modal guardianship method for autistic children, comprising the following steps: Step 1: Obtain video data and voice data through an image acquisition device and a microphone, and preprocess the data; Step 2: After extracting key points from the video data through the OpenPose algorithm, obtain action features; after processing the voice data through the VGGish network, obtain voice features; after processing the video data through the CNN network, obtain facial features; Step 3: After aligning the action features, voice features, and facial features through feature alignment processing, fuse them according to the multi-modal attention fusion mechanism to obtain comprehensive features; Step 4: Input the comprehensive features into the LSTM network to analyze whether the child's behavior is abnormal; if it is abnormal, determine the type of abnormal behavior, mark it for alarm, and then proceed to the next step; otherwise, directly proceed to the next step; Step 5: Process the comprehensive features through speech recognition to generate text, and determine whether the text content is speech interacting with the large language model through semantic analysis; if it is speech text interacting with the large language model, after being processed by the large language model to generate a corresponding response, generate speech to interact with the child; if it is non-interactive speech, proceed to the next step; Step 6: Determine whether there is an alarm mark; if there is an alarm mark, generate corresponding speech according to the type of abnormal behavior by the large language model and play it; if there is no alarm mark, directly proceed to the next step; Step 7: Store the data and end the steps.

[0005] Further, the video data preprocessing in step 1 includes identifying the child target in the video data, automatically numbering it, and frame-selectively displaying the video data with the corresponding number.

[0006] Further, the preprocessing of the voice data in step 1 includes pre-emphasis, framing, and windowing operations on the original voice.

[0007] Further, the action features in step 2 detect key points in the images in the video data through the OpenPose algorithm, extract the key points in each frame of the image; according to the human bone structure, connect each pair of key points; by processing each pair of key points, obtain the action features 。

[0008] Further, the voice features in step 2 perform a fast Fourier transform (FFT) on the preprocessed voice data to convert the time-domain signal into a frequency-domain signal; subsequently, through a Mel filter bank and logarithmic operations, convert the spectral energy in the frequency domain into energy at Mel frequencies to obtain a Mel spectrogram; finally, input the Mel spectrogram into the VGGish network trained on the AudioSet dataset for processing to obtain the voice features 。

[0009] Further, the process of obtaining the facial features in step 2 includes reading the video data frame by frame and converting it into a time-continuous initial picture; subsequently, perform face detection on each frame of the picture, if there is a face in the picture, retain it, otherwise delete it; crop the face part in the retained picture, and unify the size of the cropped picture through picture scaling; perform normalization and grayscale processing on the processed picture, and output the picture sequence in time series; finally, input the picture sequence into the CNN network for facial feature extraction to obtain the facial features 。

[0010] Further, the process of obtaining the comprehensive feature in step 3 includes: Step 31: First, obtain the speech-face multimodal attention emotion feature vector, speech-action multimodal attention emotion feature vector, and action-face multimodal attention emotion feature vector according to the multimodal attention mechanism; Step 32: Subsequently, concatenate the six groups of cross-modal interaction emotion features obtained in step 31 to obtain three single-modal emotion features, including speech emotion feature , action emotion feature and face emotion feature ; Step 33: Concatenate and fuse the speech emotion feature , action emotion feature and face emotion feature to obtain the comprehensive feature vector .

[0011] Further, the speech-face multimodal attention emotion feature vector obtained in step 31 is expressed as: , ,

[0012] where is the matrix after the speech feature completes linear transformation through the parameter matrix, is the trainable parameter vector matrix; is the face feature completes linear transformation through the parameter matrix, is the trainable parameter vector matrix; represents the adjustment factor of the attention mechanism; The speech-action multimodal attention emotion feature vector is expressed as: , ,

[0013] where is the action feature, is the matrix after the action feature completes linear transformation through the parameter matrix, is the trainable parameter vector matrix; The action-face multimodal attention emotion feature vector is expressed as: , ,

[0014] Further, the emotional features of the three unimodals in step 32 are represented as: , , ,

[0015] Among them, represents the splicing operation, respectively represent the three emotional features after fusion; In step 33, and are cascaded and fused to obtain its comprehensive feature vector , which is represented as: , Among them, represents the vector cascading operation.

[0016] A multimodal monitoring system for autistic children includes: An acquisition module that collects video data and voice data of autistic children, preprocesses the data, and numbers each child; An action feature extraction module that processes the video data through OpenPose to obtain key points in each frame and obtains action features; A voice feature extraction module that processes the voice data through the VGGish network to obtain voice features; a facial feature extraction module that processes the video data through the CNN network to obtain facial features; A feature fusion module that performs feature alignment on the three features and then realizes feature fusion through a multimodal attention mechanism to obtain a comprehensive feature vector; A behavior analysis module that processes the comprehensive feature vector through LSTM to judge whether the child has abnormal behavior. If abnormal behavior occurs, an abnormal behavior alarm message will be sent; A large language model voice interaction module that processes the comprehensive feature vector through speech recognition technology to generate corresponding text, then uses semantic analysis technology to analyze the text information, judges the text for interaction with the large language model, and after being processed by the large language model to generate a corresponding reply, generates corresponding voice to interact with the child, and when the child has abnormal behavior, generates voice according to the child's facial features and the types of abnormal behavior; A monitoring display module that sets a monitoring frame for each child according to the child's number, and changes the color of the monitoring frame to mark when the child has abnormal behavior to notify the monitoring personnel that the child has abnormal behavior; A data storage and analysis module that stores the child's behavior data and voice interaction records, and then generates an analysis report through data analysis.

[0017] The beneficial effects of the present invention are as follows: By applying the OpenPose algorithm to accurately extract the action features of children from video data, using a CNN network to obtain the facial features of children from video data, and also using the VGGish network to obtain the voice features from voice data, and fusing the action features, facial features and voice features, the comprehensive features are finally obtained. Compared with the single-modal data processing method in the prior art, it can capture the behavior information and emotional information of autistic children more comprehensively and meticulously, not only significantly improving the accuracy and reliability of behavior analysis, but also improving the reliability of the voice interaction of the large language model; The voice interaction through the large language model effectively fills the gap in the prior art that only warnings can be issued when abnormal behaviors occur and there is a lack of real-time interactive comfort. It can conduct preliminary emotional comfort and behavior guidance for children before the arrival of the guardians, reducing the possible harm caused by abnormal behaviors; Finally, through the color change of the monitoring frame of the monitoring display module, the guardians can quickly know which child has abnormal behaviors and comfort them in time. Description of the Drawings

[0018] Figure 1 It is a schematic flowchart of the implementation manner of the multi-modal monitoring method in Embodiment 1; Figure 2 It is a processing flowchart of the action features in Embodiment 1; Figure 3 It is a processing flowchart of the voice features in Embodiment 1; Figure 4 It is a processing flowchart of the facial features in Embodiment 1; Figure 5 It is a schematic diagram of the cascaded fusion of the three single-modal emotional features in Embodiment 1; Figure 6 It is a diagram of the multi-modal monitoring device in Embodiment 1 and the connection between them. Detailed Embodiment

[0019] The following uses specific specific examples to illustrate the implementation manner of the present invention. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0020] It should be noted that the illustrations provided in the following embodiments only schematically illustrate the basic concept of the present invention. Therefore, only the components related to the present invention are shown in the figures, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0021] As Figures 1 to 5 shown, a multi-modal monitoring method for autistic children includes the following steps: Step 1: Obtain video data and voice data through an image acquisition device and a microphone, and preprocess the data; Step 2: After extracting key points from the video data through the OpenPose algorithm, obtain action features; after processing the voice data through the VGGish network, obtain voice features ; after processing the video data through the CNN network, obtain facial features ; Step 3: Align the action features , voice features and facial features through feature alignment processing, and then fuse them according to the multi-modal attention fusion mechanism to obtain comprehensive features ; Step 4: Input the comprehensive features into the LSTM network to analyze whether the child's behavior is abnormal; if it is abnormal, judge the type of abnormal behavior, and after making an alarm mark, enter the next step; otherwise, directly enter the next step; Step 5: Convert the comprehensive features through speech recognition processing to generate text, and judge whether the text content is speech interacting with the large language model through semantic analysis; if it is speech text interacting with the large language model, after generating a corresponding reply through the large language model processing, generate speech to interact with the child; if it is non-interactive speech, enter the next step; Step 6: Judge whether there is an alarm mark; if there is an alarm mark, generate corresponding speech according to the type of abnormal behavior by the large language model and play it; if there is no alarm mark, directly enter the next step; Step 7: Store the data and end the steps.

[0022] The video data preprocessing in Step 1 includes identifying the child target in the video data, automatically numbering it, and generating different numbered monitoring frames in the video data for display to facilitate observing the position of the child in the video data.

[0023] The preprocessing of the speech data in step 1 includes pre-emphasis, framing and windowing operations on the original speech. The high-resolution cameras are installed at multiple locations in the children's activity area to collect video data with a resolution of 1080P and a frame number of 30 frames, and the children's speech data is collected in real time through a microphone with a sampling rate of 16kHz.

[0024] To obtain the action features in step 2, the video data must first be pre-processed by decoding, framing, denoising, normalizing, and adjusting the frame rate. The OpenPose algorithm is then used to detect key points in the images in the video data, and the key points in each frame are extracted. The key point pairs are connected according to the human skeletal structure. The parameters of OpenPose that define the connection relationship between the key points of various parts of the human body are used to calculate and process each pair of key points, including the direction, length, and angle between the key points, to obtain the action features. It should be noted that the training data set of the OpenPose algorithm can collect current video data for manual annotation, or use the existing COCO data set, MPII Human Pose data set, etc.

[0025] When calculating the direction, length and angle between key point pairs, you first need to obtain the coordinates of the key point pairs, including the coordinates of the starting point key point , the coordinates of the end point key point , the calculation formula of the vector is: Vector = , The length of the vector is: Length = , For a limb structure composed of three key points, such as the combined action of the shoulder, elbow, and wrist, it is necessary to further calculate the angle of the middle key point. , , ,vector and The angle calculation formula is: , angle for: , The vectors and angles obtained by the above calculations are combined to form key point connection features. These key point connection features can fully describe the posture and movement of the human body, such as the extension degree of the arms and the bending angle of the legs. It is represented as a feature vector, which contains information such as the length and angle of each connection vector.

[0026] The speech features in step 2 are obtained by performing a Fast Fourier Transform (FFT) on the preprocessed speech data to convert the time-domain signal into a frequency-domain signal. Subsequently, through a Mel filter bank and logarithmic operation, the spectral energy in the frequency domain is converted into energy at Mel frequencies to obtain a Mel spectrogram. Finally, the Mel spectrogram is input into a VGGish network trained on the AudioSet dataset for processing to obtain the speech features.

[0027] Specifically, the collected speech data is first preprocessed, including operations of pre-emphasis, framing, and windowing on the original speech data in the speech dataset. Among them, since the high-frequency part of the speech signal usually attenuates faster than the low-frequency part, the pre-emphasis step can enhance the energy of the high-frequency part and make the spectrum of the signal flatter. On the other hand, the speech signal is a non-stationary signal whose characteristics change over time. The framing operation divides the speech signal into multiple short-time frames, and the signal characteristics within each frame are considered approximately stationary, enabling independent processing and analysis of each frame, thereby more accurately capturing the dynamic characteristics of the speech signal. The windowing operation is to add a Hamming window, which can reduce the discontinuity of the signal at the frame edges and make the signal at both ends of the frame smoothly transition to zero. After the preprocessing is completed, a Fast Fourier Transform (FFT) is performed on each windowed frame to convert the time-domain signal into a frequency-domain signal, and the spectral energy is converted into energy at Mel frequencies through a Mel filter bank and logarithmic operation to obtain a Mel spectrogram. Then, the obtained Mel spectrogram is input into a VGGish network trained on the AudioSet dataset for processing to obtain the speech features.

[0028] The process of obtaining facial features in step 2 includes reading the video data frame by frame and converting it into a sequence of temporally continuous initial pictures. Subsequently, face detection is performed on each frame of the picture. If a face exists in the picture, it is retained; otherwise, it is deleted. The face part in the retained picture is cropped, and the size of the cropped picture is unified through picture scaling. The processed pictures are normalized and grayscaled, and a sequence of pictures is output according to the time sequence. Finally, the sequence of pictures is input into a CNN network for facial feature extraction to obtain the facial features.

[0029] The multi-modal attention fusion mechanism in step 3 is a method that weights the features of different modalities through the interaction between modalities to obtain a more representative and relevant comprehensive feature representation. The process of obtaining the comprehensive features includes: Step 31: First, obtain the speech-face multi-modal attention emotion feature vector, speech-action multi-modal attention emotion feature vector, and action-face multi-modal attention emotion feature vector according to the multi-modal attention mechanism. Step 32: Subsequently, the six groups of cross-modal interaction emotion features obtained in Step 31 are matrix-concatenated to obtain three single-modal emotion features, including speech emotion features , action emotion features and facial emotion features; Step 33: The speech emotion features, action emotion features, and facial emotion features are cascaded and fused to obtain a comprehensive feature vector .

[0030] Taking the face-speech combination as an example in Step 31, the fusion mechanism can obtain the self-attention emotion features of speech and face, as well as the features that need to be focused on in the self-attention emotion features of speech and face; subsequently, high-probability weights will be assigned to the features that are focused on, and low-probability weights will be assigned to the features with lower attention. Specifically, the obtained speech-face multi-modal attention emotion feature vector is expressed as: , , where, , is the speech feature The matrix after linear transformation through the parameter matrix, is the trainable parameter vector matrix; it should be noted that the vector obtained through matrix linear transformation can map the features of different modalities to the same feature space, facilitating interaction; is the facial feature The matrix after linear transformation through the parameter matrix, is the trainable parameter vector matrix; represents the adjustment factor of the attention mechanism, and its role is to make The result of the function normalization is more balanced and stable, and the occurrence of extreme values is avoided; The speech-action multi-modal attention emotion feature vector is expressed as: , , where, is the speech feature The matrix after linear transformation through the parameter matrix, is the trainable parameter vector matrix; The action-face multi-modal attention emotion feature vector is expressed as: , .

[0031] The three single-modal emotion features in Step 32 are expressed as: , , , Among them, represents a splicing operation, respectively represent three kinds of emotional features after fusion; In step 33, , and are cascaded and fused to obtain their comprehensive feature vector , which is expressed as: , Among them, represents a vector concatenation operation.

[0032] In step 4, after obtaining the comprehensive feature vector , use the LSTM network classification to judge whether the child has abnormal behaviors and the types of abnormal behaviors; among them, the training data set of the LSTM network comes from the actually collected behavior video data of autistic children and the publicly available child behavior data sets, such as the COCO, MPII databases, etc. The training data set of the LSTM network contains more than 10,000 samples; the samples used for training also need to be manually annotated by personnel, and the annotation content includes the coordinate information of key points and the corresponding behavior categories, such as normal behaviors and types of abnormal behaviors. In this example, the abnormal behaviors mainly include self-harm behaviors, aggressive behaviors, hyperactivity, and stereotyped behaviors. It should be noted that when it is detected that the child has abnormal behaviors, an alarm mark will be made, including controlling the monitoring frame displaying real-time video data to change color for warning. In this example, the green monitoring frame is converted to red to prompt the guardians.

[0033] The large language model in step 5 adopts GPT. After the comprehensive feature vector is processed by the existing speech recognition technology to generate text, semantic analysis is used to judge whether it is a voice interacting with GPT; for the interactive voice, GPT generates an automatic text reply and generates voice to interact with the child.

[0034] When there is an alarm mark for abnormal behavior in step 6, corresponding voices are generated according to the types of abnormal behaviors and the facial features of the child to soothe and guide the child; for example: When self-harm behaviors are detected, such as detecting head banging, long-term scratching, etc., the facial features will be analyzed simultaneously to judge the emotional state of the child; if a painful or anxious expression is detected, the system will generate soothing voices, such as "Don't hurt yourself, the teacher will come to help you", etc.; When an aggressive behavior is detected, such as a child pushing or hitting someone, the system will combine facial expression analysis to determine whether the child is in an angry or excited state; based on the results of facial feature analysis, the system will generate guiding voices, such as "Don't hit people. We can play together", etc.; When an overactive behavior is detected, such as a child running quickly or jumping for a long time in the classroom, and at the same time, through facial feature recognition, it is identified that the child is in an excited or overactive emotional state, a reminder voice will be generated, such as "Please sit still and don't run around", etc.; When a stereotyped behavior is detected, such as repetitive hand movements or fixed body postures, the system will combine facial feature analysis to determine whether the child is focused on the current behavior, and then generate guiding voices, such as "We can try a new game to see if it's more interesting", etc.

[0035] As Figure 6 shown, a multimodal monitoring system for children with autism includes: An acquisition module that collects video data and voice data of children with autism, preprocesses the data, and numbers each child; An action feature extraction module that processes the video data through OpenPose to obtain key points in each frame and obtains action features; A voice feature extraction module that processes the voice data through the VGGish network to obtain voice features; a facial feature extraction module that processes the video data through the CNN network to obtain facial features; A feature fusion module that performs feature alignment processing on the three features and then realizes feature fusion through a multimodal attention mechanism to obtain a comprehensive feature vector; A behavior analysis module that processes the comprehensive feature vector through LSTM to determine whether a child exhibits abnormal behavior. If abnormal behavior occurs, an abnormal behavior alarm message will be sent; A large language model voice interaction module that processes the comprehensive feature vector through speech recognition technology to generate corresponding text, then uses semantic analysis technology to analyze the text information, determines the text for interaction with the large language model, and after being processed by the large language model to generate a corresponding reply, generates corresponding voice to interact with the child, and when the child exhibits abnormal behavior, generates voice according to the child's facial features and the types of abnormal behavior; A monitoring display module that sets a monitoring frame for each child according to the child's number, and when a child exhibits abnormal behavior, the monitoring frame changes color to mark to notify the monitoring personnel that the child has abnormal behavior; A data storage and analysis module that stores the child's behavior data and voice interaction records, and then generates an analysis report through data analysis.

[0036] In the implementation process, the action features of children are accurately extracted from video data by using the OpenPose algorithm, the facial features of children are obtained from video data by using the CNN network, and the voice features are obtained from voice data by using the VGGish network. Then, the action features, facial features and voice features are fused to finally obtain comprehensive features. Compared with the single-modal data processing method in the prior art, this series of operations can capture the behavior information and emotional information of autistic children more comprehensively and meticulously, not only significantly improving the accuracy and reliability of behavior analysis, but also enhancing the reliability of GPT voice interaction. In addition, through GPT voice interaction, the gap in the prior art that only warnings can be issued when abnormal behaviors occur and there is a lack of real-time interactive comfort is effectively filled. It can conduct preliminary emotional comfort and behavior guidance for children before the arrival of guardians, reducing the potential harm caused by abnormal behaviors. Finally, through the color change of the monitoring frame of the monitoring display module, guardians can quickly know which child has abnormal behaviors and comfort them in time. Generally speaking, the present invention not only improves the monitoring efficiency, reduces the work burden of guardians, enhances the scientific nature and reliability of monitoring, helps with the behavior correction and rehabilitation training of autistic children, improves their quality of life and social skills, but also promotes social attention and support for autistic children, having important social value and application prospects.

[0037] The above description is only a specific example of the present invention and does not constitute any limitation to the present invention. Obviously, for professionals in the field, after understanding the content and principle of the present invention, various modifications and changes in form and details may be made without departing from the principle and structure of the present invention. However, these corrections and changes based on the idea of the present invention are still within the protection scope of the claims of the present invention.

Claims

1. A multimodal monitoring method for children with autism, characterized in that It includes the following steps: Step 1: Obtain video data and voice data through an image acquisition device and a microphone, and preprocess the data; Step 2: After extracting key points from the video data through the OpenPose algorithm, obtain action features; after processing the voice data through the VGGish network, obtain voice features; after processing the video data through the CNN network, obtain facial features; Step 3: After aligning the action features, voice features, and facial features through feature alignment processing, fuse them according to the multi-modal attention fusion mechanism to obtain comprehensive features; Step 4: Input the comprehensive features into the LSTM network to analyze whether the behavior of the child is abnormal; if it is abnormal, judge the type of abnormal behavior, and after making an alarm mark, enter the next step; otherwise, directly enter the next step; Step 5: Convert the comprehensive features through speech recognition processing to generate text, and judge whether the text content is a voice for interacting with the large language model through semantic analysis; if it is a voice text for interacting with the large language model, after being processed by the large language model to generate a corresponding reply, generate voice to interact with the child; if it is a non-interactive voice, enter the next step; Step 6: Judge whether there is an alarm mark; if there is an alarm mark, generate and play the corresponding voice according to the type of abnormal behavior by the large language model; If there is no alarm mark, directly enter the next step; Step 7: Store the data and end the steps.

2. The multimodal monitoring method for autistic children according to claim 1, wherein, The preprocessing of the video data in Step 1 includes identifying the child target in the video data, automatically numbering it, and frame-selecting and displaying the video data with the corresponding number.

3. The multimodal monitoring method for autistic children according to claim 1, wherein, The preprocessing of the voice data in Step 1 includes pre-emphasizing, framing, and windowing the original voice.

4. The multimodal monitoring method for children with autism according to claim 2, wherein The action features in Step 2 are obtained by detecting key points in the images in the video data through the OpenPose algorithm, extracting the key points in each frame of the image; according to the human bone structure, connecting pairs of key points; and obtaining action features by processing each pair of key points.

5. A multimodal monitoring method for autistic children according to claim 3, characterized in that, The voice features in Step 2 are obtained by performing a fast Fourier transform (FFT) on the preprocessed voice data to convert the time-domain signal into a frequency-domain signal; then, through a Mel filter bank and logarithmic operation, converting the spectral energy in the frequency domain into energy at Mel frequencies to obtain a Mel spectrogram; finally, inputting the Mel spectrogram into the VGGish network trained on the AudioSet dataset for processing to obtain voice features.

6. The multimodal monitoring method for children with autism according to claim 2, characterized in that, The process of obtaining facial features in Step 2 includes reading the video data frame by frame and converting it into a temporally continuous initial picture; then, performing face detection on each frame of the picture, retaining the picture if there is a face in it, and deleting it otherwise; cropping the face part in the retained picture and unifying the size of the cropped picture through picture scaling; performing normalization and grayscale processing on the processed picture, and outputting a picture sequence according to the time sequence; Finally, input the picture sequence into the CNN network for facial feature extraction to obtain facial features.

7. A multimodal monitoring method for children with autism according to claim 1, characterized in that, The process of obtaining comprehensive features in Step 3 includes: Step 31: First, obtain the speech-face multi-modal attention emotion feature vector, speech-action multi-modal attention emotion feature vector, and action-face multi-modal attention emotion feature vector according to the multi-modal attention mechanism; Step 32: Subsequently, the six groups of cross-modal interaction emotion features obtained in Step 31 are matrix-concatenated to obtain three types of unimodal emotion features, including speech emotion features , action emotion features and facial emotion features ; Step 33: Cascade and fuse the speech emotion features , action emotion features and facial emotion features to obtain a comprehensive feature vector .

8. The multimodal monitoring method for autistic children according to claim 7, characterized in that, The speech-face multi-modal attention emotion feature vector obtained in step 31 is expressed as: , , Among them, , is the matrix after linear transformation completed by the parameter matrix for voice features , and , is the trainable parameter vector matrix; , is the matrix after linear transformation completed by the parameter matrix for facial features , and , is the trainable parameter vector matrix; represents the adjustment factor of the attention mechanism; The speech-action multi-modal attention emotion feature vector is expressed as: , , Among them, , is a voice feature, which is a matrix after linear transformation through a parameter matrix, , is a trainable parameter vector matrix; The action-face multi-modal attention emotion feature vector is expressed as: , 。 9. The multimodal monitoring method for autistic children according to claim 8, wherein, The emotion features of the three single modalities in step 32 are expressed as: , , , Among them, represents a splicing operation, respectively represent three emotional features after fusion; In step 33, , and are cascaded and fused to obtain their comprehensive feature vector , expressed as: , Among them, represents a vector concatenation operation.

10. A multimodal monitoring system for children with autism, characterized in that, Including: An acquisition module that collects video data and speech data of autistic children, preprocesses the data, and assigns a number to each child; An action feature extraction module that processes the video data through OpenPose to obtain key points in each frame and obtains action features; A speech feature extraction module that processes the speech data through the VGGish network to obtain speech features; A face feature extraction module that processes the video data through the CNN network to obtain face features; A feature fusion module that performs feature alignment on the three features and then realizes feature fusion through the multi-modal attention mechanism to obtain a comprehensive feature vector; A behavior analysis module that processes the comprehensive feature vector through LSTM to judge whether the child shows abnormal behavior. If abnormal behavior occurs, an abnormal behavior alarm message will be sent; A large language model speech interaction module that processes the comprehensive feature vector through speech recognition technology to generate corresponding text, then uses semantic analysis technology to analyze the text information, judges the text for interacting with the large language model, and after being processed by the large language model to generate a corresponding reply, generates corresponding speech to interact with the child, and when the child shows abnormal behavior, generates speech according to the child's face features and the type of abnormal behavior; A guardianship display module that sets a guardianship frame for each child according to the child's number, and changes the color of the guardianship frame to mark when the child shows abnormal behavior to notify the guardianship personnel that the child has abnormal behavior; A data storage and analysis module that stores the child's behavior data and speech interaction records, and then generates an analysis report through data analysis.

Citation Information

Patent Citations

  • Mobile terminal child visual attention abnormity screening method based on multi-modal data learning

    CN115761908A

  • Autistic child rehabilitation training management method and system based on artificial intelligence

    CN118888080A

  • Language evaluation method and device for autistic children and medium

    CN119028594A

  • Autism multi-source fusion diagnosis method, device and equipment and storage medium

    CN119132564A

  • Children autism tendency detection method based on voice MFCC characteristics

    CN119170056A