Autism early screening system based on human-computer interaction

Multimodal data is collected through human-computer interaction system, and early screening of autism is used to use neural network models, solving the problems of adaptability and efficiency in the existing technology, and achieving efficient and accurate screening of autism.

CN114974572BActive Publication Date: 2025-08-19JINAN ZHONGKE UBIQUITOUS INTELLIGENT COMPUTING RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210631013.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-06
Publication Date
2025-08-19
Estimated Expiration
2042-06-06

AI Technical Summary

Technical Problem

The existing early screening methods for autism rely on foreign normal models, long measurement time, require guardian observation, rely on high qualifications of practitioners, low participation in children, and single data modes, making it difficult to achieve high reliability and efficient screening.

Method used

A screening system based on human-computer interaction is adopted to obtain multimodal data through the data acquisition module, and autism diagnosis is performed using neural network models, including limb backbone detection, body language recognition, face detection, expression recognition, iris detection, etc. The ability of the subjects is evaluated in combination with interactive scenarios and output screening results.

Benefits of technology

It has achieved local adaptive and efficient screening, shortened the test time to 30 minutes, reduced dependence on guardians, improved children's participation, improved screening accuracy and data quality, and the model can be updated quickly iteratively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114974572B_ABST
    Figure CN114974572B_ABST
Patent Text Reader

Abstract

The present invention discloses an early screening system for autism based on human-computer interaction, comprising: a data acquisition module, which collects interactive action videos and audios of subjects through human-computer interaction; a data perception module, which performs limb backbone detection and body language recognition on the interactive action video data; performs face detection, facial key point detection and expression recognition on the interactive action video data; performs eye key point detection on the interactive action video data, and estimates the line of sight direction in combination with the facial key points; performs iris detection on the interactive action video data, and performs face distance measurement based on the iris size; a data analysis module, which counts the number of correct responses of the subjects based on the interactive scene guidance and the data perception results, and evaluates various abilities of the autism index; predicts the results of auxiliary diagnosis of autism based on the trained neural network model; and a data output module, which outputs the early screening results of autism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of early screening of autism, and in particular to an early screening system for autism based on human-computer interaction. Background Art

[0002] The statements in this section merely mention background art related to the present invention and do not necessarily constitute prior art.

[0003] Autism Spectrum Disorder (ASD) is a group of relatively serious neurodevelopmental disorders in children. Patients have difficulty developing normal social relationships and exhibit restricted or repetitive behaviors. Currently, early screening for ASD is primarily based on scale-based testing. There are more than seven mainstream ASD screening scales, each with its own characteristics. The selection of scales varies from hospital to hospital, medical institution to medical practitioner, and so on. Overall, scale-based early screening for ASD faces the following challenges that need to be addressed:

[0004] (1) Most of the norms are formed based on foreign ASD research and lack local adaptability.

[0005] (2) The measurement time is long. Filling out the scale usually takes more than one hour. Some scales, such as the Autism Diagnostic Interview-Revised (ADI-R) test, even take 2-2.5 hours.

[0006] (3) The test results are dependent on long-term observation by guardians or caregivers, and subjective factors may affect the test results.

[0007] (4) Requires practitioners to have higher qualifications.

[0008] (5) Low participation of children or test subjects.

[0009] (6) Screening norms are determined by experimental research rules and are updated slowly.

[0010] In addition to scale-based ASD early screening methods, some researchers have proposed using machine learning solutions for ASD early screening. However, this solution still has some problems that need to be solved:

[0011] (1) Current machine learning usually inputs single-modality data, such as eye movement data and images of subjects, which are insufficiently responsive to relevant information such as body movements, facial expressions, voice, eye contact, social distance, and other key features of ASD, making it difficult to support the model to achieve high reliability and efficiency.

[0012] (2) The lack of interactivity and immersive participation of subjects makes it difficult to complete the necessary high-quality data collection and model iteration in machine learning. Summary of the Invention

[0013] In order to address the deficiencies of the existing technology, the present invention provides an early screening system for autism based on human-computer interaction; the subjects have a high degree of participation, the screening process is interesting, which helps to stimulate the involvement of child subjects and improve screening accuracy and data quality.

[0014] The autism early screening system based on human-computer interaction includes:

[0015] A data acquisition module is configured to: collect the interactive action video and audio of the subject through human-computer interaction;

[0016] The data perception module is configured to: perform limb backbone detection and body language recognition on interactive action video data; perform face detection, facial key point detection, and expression recognition on interactive action video data; perform eye key point detection and gaze direction estimation on interactive action video data in combination with facial key points; perform denoising and quantization processing on audio data; perform iris detection on interactive action video data and perform face distance measurement based on iris size;

[0017] The data analysis module is configured to: count the number of correct responses of the subject based on the interactive scenario instructions and data perception results, and evaluate various abilities of the autism index; use the sequence of body language correctness results, expression correctness results, deviation value sequence of eye direction and target object displayed on the screen, and deviation value sequence of social distance and expected distance generated by the subject during the test as feature data, and predict the results of autism auxiliary diagnosis based on the trained neural network model;

[0018] The data output module is configured to: output the subject's body language ability, expression control ability, eye control ability, volume control ability and social distance ability; and output the early screening results of autism.

[0019] Compared with the prior art, the present invention has the following beneficial effects:

[0020] (1) The initial model is trained and generated with the help of the existing neural network model, and data is collected to update the model to form a local adaptive neural network model.

[0021] (2) Using the screening method of the present invention, the subject can obtain the screening results by interactively performing the test for about 30 minutes.

[0022] (3) The present invention does not need to rely too much on long-term observation and questionnaire filling by guardians or caregivers, but mainly relies on the objective interaction performance of the subjects and the reaction characteristics during the interaction process.

[0023] (4) The present invention does not rely on the experience of practitioners, and the computer completes autonomous detection, autonomous identification, and autonomous decision-making.

[0024] (5) In the method of the present invention, the subject is the main participant, and the screening results are generated by the subject's own interaction data.

[0025] (6) The screening model of the present invention is generated by data drive, and as the data is collected, the screening model can be iterated and updated more quickly.

[0026] (7) Through human-computer interaction technology, multimodal data such as images, audio, voice, coordinate sequences, and interaction results can be obtained, which helps to build an efficient and reliable screening model.

[0027] (8) The subjects have a high level of participation, and the screening process is interesting, which helps to stimulate the involvement of child subjects and improve the screening accuracy and data quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0029] Figure 1 This is a system function module diagram of Example 1;

[0030] Figure 2 This is a diagram of the system hardware architecture of Example 1;

[0031] Figure 3 This is a software architecture diagram of Example 1;

[0032] Figure 4 This is a software flow chart of Example 1;

[0033] Figure 5 This is the first convolutional neural network structure of Example 1;

[0034] Figure 6(a) and Figure 6(b) are the second convolutional neural network structure of Example 1. DETAILED DESCRIPTION

[0035] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.

[0036] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0037] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.

[0038] All data in this embodiment is obtained in compliance with laws and regulations and based on the consent of the user, and is used legally.

[0039] Example 1

[0040] This embodiment provides an early screening system for autism based on human-computer interaction;

[0041] like Figure 1 As shown, the autism early screening system based on human-computer interaction includes:

[0042] A data acquisition module is configured to: collect the interactive action video and audio of the subject through human-computer interaction;

[0043] The data perception module is configured to: perform limb backbone detection and body language recognition on interactive action video data; perform face detection, facial key point detection, and expression recognition on interactive action video data; perform eye key point detection on interactive action video data and estimate the gaze direction based on the facial key points; perform denoising and quantization on audio data; perform iris detection on interactive action video data and perform face distance measurement based on iris size;

[0044] The data analysis module is configured to: count the number of correct responses of the subject based on the interactive scenario guidance and data perception results, and evaluate various abilities of the autism index; use the sequence of body language correctness results, expression correctness results, eye direction deviation value sequence and screen fruit display value sequence, and social distance deviation value sequence between expected distance generated by the subject during the test as feature data, and predict the results of autism auxiliary diagnosis based on the trained neural network model;

[0045] The data output module is configured to: output the subject's body language ability, expression control ability, eye control ability, volume control ability and social distance ability; and output the early screening results of autism.

[0046] Furthermore, the method of collecting the interactive action video and audio of the subject through human-computer interaction specifically includes:

[0047] The subject's limb movement data is collected through human-computer interaction, the specified limb movements of the virtual character are played on the screen, voice is played to guide the subject to imitate, and a video of the subject imitating the virtual character's limb movements is collected through the camera.

[0048] Furthermore, the method of collecting the interactive action video and audio of the subject through human-computer interaction specifically includes:

[0049] The facial expression data of the subjects are collected through human-computer interaction, the set facial expressions of the virtual characters are played on the screen, voice is played to guide the subjects to imitate, and the video of the subjects imitating the facial expressions of the virtual characters is collected through the camera.

[0050] Furthermore, the method of collecting the interactive action video and audio of the subject through human-computer interaction specifically includes:

[0051] The eye movement data of the subjects are collected through human-computer interaction. The subjects are allowed to control the virtual mouse with their eyes, and voice is played to guide the subjects to look at different target objects. In this process, the distance deviation between the position of the mouse controlled by the eyes and the target object on the screen is collected.

[0052] Furthermore, the method of collecting the interactive action video and audio of the subject through human-computer interaction specifically includes:

[0053] The audio data of the subjects are collected through human-computer interaction. The speakers play sounds of different volumes to communicate with the subjects, and the deviation between the subjects' speech volume and the expected volume in different volume scenarios is collected.

[0054] Furthermore, the method of collecting the interactive action video and audio of the subject through human-computer interaction specifically includes:

[0055] The social distance of the subjects is collected through human-computer interaction. People of different identities are played on the screen, and the subjects are guided to maintain a set social distance with people of different identities. The social distance adjustment process of the subjects and the deviation from the set standard social distance are collected.

[0056] Furthermore, the performing of limb backbone detection and body language recognition on the interactive action video data specifically includes:

[0057] (11): Extract limb feature key points and hand feature key points from limb motion data; use the BlazePose model to extract the limb feature key points of the human body, and use the BlazePose model to extract the hand feature key points of the human body;

[0058] (12): Fusing the key points of limb features with the key points of hand features to obtain fused features; the feature fusion is performed in a splicing manner;

[0059] (13): Input the fused features into the first trained convolutional neural network and output the body language prediction label;

[0060] (14): Record the sequence of body language prediction results and the standard body language required to be imitated. When the comparison results are completely consistent, it is recorded as passing the body language imitation test.

[0061] Furthermore, the training process of the trained first convolutional neural network includes:

[0062] Constructing a first training set; the first training set is a video with known body language labels;

[0063] The first training set is input into the first convolutional neural network, and the network is trained to obtain a trained first convolutional neural network.

[0064] It should be understood that the BlazePose model uses a combination of heatmaps, offsets, and regression to stack a small encoder-decoder heatmap-based network and a subsequent regression encoder network. Heatmap and offset losses are only used in the training phase, and the corresponding output layer is deleted from the model before running inference. Skip connections are used between network stages to achieve a balance between high-level and low-level features. The model uses relatively rigid face detection (using the BlazeFace network) instead of human body detection, which allows the network to detect limbs more lightweight. The model effectively uses an encoder-decoder network structure to predict the heatmaps of all relevant nodes, and then uses another encoder to directly regress the coordinates of all joint points, that is, the 33 key points of the detected human body.

[0065] The BlazePlam model, including the palm detector and hand keypoint model, also achieves real-time inference speed and high prediction quality on mobile GPUs.

[0066] Detecting hands is a relatively complex task due to the variety of hand sizes and occlusions. To address these issues, the detector in this model has been improved in the following aspects: Training a palm detector that excludes fingers: Detecting a palm is simpler than detecting the entire hand including fingers. Palms are relatively smaller, so non-maximum suppression (NMS) has fewer side effects. For example, in the case of a handshake, two palms are less likely to be filtered out by NMS than two hands.

[0067] Using an encoder-decoder similar to FPN (Feature Pyramid Networks) can extract multi-scale information. Even for small object detection, a larger scene context is used; Focal Loss is used to deal with the excessive anchor boxes caused by huge scale changes.

[0068] It should be understood that the above steps can generate 33 pose key points, 21 lefthand key points, and 21 righthand key points in 3D coordinates. The 3D coordinates here are the relative coordinate values (x, y, z) of each point in the image, where x and y are the relative positions of the key points in the image, normalized to coordinates between 0 and 1 according to the image width and height, respectively. z is the depth information of the key point in the image, with the depth at the midpoint of the hip as the origin. The smaller the value, the closer the marker is to the camera, and the size of z is roughly the same as the ratio of x. Stored in the form of a row vector according to one image. The pose feature of each image is 33*3=99 dimensions, and the hand feature is 42*3=126 dimensions, that is, the feature points of one image correspond to a 225-dimensional vector. The access data takes the value of the first dimension as a label (i.e., a number between 0 and 6), that is, one image is stored as a 226-dimensional row vector, where the first dimension is the label and the second to 226 dimensions are the feature point coordinates.

[0069] It should be understood that if Figure 5 As shown, the first convolutional neural network includes a depth-separable convolutional layer, a convolutional layer, a fully connected layer and a softmax function layer connected in sequence.

[0070] The depthwise separable convolutional layer is a one-dimensional convolutional neural network with 7 convolution kernels, a convolution kernel size of 5, and a convolution step of 1; the convolution layer is a one-dimensional convolutional neural network with 7 convolution kernels, a convolution kernel size of 5, and a convolution step of 1; the fully connected layer is a fully connected neural network with 7 output nodes, corresponding to the classification results of body language.

[0071] Furthermore, face detection, facial key point detection and expression recognition are performed on the interactive action video data, specifically including:

[0072] (21): Perform face detection from facial expression data and resize the detected face image to 48*48 pixels;

[0073] (22): Input the face image into the trained second convolutional neural network to obtain the expression prediction result;

[0074] (23): Record the sequence of expression prediction results and the standard expression required to be imitated. When the comparison results are completely consistent, it is recorded as passing the expression imitation test.

[0075] Furthermore, the training process of the trained second convolutional neural network includes:

[0076] Constructing a second training set; the second training set is a video with known expression labels;

[0077] The second training set is input into the second convolutional neural network, and the network is trained to obtain a trained second convolutional neural network.

[0078] It should be understood that, as shown in Figures 6(a) and 6(b), the second convolutional neural network includes sequentially connected depthwise separable convolutional layers, the output end of the depthwise separable convolutional layer is connected to the input end of the convolutional layer n1, and the output end of the convolutional layer n1 is respectively connected to the input end of the convolutional layer n2 and the input end of the first residual block;

[0079] The output end of the convolutional layer n2 and the output end of the first residual block are both connected to the input end of the first adder; the output end of the first adder is respectively connected to the input end of the convolutional layer n3 and the input end of the second residual block;

[0080] The output end of the convolutional layer n3 and the output end of the second residual block are both connected to the input end of the second adder; the output end of the second adder is respectively connected to the input end of the convolutional layer n4 and the input end of the third residual block;

[0081] The output end of the convolution layer n4 and the output end of the third residual block are both connected to the input end of the third adder; the output end of the third adder is connected to the softmax function layer through the convolution layer n5 and the average pooling layer.

[0082] The second convolutional neural network consists of 4 modules. The basic module contains two layers of 2D convolutional neural networks. The number of convolution kernels is 8 and the size of convolution kernel is 3*3.

[0083] The first residual block consists of a 4-layer convolutional neural network, and the output feature depth is 16. This feature is added to the skip connection containing a convolutional layer to obtain the output of the new first residual block.

[0084] The structures of the second residual block and the third residual block are similar to those of the first residual block, except that the number of convolution kernels in the last layer of the second residual block is 64, and the number of convolution kernels in the last layer of the third residual block is 128.

[0085] The feature size after three residual modules is 1×3×3×7. The feature is average pooled and the pooled result is converted into the probability of predicting various expressions through the softmax layer.

[0086] Furthermore, the eye key point detection is performed on the interactive action video data, and the gaze direction is estimated in combination with the facial key points; specifically including:

[0087] (31): Use the face detection algorithm Blaze Face to perform facial mesh detection and separate the eye image based on the facial key points;

[0088] (32): Input the eye image into the iris recognition algorithm and output the eye contour estimation result and iris positioning result;

[0089] (33): Based on the eye contour coordinates and iris positioning coordinates, the trained third convolutional neural network is used to predict the coordinate position of the subject's eyes looking at the screen;

[0090] (34): Record the distance difference sequence between the predicted coordinate position and the displayed target object position. When the distance difference is less than 100 pixels, it means that the subject has passed the test of gaze control.

[0091] Furthermore, the training process of the trained third convolutional neural network includes:

[0092] Constructing a third training set; the third training set is the eye contour coordinates and iris positioning coordinates of the known coordinate positions of the eyes looking at the screen;

[0093] The third training set is input into the third convolutional neural network, and the network is trained to obtain a trained third convolutional neural network.

[0094] The BlazeFace model, based on MobileNet, combines a more compact and lightweight feature extraction method with a novel anchor box mechanism optimized for efficient operation on mobile GPUs, and a weighted approach that replaces non-maximum suppression to ensure detection stability. This enables ultra-fast, high-performance face detection on mobile devices. The BlazeFace model improves upon MobileNet by enhancing the computational efficiency and receptive field of MobileNet's depthwise separable convolutions. Based on this, an effective feature extractor is constructed, and the post-processing of the anchor box mechanism is improved. The network outputs the predicted locations of 468 facial landmarks.

[0095] The iris detection model has a two-output branch structure, detecting both eye contour and iris location from a single image. The iris detection algorithm consists of three main modules: the feature encoding module, a MobileNet network structure that extracts features from eye images; the eye contour estimation module, which primarily consists of four layers of 1x1 convolutional kernels and 3x3 convolutional layers; and the iris location estimation module, which also consists of four layers of 1x1 convolutional kernels and 3x3 convolutional layers.

[0096] It should be understood that the neural network model used to predict eye movement information to screen coordinates is similar to the network structure for predicting body language mentioned above. The difference is that the prediction results are two-dimensional numbers, corresponding to the x-coordinate and y-coordinate on the screen respectively.

[0097] Furthermore, the denoising and quantization processing of the audio data specifically includes:

[0098] The voice activity detection algorithm is used to remove noise from the audio data, and then the noise-removed data is quantized.

[0099] Furthermore, the iris detection is performed on the interactive action video data, and the face distance is measured according to the iris size; specifically including:

[0100] (41) Calculate the diameter d′ based on the iris positioning result, where the unit of the diameter is millimeter.

[0101] (42) The focal length information is obtained from the exchangeable image file EXIF of the image, and the focal length is f, in millimeters.

[0102] (43) Based on the diameter and focal length information, calculate the distance s between the face and the camera in centimeters. The specific calculation formula is:

[0103]

[0104] It should be understood that the standard size of the human iris is 1.17 centimeters. The distance between a person and the camera can be estimated by using the standard iris size, focal length information, and the iris size presented on the screen.

[0105] Furthermore, the correct responses of the subjects were counted based on the interactive scenario guidance and data perception results, and the autism index was evaluated for various abilities. Specifically, the following were included:

[0106] Based on the interaction process of the scene, the scores of various abilities are calculated according to the degree of completion of body language interaction, facial expression interaction, eye contact interaction, conversation volume control, and social distance control.

[0107] Furthermore, the correct responses of the subjects were counted based on the interactive scenario guidance and data perception results, and the autism index was evaluated for various abilities. Specifically, the following were included:

[0108] (51) In the body language test, six types of body language without action are preset, and 10 imitation instructions are given. Each completion is scored 10 points. If the subject fails to imitate three types of body language in a row, it is considered a failure and the scoring is stopped.

[0109] (52) In the expression test, 6 expressions without actions were preset, and 10 imitation instructions were given. Each completion was scored 10 points. If the subject failed to imitate 3 expressions in a row, it was considered a failure and the scoring was stopped.

[0110] (53) In the social distance control test, three types of social task roles are preset. The subjects complete the task according to the distance instructions to maintain an appropriate social distance with two types of roles. The average difference between the social distance and the standard distance is calculated, and the deviation is normalized from large to small to 0 to 100 points;

[0111] (54) In the speech volume test, the subjects were guided to have a conversation by setting up a quiet indoor environment and a noisy outdoor environment. The standard volume in the current environment was prompted. The average value of the difference between the speech volume and the standard volume was calculated, and the deviation was normalized from large to small to a score of 0 to 100.

[0112] (55) In the eye control test, various types of fruits will appear on the screen 10 times, and the distance between the mouse position controlled by the line of sight and the target object position will be counted. If the deviation is less than 100 pixels, the test is passed, and the score is calculated according to the number of times it is passed.

[0113] Furthermore, the test uses the test subject's body language correctness result sequence, facial expression correctness result sequence, eye gaze direction and screen target object deviation value sequence, and social distance and expected distance deviation value sequence as feature data to predict the autism auxiliary diagnosis result based on the trained neural network model; wherein the training process of the trained neural network model includes:

[0114] Constructing a fourth training set; the fourth training set includes: feature data of known autism diagnostic labels;

[0115] The fourth training set is input into the neural network model, and the model is trained to obtain a trained neural network model.

[0116] An early screening neural network model was constructed by using body language prediction sequences, expression prediction sequences, distance sequences between the eye-controlled mouse and the target object on the screen, deviation sequences between speech volume and expected volume, and deviation sequences between social distance and expected distance, with confirmed children and normal children as classification labels.

[0117] The embodiment of the present application also needs to determine the characteristic indicators of ASD core symptoms: based on DSM-5 and the recommendations of ASD screening and assessment practitioners, five characteristic indicators reflecting the core symptoms of ASD are established, including body language, facial expressions, social distance control, speech volume control, and eye control.

[0118] Edge computing solution and hardware and software system architecture: An edge computing solution was determined, employing an edge-terminal hardware architecture and a client-server software architecture to enable efficient operation of computer vision algorithms on edge servers. The terminal uses the Android operating system, primarily running interactive terminal applications responsible for stimulating interaction and multimodal data collection. The edge uses the Linux operating system, primarily running computer vision algorithms and multimodal data processing algorithm services, responsible for processing multimodal data collected by the terminal and providing feedback on early ASD screening results. Furthermore, the edge software design implements on-demand loading of computer vision models to avoid memory usage issues caused by loading multiple perception models.

[0119] Interactive scenario design: Design five scenarios based on the five core ASD symptom indicators to guide subjects to participate in the interaction.

[0120] Design ideas for interactive body language scenarios: Design a virtual character capable of seven types of body language: "Keep Quiet," "Greet," "Admiration," "Helpless," "Reject," "Think," and "No Action." Use voice guidance to guide children to imitate, while a camera captures the child's portrait for real-time display. An algorithm evaluates the imitation results, which are then displayed in real time on the webpage.

[0121] Design ideas for interactive facial recognition scenarios: Design a virtual character capable of expressing seven different expressions: "disgust," "fear," "anger," "happiness," "sadness," "surprise," and "no expression." Use voice guidance to guide children to imitate, while a camera captures the child's portrait for real-time display. An algorithm evaluates the imitation results, which are then displayed in real time on the webpage.

[0122] Design ideas for the eye contact interaction scenario: Design a virtual character to guide the subject to control the "virtual mouse" by looking at the designated object. Once the subject is able to use the "virtual mouse" with their eyes, start using voice prompts to guide the subject to look at different types of fruit. When the subject controls the eye contact correctly, they will receive positive feedback indicating that the instruction has been completed. During this process, the child's eye movement information will be collected;

[0123] Speech volume control scenario design ideas: Design two scenarios: "Indoor Conversation" and "Outdoor Conversation". The virtual avatar guides the subject to speak at an appropriate volume. The conversation volume is displayed on the interactive interface. Consider the subject's ability to control the speech volume in different scenarios and provide certain feedback guidance.

[0124] Design ideas for social distance interaction test scenarios: Design virtual human images with different identities, including "parents", "partners", and "strangers", and guide the subjects to maintain a certain social distance with people of different identities. By collecting the subjects' social distance adjustment process and social distance maintenance process, relevant data reflecting social distance control ability are collected.

[0125] Vision-based body movement perception and body language recognition algorithm:

[0126] The limb motion perception algorithm can obtain the topological map of the human backbone, which helps to extract the characteristics of human motion. This paper refers to the BlazePose network structure and pre-trained model to achieve this task.

[0127] First, static topology design is performed: a static limb backbone topology diagram is constructed based on the basic human joints that can express limb movement. The 3D coordinates of the intersection points of the topology diagram are represented by (x, y, z) in the Euclidean coordinate system. The key points of the human backbone are planned for reference, such as Figure 2 shown.

[0128] Then, human target detection is performed: a deep learning algorithm solution for target detection is used to establish a human detection algorithm model. The bounding box of the selected human body is obtained in the inference stage, and a static human backbone topology map is preset in the bounding box to ensure the integrity of the key points of the human backbone.

[0129] Finally, human key point detection is performed: a deep learning regression algorithm model is constructed, and the input data is the image data that has been cropped with a bounding box after human target detection; the output data (prediction data) is the predicted value of the topological structure point in the three-dimensional coordinates and the visibility confidence of the topological structure point. After obtaining the regression key point prediction, the visible points are matched to the preset static topology map, and the invisible key points are updated according to the connection relationship of the topology map.

[0130] The body language recognition algorithm can identify the body language expressions of people in RGB images containing human gestures. This algorithm includes seven types of body language: "Keep Quiet," "Greet," "Admiration," "Helpless," "Rejection," "Thinking," and "No Action." This algorithm is implemented by establishing a neural network classification task. First, image data for each of the seven types of body language is collected, with the corresponding expressions representing index values (1 to 7) for each of the seven types of body language. The image data is then processed using the aforementioned body movement perception algorithm to obtain body movement coordinate data. A neural network model is then designed, with the body movement coordinate data as input and the body language index values as output labels. The cross-entropy loss function is used as the loss function. This data is then fed into the neural network for model training, generating an algorithmic model for body language recognition.

[0131] Vision-based facial feature perception and expression recognition algorithm: The 3D feature information of the face can well reflect the deflection of the face and the changes in facial expressions, which is of great significance for discovering the subtle features of the face. Therefore, the present invention uses the BlazeFace network structure and pre-trained weights to realize facial feature recognition.

[0132] First, a static topology structure is designed: based on the facial features of the face, a static face mesh topology map is established, and the coordinates of the mesh intersection points are represented by (x, y, z) in the Euclidean coordinate system.

[0133] Then face detection and key point matching are performed: single image data is obtained from the time-series video, face detection is performed on the single data, and the face bounding box and estimated coordinates of the face key points are generated.

[0134] Facial key points can be used to match the key points in the facial 3D structure in a rotated state, and the coordinates of the mesh structure are estimated based on the deformation relationship of the key points to achieve the estimation of the entire face 3D information.

[0135] The facial expression recognition algorithm can identify the facial expressions of people in RGB images containing human gestures. This algorithm covers seven categories of expressions: disgust, fear, anger, happiness, sadness, surprise, and neutral expression. This algorithm is implemented by establishing a neural network classification task. First, image data for each of the seven categories of expressions is collected, with the corresponding expressions representing index values (1 to 7). The image data is then processed using the facial feature perception algorithm to obtain coordinate data for body movements. A neural network model is then designed, with the input data consisting of the coordinates of key facial points and the output labels being the expression index values. The loss function uses the cross-entropy loss function. This data is then fed into the neural network for model training, generating an algorithmic model for facial expression recognition.

[0136] Eye movement feature capture based on iris detection:

[0137] Eye movement feature extraction is mainly achieved through the iris detection algorithm. The first step of this algorithm relies on 3D facial mesh detection in BlazeFace. This algorithm applies high-assurance facial landmarks to generate a rough mesh of the facial structure. From these meshes, we can separate the eye image from the original image and apply the image to the iris tracking model. The problem is then divided into two parts: eye contour estimation and iris positioning. A multi-task model is designed that has a unified encoder for each task, and each task has a separate component. This design allows us to use task-specific training data. The image data is annotated with eyelid information and iris contour information. These two types of information are labels for the two learning tasks. The specific recognition process is as follows: Figure 3As shown in Figure 1, the input of the model is the cropped image of the human eye area. In the prediction result, the human eye contour is generated by a separate task 1 decoder, and the iris contour is generated by a task 2 decoder.

[0138] Speech and human voice detection and volume quantification: When interacting with speech volume scenarios, it is necessary to quantify the subject's volume in real time. However, in actual implementation, the volume of ambient noise often exceeds the volume of human voice, resulting in incorrect volume detection. The present invention introduces a voice activity detection algorithm (VAD) to first remove noise and then quantify the sound volume.

[0139] Face ranging algorithm based on monocular camera iris detection: This invention uses a face ranging model to determine the distance between the human eye and the camera with an error of no more than 10% without any additional equipment. This relies on the consistency of the human iris' horizontal size. Studies have shown that the human eye size remains constant at 11.7 ± 0.5 mm and has a simple geometric structure.

[0140] A pinhole camera projects an image onto a square sensor. The distance between the eye and the camera can be estimated using the camera's focal length, which can be obtained from the image's EXIF metadata. Other camera parameters are intrinsic. Given the focal length, the distance between the eye and the camera can be directly calculated by comparing the actual eye size with the pixel size of the imaged eye.

[0141] Through human-computer interaction data collection, a data set with local characteristics can be constructed, which is helpful in generating local ASD screening norms.

[0142] Using the screening method of the present invention, the subject can obtain the screening results by interactively performing the test for about 30 minutes.

[0143] The present invention does not need to overly rely on long-term observation and questionnaire filling by guardians or caregivers, and mainly relies on the objective interaction performance of the subjects and the reaction characteristics during the interaction process.

[0144] The present invention does not rely on the experience of practitioners, and the computer completes independent detection, independent identification, and independent decision-making.

[0145] In the method of the present invention, the subject is the main participant, and the screening result is generated by the subject's own interaction data.

[0146] The screening model of the present invention is generated by data drive, and as data is collected, the screening model can be updated more quickly through iteration.

[0147] The present invention takes into account multimodal data such as images, audio, voice, coordinate sequences, and interaction results, which helps to build a high-efficiency and high-reliability screening model.

[0148] The present invention can not only obtain screening results, but also refine them into quantitative evaluation indicators such as body language, facial expressions, social distance control, speech volume, and eye control.

[0149] The subjects had a high level of participation, and the screening process was interesting, which helped to stimulate the involvement of the child subjects.

[0150] like Figure 2 As shown in the figure, the system hardware consists of three parts: peripheral devices, terminal devices, and edge devices. The peripheral devices include a monocular camera, microphone, speaker, and touch screen; the terminal device includes a Rockchip RK3399 CPU, 6GB of memory, 128G of storage space, an Android motherboard with a network card, and other devices, and its operating system is the Android operating system; the edge device mainly includes an Intel i7 CPU, 16GB of memory, a 1T hard drive, a GTX2080S graphics card, a PC motherboard with a network card, and other devices, running the Linux operating system. The communication between the edge and the terminal adopts a LAN routing solution.

[0151] like Figure 3 As shown in the figure, the system software consists of an interactive scene layer, a data acquisition layer, a data storage layer, an intelligent perception layer, a data analysis layer, and a user feedback layer.

[0152] Among them, the interactive scene layer includes body language interaction scenes, facial expression interaction scenes, eye contact interaction scenes, dialogue scenes, and social distance interaction scenes. These scenes are presented to the children in the form of cartoon images. Children use cartoon interactive scenes to click buttons, imitate facial expressions, imitate body language, control the virtual mouse with their eyes, have conversations, and actively adjust social distance to complete the interactive tasks designed in the scenes.

[0153] The data acquisition layer includes data in multiple modalities such as images, videos, audio, and interactive feedback. These data can be refined into video data containing limb movement information, image data containing expression information, facial video data containing eye movement information, audio data containing voice volume information, and image data containing dry iris pixel information and camera focal length information.

[0154] The data storage layer primarily implements raw data storage and structured interactive data storage. Audio, video, and image data acquired by acquisition devices are stored in MinIO object storage, providing data support for subsequent model iterations. 3D limb coordinate data, face 3D coordinate data, eye coordinate data, expression data, volume data, and distance data calculated by the perception algorithm are stored in a MySQL structured database.

[0155] Intelligent perception layer, including general perception and scenario-based perception. General perception algorithms include face detection, limb backbone detection, iris detection, face 3D mesh detection, face tracking, and iris tracking.

[0156] A general perception algorithm can estimate the coordinate state of various parts of the human body in a 3D scene. When there is a need to protect the privacy of the subject, the above object storage method will be stopped, and the general perception algorithm will be used to obtain and store 3D information;

[0157] Scenario-based perception algorithms are mainly used for algorithm support in interactive scenarios, including body language recognition, expression recognition, gaze direction estimation, face distance estimation, voice denoising, audio quantization, etc., to achieve automatic evaluation and automatic response of interactive scene interaction results.

[0158] Data analysis layer, including: ASD comprehensive assessment and feature data modeling and analysis;

[0159] ASD comprehensive assessment refers to calculating the scores of various abilities according to certain rules based on the interaction process of the scene, the degree of completion of body language interaction, facial expression interaction, eye contact interaction, conversation volume control, and social distance control.

[0160] The feature data modeling and analysis method uses collected time-series data of limb coordinates, 3D facial coordinates, and eye coordinates as input, and uses confirmed and normal children as classification labels to construct an early screening neural network model. This model's predictions complement the comprehensive ASD assessment model, primarily addressing biased results caused by participants' inattention during assessments.

[0161] The user feedback layer includes ASD early screening results and body language ability, expression control ability, eye control ability, speech volume control ability, and social distance control ability based on interactive statistical analysis.

[0162] In addition to the various software layers, the system's operational health module runs through the entire process. Meanwhile, the data storage layer, intelligent perception layer, and data analysis layer have logging functions, primarily used for recording operational logs and debugging development.

[0163] Since the intelligent perception layer and data analysis layer involve a large number of perception models and data analysis models, the amount of calculation is large. At the same time, the terminal / edge communication based on the C / S architecture has a certain dependence on access bandwidth. Therefore, a load balancing module is added to the intelligent perception layer and data analysis layer to ensure the stability of computing power distribution and data access.

[0164] like Figure 4 As shown, the specific data processing flow is as follows:

[0165] The subjects complete the interactive tasks in front of the terminal according to the prompts. During the interaction process, interactive data and audio and video data collected by peripheral devices are generated. These data are collectively referred to as process data. The process data is input into the perception algorithm, which can calculate the statistical data of the degree of interaction completed by the subjects, and at the same time generate feature data from images, audio data, and video data. The evaluation model based on interactive input rule judgment produces evaluation results covering the key symptoms of ASD, including limb movement ability, expression ability, eye control ability, speech volume control ability, and distance maintenance ability. For feature data, it is imported as input data into the trained screening classification model to obtain screening results.

[0166] Based on the present invention, those skilled in the art can use other forms of interaction to obtain subject data, such as using VR equipment to build a virtual interactive scene for data collection and early screening of ASD.

[0167] Based on the present invention, those skilled in the art can use other types of edge computing servers for data processing and interactive support, or use high-computing-power terminals to migrate the edge algorithms involved in the present invention to the terminal for operation.

[0168] Based on the present invention, those skilled in the art can replace the interactive scene content and design an ASD early screening system with the same concept.

[0169] Those skilled in the art can replace the perception algorithm model structure or weights based on the present invention to achieve the same or similar visual perception results.

[0170] The current basis for early screening is as follows:

[0171] (1) Children with ASD generally lack the ability to imitate. Body language and expression testing scenarios mainly stimulate the subjects to imitate and can reflect the subjects' imitation ability.

[0172] (2) Children with ASD lack the ability to proactively respond to instructions.

[0173] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. The early screening system for autism based on human-computer interaction is characterized by: include: A data acquisition module is configured to: collect the interactive action video and audio of the subject through human-computer interaction; The method of collecting the subject's interactive action video and audio through human-computer interaction includes: collecting the subject's social distance through human-computer interaction, displaying people of different identities on the screen, guiding the subject to maintain a set social distance with people of different identities, and collecting the subject's social distance adjustment process and deviation from the set standard social distance; The data perception module is configured to: perform limb backbone detection and body language recognition on interactive action video data; perform face detection, facial key point detection, and expression recognition on interactive action video data; perform eye key point detection on interactive action video data and estimate gaze direction based on facial key points; perform denoising and quantization on audio data; perform iris detection on interactive action video data and perform face distance measurement based on iris size; The limb backbone detection and body language recognition of the interactive action video data specifically includes: extracting limb feature key points and hand feature key points from the limb movement data; using the Blaze Pose model to extract the limb feature key points of the human body, and using the Blaze Plam to extract the hand feature key points of the human body; fusing the limb feature key points with the hand feature key points to obtain a fused feature; the feature fusion is performed in a splicing manner; the fused feature is input into the trained first convolutional neural network to output a body language prediction label; and a sequence of the body language prediction results and the standard body language required to be imitated are recorded. When the comparison results are completely consistent, it is recorded as passing the body language imitation test. The method of detecting eye key points on interactive action video data and estimating gaze direction in combination with facial key points specifically includes: The Blaze Face face detection algorithm is used to perform facial mesh detection and isolate the eye image based on facial key points. The eye image is input into the iris recognition algorithm, which outputs eye contour estimation and iris location results. Based on the eye contour coordinates and iris location coordinates, a trained third convolutional neural network is used to predict the coordinate position of the subject's eyes looking at the screen. The distance difference sequence between the predicted coordinate position and the displayed target object position is recorded. When the distance difference is less than the set pixel, it indicates that the subject has passed the gaze control test. The method includes performing iris detection on the interactive action video data and measuring the distance of the face based on the iris size. The method includes calculating the diameter of the iris positioning result; obtaining the focal length information from the exchangeable image file EXIF of the image; and calculating the distance between the face and the camera based on the diameter and focal length information. The data analysis module is configured to: count the number of correct responses of the subject based on the interactive scenario instructions and data perception results, and evaluate various abilities of the autism index; use the sequence of body language correctness results, expression correctness results, deviation value sequence of eye direction and target object displayed on the screen, and deviation value sequence of social distance and expected distance generated by the subject during the test as feature data, and predict the results of autism auxiliary diagnosis based on the trained neural network model; The data output module is configured to: output the subject's body language ability, expression control ability, eye control ability, volume control ability and social distance ability; and output the early screening results of autism.

2. The autism early screening system based on human-computer interaction as claimed in claim 1, characterized in that: Perform face detection, facial key point detection, and expression recognition on interactive action video data, including: Perform face detection on facial expression data and resize the detected face image to 48*48 pixels. Input the facial image into the trained second convolutional neural network to obtain the expression prediction result; Record the sequence of expression prediction results and the standard expression required to be imitated. When the comparison results are completely consistent, it is recorded as passing the expression imitation test.

3. The autism early screening system based on human-computer interaction as claimed in claim 1, characterized in that: According to the interactive scenario guidance and data perception results, the correct response times of the subjects are counted, and the autism index is evaluated for various abilities; Specifically include: Based on the interaction process of the scene, the scores of various abilities are calculated according to the degree of completion of body language interaction, facial expression interaction, eye contact interaction, conversation volume control, and social distance control.

4. The autism early screening system based on human-computer interaction as claimed in claim 1, characterized in that: The method of collecting the interactive action video and audio of the subject through human-computer interaction specifically includes: The subject's limb movement data is collected through human-computer interaction, the specified limb movements of the virtual character are played on the screen, voice is played to guide the subject to imitate, and a video of the subject imitating the virtual character's limb movements is collected through the camera.

5. The autism early screening system based on human-computer interaction as claimed in claim 1, characterized in that: The method of collecting the interactive action video and audio of the subject through human-computer interaction specifically includes: The facial expression data of the subjects are collected through human-computer interaction, the set facial expressions of the virtual characters are played on the screen, voice is played to guide the subjects to imitate, and the video of the subjects imitating the facial expressions of the virtual characters is collected through the camera.

6. The autism early screening system based on human-computer interaction as claimed in claim 1, characterized in that: The method of collecting the interactive action video and audio of the subject through human-computer interaction specifically includes: The eye movement data of the subjects are collected through human-computer interaction. The subjects are asked to control the virtual mouse with their eyes, and voice is played to guide the subjects to look at different target objects. In this process, the distance deviation between the position of the mouse controlled by the eyes and the target object on the screen is collected.

7. The autism early screening system based on human-computer interaction as claimed in claim 1, characterized in that: The method of collecting the interactive action video and audio of the subject through human-computer interaction specifically includes: The audio data of the subjects are collected through human-computer interaction. The speakers play sounds of different volumes to communicate with the subjects, and the deviation between the subjects' speaking volume and the expected volume in different volume scenarios is collected.

Citation Information

Patent Citations

  • Autism early-period evaluation device based on indicative language paradigm, and system

    CN110364260A

  • Autism rehabilitation training and ability evaluation system and method based on virtual reality

    CN110890140A