Construction method of autism spectrum disorder recognition model based on deep learning
By monitoring and analyzing the gaze video rate and social language data of autistic patients in real time in the autistic spectrum disorder identification model, the problem of difficulty in accurately judging the patient's gaze reduction status in the prior art is solved, and efficient and accurate identification of autistic spectrum disorder is achieved.
Patent Information
- Application Number
- CN202510472860.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-16
AI Technical Summary
When analyzing the patient's eye gaze, the existing autism spectrum disorder recognition model lacks dynamic monitoring of the injecting video rate and real-time comparison with the threshold, making it difficult to accurately judge the state of decreased gaze and cannot accurately record the gaze duration.
By obtaining conversation videos of autistic patients and continuously monitoring the gaze video rate, setting the gaze video rate threshold. When the gaze video rate is detected to be lower than the threshold, it is determined that the gaze gaze decrease state and the gaze duration is recorded. At the same time, social language analysis was performed on the video, facial response characteristics and limb posture characteristics were identified, and recognition models were constructed based on eye gaze data.
Accurate quantification and dynamic monitoring of gaze behavior in autistic patients is achieved, the accuracy of assessment of social communication disorders is improved, and the efficiency and accuracy of identifying autistic spectrum disorders is enhanced.
Smart Images

Figure CN119993467A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedical information technology, and in particular to a method for constructing an autism spectrum disorder recognition model based on deep learning. Background Art
[0002] One of the core symptoms of autism spectrum disorder is social communication disorder, including difficulties in social interaction and limited verbal and non-verbal communication abilities. Existing methods for building autism spectrum disorder identification models often involve multiple types of data, including behavioral data, brain imaging data, psychological and emotional states, external behavioral patterns, achievement development trajectories and other multi-dimensional information. For example, some studies use data on children's head, trunk and foot movements, as well as head movement characteristics for judgment; others combine resting-state functional magnetic resonance imaging (rsfMRI) data with machine learning methods. However, when analyzing the gaze of autistic patients, traditional methods often simply count the number of gazes or the duration of gazes, lack dynamic monitoring of gaze frequency and real-time comparison with the threshold, making it difficult to accurately judge the state of reduced gaze, and unable to accurately record the duration of gaze. Summary of the invention
[0003] Based on this, it is necessary to provide a method for constructing an autism spectrum disorder identification model based on deep learning to solve at least one of the above technical problems.
[0004] To achieve the above purpose, a method for constructing an autism spectrum disorder recognition model based on deep learning is provided, the method comprising the following steps: Step S1: obtaining a conversation video of an autistic patient; continuously measuring the eye gaze frequency in the conversation video of the autistic patient; when detecting that the eye gaze frequency is lower than an eye gaze frequency threshold, determining that the state is an eye gaze reduction state and recording the eye gaze duration; Step S2: Performing social language analysis on the conversation video of the autistic patient to obtain social language data; identifying the facial reaction features of the patient in the conversation video of the autistic patient based on the social language data and recording the facial muscle reaction data; identifying the body posture features in the conversation video of the autistic patient based on the social language data and recording the body posture method data; Step S3: evaluating communication disorder based on the gaze duration and facial muscle reaction data to obtain a communication disorder pattern; evaluating social interaction disorder based on the gaze duration and body posture data to obtain a social interaction disorder pattern; Step S4: Obtain an autism spectrum disorder history library; perform feature matching on the social interaction disorder pattern and the communication disorder pattern with the autism spectrum disorder history library respectively to generate interaction disorder matching data and communication disorder matching data; and construct an autism spectrum disorder recognition model based on the interaction disorder matching data and the communication disorder matching data.
[0005] The present invention can convert the patient's gaze behavior into specific quantifiable data, namely, gaze frequency and gaze duration, by acquiring autistic patients' conversation videos and continuously monitoring the frequency of gazes therein. This process provides an objective, accurate and operable data basis for subsequent evaluation, avoiding the deviation caused by subjective judgment; setting a gaze frequency threshold, when it is detected that the gaze frequency is lower than the threshold, it can timely and accurately determine that the patient is in a state of reduced gaze, and record the duration of gaze. This helps to capture the abnormal changes in the gaze behavior of autistic patients in social interactions, and provides a key basis for further analyzing their communication disorders. Social language analysis is performed on the conversation videos of autistic patients to obtain social language data, and at the same time, the facial reaction features of the patient are identified based on the data and recorded as facial muscle reaction data, and the body posture features are identified and recorded as body posture mode data. This process realizes the comprehensive collection of multi-dimensional data from language communication, facial expressions to body movements, providing rich and complete information for subsequent comprehensive evaluation; social language data can reflect the patient's language expression and communication ability, facial muscle reaction data can reflect the patient's emotional response to social situations, and body posture data reveals the patient's body language performance in social interaction. The comprehensive analysis of these three types of data can comprehensively and objectively reflect the comprehensive ability of autistic patients in the social process, and provide strong support for accurately identifying and evaluating their degree of social impairment. Communication impairment assessment is based on gaze duration and facial muscle reaction data to obtain communication impairment patterns; social interaction impairment assessment is based on gaze duration and body posture data to obtain social interaction impairment patterns. This comprehensive evaluation method based on multi-dimensional data can accurately identify the specific impairment patterns of autistic patients in communication and social interaction, and provide clear directions and basis for subsequent personalized intervention; by integrating multiple data for evaluation, the limitations of single indicator evaluation are avoided. For example, evaluation based only on language communication ability may ignore the patient's impairment in non-verbal social interaction, while combining data such as facial expressions and body movements can more comprehensively reflect the patient's social ability and improve the accuracy and reliability of the evaluation. Obtain an autism spectrum disorder history library, and perform feature matching of social interaction disorder patterns and communication disorder patterns with the history library, respectively, to generate interaction disorder matching data and communication disorder matching data, and construct an autism spectrum disorder recognition model based on the interaction disorder matching data and the communication disorder matching data. Therefore, the present invention uses data processing technology and deep learning technology to evaluate the communication disorder and social interaction disorder of autistic patients, and construct an autism spectrum disorder recognition model, thereby improving the recognition efficiency and accuracy of autism spectrum disorders.
[0006] Preferably, the step S1 of obtaining the conversation video of the autistic patient includes: In multiple preset scenarios, the conversation videos of autistic patients were recorded. During the recording process, the patient was at the center of the camera, and the horizontal angle between the camera and the patient's eyes was kept within 10°. Perform a preliminary screening of the recorded videos to remove the clips with blurry images for more than 5 seconds due to camera shaking, while retaining the clips in which the patient's facial expressions and body movements are shown; The retained video clips were cropped, and the video frames of the first 3 seconds and the last 2 seconds before the patient entered the conversation state were cropped; The cropped video segments are spliced together according to the timestamps of the patient's conversation. The total duration of the spliced video should be no less than 1 minute to obtain a conversation video of an autistic patient.
[0007] In the present invention, the patient is at the center of the lens and the angle between the lens and the patient's eye level is maintained within 10°, which can ensure that the acquisition angle of video data in different scenes is consistent, reduce the interference caused by the difference in shooting angles on subsequent analysis, and enable the model to more accurately capture the patient's facial expressions, body movements and other key information; remove the fragments with blurred images for more than 5 seconds due to lens shaking, which can effectively avoid the blurred images interfering with the model's recognition of the patient's behavioral characteristics, and ensure that the clarity of the retained video fragments meets the analysis requirements, thereby improving the data quality; retain the fragments in which the patient's expression changes and body movements appear, and can accurately focus on the patient's key behavioral performance during the conversation. These fragments often contain important clues for the identification of autism spectrum disorders; cut off the video frames of the first 3 seconds and the last 2 seconds when the patient does not enter the conversation state, and can remove redundant information unrelated to the conversation, so that the video fragments are more focused on the patient's real behavioral performance during the conversation, and further improve the pertinence and effectiveness of the data; splicing according to the timestamp of the patient's conversation can integrate the effective conversation fragments in different scenes into a complete video, and ensure that the total duration of the spliced video is not less than 1 minute, providing a continuous conversation sample of sufficient duration for the deep learning model.
[0008] Preferably, the eye gaze frequency in the continuous autistic patient conversation video in step S1 includes: Analyze the conversation videos of autistic patients frame by frame to determine the patient's eye direction in each frame; In each frame, determine whether the patient's eye direction forms an angle less than 30° with the camera direction. If the eye direction forms an angle less than 30° with the camera direction, record the frame as a valid gaze. Take every 10 seconds as a statistical unit, count the number of frames with effective gaze in each statistical unit, and get the gaze frequency of each statistical unit; Divide the entire video into the above statistical units, calculate the gaze frequency of each statistical unit, and record it as the frequency value of the statistical unit; The frequency values of all statistical units are summarized and the average eye gaze frequency of the entire video is calculated as the eye gaze frequency in the conversation video of the autistic patient.
[0009] The present invention can accurately obtain the patient's gaze change information during the conversation by analyzing the patient's eye direction frame by frame, providing an accurate basis for subsequent judgment of the gaze situation; taking the angle between the eye direction and the lens direction less than 30° as the judgment standard of effective gaze, the gaze behavior is quantified into a countable number of frames, providing a specific and operational quantitative indicator for subsequent analysis; taking every 10 seconds as a statistical unit, the number of frames of effective gaze in each statistical unit is counted to obtain the gaze frequency of each statistical unit, which can more carefully observe the patient's gaze situation in different time periods and avoid the overall statistics from covering up local features; summarizing the frequency values of all statistical units, and calculating the average gaze frequency of the entire video, providing a comprehensive and quantified target feature for the autism spectrum disorder recognition model.
[0010] Preferably, the communication disorder assessment based on gaze duration and facial muscle reaction data in step S3 includes: Determine the start time and end time of the gaze duration; determine the start time and end time of the facial muscle reaction data; The duration of gaze fixation was compared with the start and end time of the facial muscle reaction data; if the difference between the start time of gaze fixation and the start time of facial muscle reaction was less than 1 second, it was marked as a synchronous reaction; if the difference between the start time of gaze fixation and the start time of facial muscle reaction was more than 1 second, it was marked as an asynchronous reaction; The ratio of gaze duration to facial muscle reaction duration was calculated; if the ratio was less than 0.5, it was considered a low ratio; if the ratio was between 0.5 and 1.5, it was considered a medium ratio; if the ratio was greater than 1.5, it was considered a high ratio; Communication disorders are assessed based on the synchronization of eye gaze and facial muscle response and the ratio range; if the response is synchronous and the ratio is low, it is marked as a slow communication mode; if the response is synchronous and the ratio is high, it is marked as a rapid communication mode; if the response is asynchronous and the ratio is low, it is marked as an incoordinated communication mode; if the response is asynchronous and the ratio is high, it is marked as a disconnected communication mode; The above communication disorder assessment results are summarized into a communication disorder model.
[0011] The present invention can accurately quantify the time characteristics of eye gaze and facial muscle reaction by determining the start time and end time of eye gaze and the start time and end time of facial muscle reaction data; the start time of eye gaze is compared with the start time of facial muscle reaction, and if the time difference does not exceed 1 second, it is marked as a synchronous reaction, and if it exceeds 1 second, it is marked as an asynchronous reaction. This quantitative method can clearly distinguish the patient's response patterns in social interactions and provide a clear classification standard for identifying communication disorders. By calculating the ratio of gaze duration to facial muscle response duration and dividing it into low ratio (less than 0.5), medium ratio (0.5 to 1.5) and high ratio (greater than 1.5), the patient's communication behavior characteristics can be further refined. According to the synchronization of gaze and facial muscle response and the ratio range, the communication disorder assessment results are divided into communication delay mode (synchronous response and low ratio), communication rush mode (synchronous response and high ratio), communication incoordination mode (asynchronous response and low ratio) and communication interruption mode (asynchronous response and high ratio). The above communication disorder assessment results are summarized as a communication disorder mode, which provides a comprehensive and quantitative target feature for the autism spectrum disorder identification model.
[0012] Preferably, the social interaction disorder assessment based on gaze duration and body posture data in step S3 includes: Determine the start time and end time of the gaze duration; determine the start time and end time of the facial muscle reaction data; set the gaze duration of more than 3 seconds as long gaze, and the gaze duration of less than 1 second as short gaze; Extract the angle between the trunk and the vertical direction and the extension angle of the knee joint from the limb posture data; set the change of the vertical angle of the trunk to 15-30 degrees as a small change of the trunk angle, and the change of the vertical angle of the trunk to more than 30 degrees as a large change of the trunk angle; The duration of gaze fixation was compared with the changes in the trunk angle and knee extension angle. If the duration of gaze fixation was long and the trunk angle had a small change, it was marked as a coordinated social interaction mode; if the duration of gaze fixation was short and the trunk angle had a large change, it was marked as a discontinuous social interaction mode. By integrating the social interaction coordination model and the social interaction disruption model, we obtain the social interaction disorder model.
[0013] The present invention can accurately calculate the duration of gaze by determining the start time and end time of gaze, and provide specific and operable quantitative indicators for subsequent analysis; determine the start time and end time of facial muscle reaction data, and capture the patient's social interaction behavior characteristics from multiple dimensions; set the gaze duration of more than 3 seconds as long gaze, and less than 1 second as short gaze, and classify the gaze behavior to provide clear classification standards for subsequent analysis; extract the angle between the trunk and the vertical direction and the extension angle of the knee joint in the limb posture data, and obtain the patient's limb behavior characteristics in social interaction; set the change of the vertical direction angle of the trunk to 15-30 degrees as the trunk angle Small changes are considered as trunk angle changes, and more than 30 degrees is considered as large changes in the trunk angle, and the changes in limb movements are quantified and classified; the duration of gaze is compared with the changes in the trunk angle and knee extension angle to correlate the relationship between gaze and limb movements; if the duration of gaze is long and the trunk angle changes slightly, it is marked as a social interaction coordination mode; if the duration of gaze is short and the trunk angle changes significantly, it is marked as a social interaction disruption mode, which can identify the different modes of patients in social interaction; the social interaction coordination mode and the social interaction disruption mode are integrated to obtain the social interaction disorder mode, which provides a comprehensive and quantitative target feature for the autism spectrum disorder identification model.
[0014] Preferably, the step S4 of matching the social interaction disorder pattern and the communication disorder pattern with the autism spectrum disorder history library includes: Extract the pattern occurrence time, pattern duration and pattern frequency of the social interaction disorder pattern and record them as the social interaction disorder characteristic data; Extract the pattern occurrence time, pattern duration and pattern frequency of the communication disorder pattern and record them as communication disorder characteristic data; The social interaction disorder data and communication disorder feature data are matched with the autism spectrum disorder history library respectively, and the similarity between each pattern and the known patterns in the autism spectrum disorder history library is determined and quantified. If the quantified similarity value exceeds 0.8, the match is considered successful. Interaction barrier matching data and communication barrier matching data are recorded based on similarity quantification values.
[0015] The present invention extracts the pattern occurrence time, pattern duration and pattern frequency of the social interaction disorder pattern, and records them as social interaction disorder feature data. This process can accurately quantify the characteristics of social interaction disorder and provide accurate input for subsequent feature matching; extract the pattern occurrence time, pattern duration and pattern frequency of the communication disorder pattern, and record them as communication disorder feature data. This process can accurately quantify the characteristics of communication disorders and further enrich the input data of the model; feature match the social interaction disorder feature data and the communication disorder feature data with the autism spectrum disorder history library respectively, determine the similarity of each pattern with the known patterns in the history library, and quantify the similarity. If the quantized similarity value exceeds 0.8, it is considered that the match is successful, and the similarity between the feature and the known pattern can be effectively evaluated to improve the accuracy of recognition; based on the quantized similarity value, the interaction disorder matching data and the communication disorder matching data are recorded. This process can provide labeled data for model training and further optimize the recognition ability of the model.
[0016] Preferably, the step S4 of constructing an autism spectrum disorder recognition model based on the interaction disorder matching data and the communication disorder matching data includes: Extracting high-dimensional feature vectors of interaction barrier matching data and communication barrier matching data; The high-dimensional feature vector is input into the deep learning model, and the MLP technology is used for feature fusion. In the deep learning model, the attention mechanism is introduced to perform weighted processing on different feature vectors to generate a comprehensive feature representation; Use the autoencoder to reduce the dimension of the comprehensive feature representation to obtain the feature dimension reduction representation data; Input the feature dimension reduction representation data into the classifier and use the decision tree for classification to generate preliminary recognition results; Use sliding window technology to smooth the preliminary recognition results to obtain processed recognition results; Perform secondary matching on the processed recognition results and the known patterns in the history library, and align the time series to obtain the secondary matching results; An autism spectrum disorder recognition model was constructed based on the secondary matching results.
[0017] The present invention inputs high-dimensional feature vectors into a deep learning model, uses MLP technology for feature fusion, and introduces an attention mechanism to perform weighted processing on different feature vectors to generate a comprehensive feature representation. This process can effectively integrate multi-source feature information and highlight important features; the autoencoder is used to reduce the dimension of the comprehensive feature representation to obtain feature dimension reduction representation data. The autoencoder can remove redundant information, retain key features, improve the computational efficiency and generalization ability of the model, and reduce the risk of overfitting; the feature dimension reduction representation data is input into a classifier, and a decision tree is used for classification to generate a preliminary recognition result. The decision tree classifier can provide clear decision logic, quickly distinguish different categories, and provide basic recognition results for subsequent processing; the sliding window technology is used to smooth the preliminary recognition results to obtain the processed recognition results. The sliding window technology can smooth the noise and mutation in the recognition results, enhance the stability and continuity of the results, and improve the reliability of the recognition results; the processed recognition results are matched with the known patterns in the history library for a second time, and the time series are aligned to obtain the secondary matching results. Secondary matching can further verify the accuracy of the recognition results, and time series alignment ensures the consistency of the results with the known patterns in the time dimension, improving the recognition accuracy of the model; an autism spectrum disorder recognition model is constructed based on the secondary matching results. Through multi-step feature processing and result optimization, the model can more accurately identify the characteristics of autism spectrum disorders and improve the accuracy and reliability of recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A schematic diagram of the steps of a method for constructing an autism spectrum disorder recognition model based on deep learning; Figure 2 A schematic flow chart of detailed implementation steps for obtaining a video of a conversation with an autistic patient in step S1; The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0019] The technical method of the present invention is described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by technicians in this field without creative work are within the scope of protection of the present invention.
[0020] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor methods and / or microcontroller methods.
[0021] It should be understood that, although the terms "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are used only to distinguish one unit from another unit. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.
[0022] To achieve this, please refer to Figure 1 to Figure 2 , a method for constructing an autism spectrum disorder recognition model based on deep learning, the method comprising the following steps: Step S1: obtaining a conversation video of an autistic patient; continuously measuring the eye gaze frequency in the conversation video of the autistic patient; when detecting that the eye gaze frequency is lower than an eye gaze frequency threshold, determining that the state is an eye gaze reduction state and recording the eye gaze duration; Step S2: Performing social language analysis on the conversation video of the autistic patient to obtain social language data; identifying the facial reaction features of the patient in the conversation video of the autistic patient based on the social language data and recording the facial muscle reaction data; identifying the body posture features in the conversation video of the autistic patient based on the social language data and recording the body posture method data; Step S3: evaluating communication disorder based on the gaze duration and facial muscle reaction data to obtain a communication disorder pattern; evaluating social interaction disorder based on the gaze duration and body posture data to obtain a social interaction disorder pattern; Step S4: Obtain an autism spectrum disorder history library; perform feature matching on the social interaction disorder pattern and the communication disorder pattern with the autism spectrum disorder history library respectively to generate interaction disorder matching data and communication disorder matching data; and construct an autism spectrum disorder recognition model based on the interaction disorder matching data and the communication disorder matching data.
[0023] The present invention can convert the patient's gaze behavior into specific quantifiable data, namely, gaze frequency and gaze duration, by acquiring autistic patients' conversation videos and continuously monitoring the frequency of gazes therein. This process provides an objective, accurate and operable data basis for subsequent evaluation, avoiding the deviation caused by subjective judgment; setting a gaze frequency threshold, when it is detected that the gaze frequency is lower than the threshold, it can timely and accurately determine that the patient is in a state of reduced gaze, and record the duration of gaze. This helps to capture the abnormal changes in the gaze behavior of autistic patients in social interactions, and provides a key basis for further analyzing their communication disorders. Social language analysis is performed on the conversation videos of autistic patients to obtain social language data, and at the same time, the facial reaction features of the patient are identified based on the data and recorded as facial muscle reaction data, and the body posture features are identified and recorded as body posture mode data. This process realizes the comprehensive collection of multi-dimensional data from language communication, facial expressions to body movements, providing rich and complete information for subsequent comprehensive evaluation; social language data can reflect the patient's language expression and communication ability, facial muscle reaction data can reflect the patient's emotional response to social situations, and body posture data reveals the patient's body language performance in social interaction. The comprehensive analysis of these three types of data can comprehensively and objectively reflect the comprehensive ability of autistic patients in the social process, and provide strong support for accurately identifying and evaluating their degree of social impairment. Communication impairment assessment is based on gaze duration and facial muscle reaction data to obtain communication impairment patterns; social interaction impairment assessment is based on gaze duration and body posture data to obtain social interaction impairment patterns. This comprehensive evaluation method based on multi-dimensional data can accurately identify the specific impairment patterns of autistic patients in communication and social interaction, and provide clear directions and basis for subsequent personalized intervention; by integrating multiple data for evaluation, the limitations of single indicator evaluation are avoided. For example, evaluation based only on language communication ability may ignore the patient's impairment in non-verbal social interaction, while combining data such as facial expressions and body movements can more comprehensively reflect the patient's social ability and improve the accuracy and reliability of the evaluation. Obtain an autism spectrum disorder history library, and perform feature matching of social interaction disorder patterns and communication disorder patterns with the history library, respectively, to generate interaction disorder matching data and communication disorder matching data, and construct an autism spectrum disorder recognition model based on the interaction disorder matching data and the communication disorder matching data. Therefore, the present invention uses data processing technology and deep learning technology to evaluate the communication disorder and social interaction disorder of autistic patients, and construct an autism spectrum disorder recognition model, thereby improving the recognition efficiency and accuracy of autism spectrum disorders.
[0024] In the embodiment of the present invention, reference Figure 1As shown, it is a schematic diagram of the steps of a method for constructing an autism spectrum disorder recognition model based on deep learning of the present invention. In this example, the method for constructing an autism spectrum disorder recognition model based on deep learning includes the following steps: Step S1: obtaining a conversation video of an autistic patient; continuously measuring the eye gaze frequency in the conversation video of the autistic patient; when detecting that the eye gaze frequency is lower than an eye gaze frequency threshold, determining that the state is an eye gaze reduction state and recording the eye gaze duration; In the embodiment of the present invention, a high-resolution camera is selected to ensure that its resolution reaches or exceeds 1920×1080 pixels, and the frame rate is set to 30 frames per second to ensure the clarity and smoothness of the video, and to accurately capture various subtle movements and expressions of the patient. The camera is fixed on a tripod to ensure its stability. The shooting angle of the camera should be facing the patient, ensuring that the patient's face and upper body occupy the main position in the picture, which is convenient for subsequent analysis. The distance between the camera and the patient should be adjusted according to the actual scene, generally maintained at about 1.5-2 meters, to ensure that the patient's whole body is included in the picture while the facial details are clearly discernible. Choose a well-lit and uniform shooting environment to avoid direct strong light or shadow interference, and ensure that the brightness and contrast of the video picture are moderate. If there is insufficient natural light, auxiliary lighting equipment, such as evenly distributed LED lights, can be used to ensure soft light and no obvious reflections. Before starting recording, warm up and test the camera to ensure that the equipment is working properly and the picture is clear and stable. Check the picture clarity, color reproduction, and whether there is distortion through the preview function of the camera or the video monitoring software connected to the computer. Start the video recording function and record the entire process of the patient's conversation with others. During the recording process, monitor the video screen in real time to ensure that there is no occlusion, jitter or other interference factors that affect the video quality. If any problems are found, pause the recording and readjust the device in time. Store the recorded video file in a storage device with sufficient capacity, such as a large-capacity hard disk or solid-state drive, to ensure the integrity and security of the video data. The video file format should be a common format, such as MP4 or AVI, for subsequent processing and analysis. Perform preliminary processing on the video, including cutting out irrelevant beginning and ending parts to ensure that the video content only contains the complete process of the patient's conversation. Use video editing software (such as Adobe Premiere Pro) to crop the video, and check whether the video parameters such as frame rate and resolution meet the requirements. If necessary, make corresponding adjustments. Use the Tobii TX300 eye tracker for eye tracking. Install and set up the Tobii Experience software to ensure that the software version is compatible with the eye tracker. Before calibration, ensure that the relative position of the eye tracker and the screen is correct, and align the two white lines above the eye tracker with the lines on the screen. Then, select the eye to be detected and perform the calibration operation according to the software prompts. During the calibration process, let the patient look at multiple fixed points on the screen and follow the small dots on the screen in turn until they break to complete the calibration. During the entire calibration process and subsequent monitoring process, the patient should try to keep the body position still and display multiple calibration targets to cover the participant's entire field of view. At the beginning and end of the experiment, and after the eye tracker equipment is moved, a verification procedure must be performed to ensure the accuracy of the calibration. The eye tracker will record the patient's gaze data in real time, including the coordinates of the gaze point, gaze duration and other information.The system calculates the gaze frequency, i.e., the number of times the patient's gaze is made per unit time, through a preset algorithm. The specific calculation method is: within a set time window (such as 1 second), count the effective number of times the patient's gaze is made, and then divide the number by the length of the time window to obtain the gaze frequency. Set a gaze frequency threshold, which is determined based on the average gaze frequency of the normal population and the statistical difference in the gaze frequency of autistic patients. For example, through statistical analysis of a large number of sample data, the average gaze frequency of the normal population is 5 times per second, while the gaze frequency of autistic patients is usually lower than a certain percentage of this value (such as 20%), so the threshold is set to 4 times per second. When the gaze frequency is detected to be lower than the set threshold, the system determines that the patient is in a state of reduced gaze, and immediately starts the timing function to record the duration of gaze. The timer accurately records the time period from the gaze frequency falling below the threshold to the restoration of the normal gaze frequency in milliseconds. This time period is the duration of reduced gaze.
[0025] Step S2: Performing social language analysis on the conversation video of the autistic patient to obtain social language data; identifying the facial reaction features of the patient in the conversation video of the autistic patient based on the social language data and recording the facial muscle reaction data; identifying the body posture features in the conversation video of the autistic patient based on the social language data and recording the body posture method data; In an embodiment of the present invention, a professional audio processing software (such as Audacity) is used to extract a voice track from a video file to ensure that the extracted voice is clear and free of noise. The extracted voice file is imported into a speech recognition tool (such as IBM Watson Speech to Text or Microsoft Azure Speech Service), the language model is set to Mandarin or English (selected according to the patient's language), and the transcription function is started to convert the voice into a text format and save it as a TXT file. The transcribed text is manually proofread to ensure that the text accuracy is above 95%. The transcribed text is encoded using a text analysis tool (such as NVivo) to mark key social language features such as greetings, questions, responses, repetitive language, etc. The frequency of occurrence of various language features in the text is counted, such as the number of times greetings are used, the number of times questions are asked, etc., and a social language data table is generated, including indicators such as vocabulary richness (the total number of different words), communication turns (the number of conversations back and forth). The video is processed frame by frame using video analysis software (such as OpenCV), and a pre-trained facial detection model (such as HaarCascadeClassifier) is loaded to locate the patient's facial area in each frame. Set the detection parameters: the minimum number of neighbors is 5, the scaling factor is 1.1, and the accuracy of face detection is not less than 95%. Track the detected facial area and record the position changes of the face in the video. In the detected facial area, use facial expression analysis tools (such as AffectivaSDK) to identify facial expressions and record the time point and duration of each expression (such as happiness, sadness, surprise, etc.). Generate a facial muscle reaction data file, record the expression category and its corresponding duration of each frame, in the format of a CSV file, containing fields such as timestamp, expression category, and duration. Use a posture estimation tool (such as OpenPose) to process the video and extract the patient's limb key points in each frame, including the head, shoulders, elbows, wrists, knees, ankles, etc. Set the parameters of OpenPose: the detection threshold is 0.5, the key point connection distance is 50 pixels, and the accuracy of key point detection is not less than 90%. Export the detected key point data in JSON format, including the key point coordinate information of each frame. Use Python scripts to parse JSON files, determine the angles and distances between key points, and identify body posture features, such as crossed arms, hands on the head, and body leaning forward. Record the time and duration of each posture, and generate a body posture data file in the form of a CSV file, which contains fields such as timestamp, posture category, and duration.
[0026] Step S3: evaluating communication disorder based on the gaze duration and facial muscle reaction data to obtain a communication disorder pattern; evaluating social interaction disorder based on the gaze duration and body posture data to obtain a social interaction disorder pattern; In the embodiment of the present invention, the gaze duration data and the facial muscle reaction data are aligned according to the timestamp to ensure that the gaze duration at each moment can match the corresponding facial muscle reaction data. For example, if the gaze duration is recorded as from the 5th second to the 10th second, the facial muscle reaction data extracts the expression category and its duration in the same time period. For each time period, the total duration of gaze, the category and duration of facial expression are counted. For example, the proportion of gaze duration in a certain time period and the number of conversions of facial expression are determined. A threshold value of gaze duration is set. For example, if the gaze duration is less than 30% of the time period, it is marked as gaze reduction. A threshold value of facial expression conversion frequency is set. For example, if the number of facial expression conversions is less than 5 times per minute, it is marked as low expression conversion frequency. According to the above two marking results, it is judged whether a communication barrier mode occurs. If both gaze reduction and expression conversion frequency are met, it is determined to be a communication barrier mode, and the start time and duration of the time period are recorded. The evaluation results are recorded in a table, including information such as timestamp, gaze duration, facial expression category and duration, whether it is a communication barrier mode, etc. For example, a row in the table is displayed as: timestamp (5 seconds-10 seconds), gaze duration (3 seconds), facial expression (happy, lasting 2 seconds; sad, lasting 1 second), communication disorder mode (yes). Align the gaze duration data with the body posture data according to the timestamp to ensure that the gaze duration at each moment matches the corresponding body posture data. For example, if the gaze duration is recorded from the 15th second to the 20th second, the body posture data extracts the body posture category and its duration in the same time period. For each time period, the total duration of gaze and the category and duration of body posture are counted. For example, determine the proportion of gaze duration in a certain time period and the number of body posture conversions in the time period. Set a threshold for gaze duration. For example, if the gaze duration is less than 30% of the time period, it is marked as reduced gaze. Set a threshold for the body posture conversion frequency. For example, if the number of body posture conversions is less than 3 times per minute, it is marked as low body posture conversion frequency. Based on the above two marking results, determine whether a social interaction disorder mode occurs. If both the reduced gaze and low frequency of body posture changes are met, it is determined to be a social interaction disorder mode, and the start time and duration of this time period are recorded. The evaluation results are recorded in a table, including timestamp, gaze duration, body posture category and duration, whether it is a social interaction disorder mode, and other information. For example, a row in the table may be displayed as: timestamp (15 seconds-20 seconds), gaze duration (4 seconds), body posture (arms crossed, lasting 3 seconds; hands holding head, lasting 1 second), social interaction disorder mode (yes).
[0027] Step S4: Obtain an autism spectrum disorder history library; perform feature matching on the social interaction disorder pattern and the communication disorder pattern with the autism spectrum disorder history library respectively to generate interaction disorder matching data and communication disorder matching data; and construct an autism spectrum disorder recognition model based on the interaction disorder matching data and the communication disorder matching data.
[0028] In an embodiment of the present invention, historical data is obtained from the open source autism spectrum disorder knowledge base AsdKB, which is a Chinese knowledge base for early screening of autism spectrum disorders and covers multiple information sources, including ICD-10 and DSM-5. Feature data related to social interaction disorder and communication disorder are extracted; the social interaction disorder pattern and the communication disorder pattern are matched with the feature data in the autism spectrum disorder history library. The specific operation is as follows: For the social interaction disorder pattern, its key features are extracted, such as low frequency of body posture conversion, short duration of eye gaze, etc., and compared with the symptom instances related to social interaction disorder recorded in AsdKB. For the communication disorder pattern, its key features are extracted, such as low frequency of facial expression conversion, reduced eye gaze, etc., and compared with the symptom instances related to communication disorder recorded in AsdKB. The similarity between features is quantified using cosine similarity. Specifically determine the cosine similarity between the feature vector of the current pattern and the feature vector in the history library, and the match is considered successful if the similarity is higher than the set threshold (such as 0.8). Record the matching results and generate interaction disorder matching data and communication disorder matching data, including the matching feature names and similarity score information. Build a recognition model based on the interaction disorder matching data and communication disorder matching data. Select the XGBoost algorithm, which can handle unbalanced data and provide high-accuracy classification results. Preprocess the data, including handling unbalanced data (such as by setting weights) and dividing the data into training sets and validation sets (such as in a 7:3 ratio). Use the training set to train the XGBoost model and optimize the model parameters. The general parameter settings are, booster: select gbtree as the booster. silent: set to 0 to print run messages. nthread: set to the maximum number of available threads. Tree boosting settings are, eta: set to 0.3 to control the learning rate. gamma: set to 0 to control the minimum loss reduction when splitting a node. max_depth: set to 6 to control the maximum depth of the tree. min_child_weight: set to 1 to control the minimum weight of leaf nodes. subsample: set to 1 to control the sample sampling ratio. colsample_bytree: set to 1 to control the feature sampling ratio of each tree. lambda: set to 1 to control the L2 regularization term. alpha: set to 0 to control the L1 regularization term. The learning task parameters are set to, objective: set to binary:logistic, for binary classification problems, output probability. eval_metric: set to auc, use AUC as the evaluation metric for the validation set.
[0029] As an example of the present invention, refer to Figure 2 As shown, the acquisition of the conversation video of the autistic patient in step S1 includes: S11: Record conversation videos of autistic patients in different preset scenarios. During the recording process, the patient is at the center of the camera, and the horizontal angle between the camera and the patient's eyes is kept within 10°; S12: Perform a preliminary screening of the recorded videos to remove the clips with blurry images for more than 5 seconds due to camera shaking, while retaining the clips in which the patient's facial expressions and body movements are shown in the video; S13: cropping the retained video clips, cropping the video frames of the first 3 seconds and the last 2 seconds in the video before the patient enters the conversation state; S14: splicing the cropped video clips according to the timestamps of the patient's conversation, and the total duration of the spliced video is no less than 1 minute to obtain a conversation video of the autistic patient.
[0030] In an embodiment of the present invention, a plurality of different social scenes are designed, such as a family environment, a school environment, and a social gathering environment. Each scene should have a clear background and interactive objects to simulate the social behavior of autistic patients in different environments. Use a high-definition camera for recording, ensure that the resolution of the camera is not less than 1920×1080 pixels, and set the frame rate to 30 frames per second. Adjust the position and angle of the camera, and use an angle measurement tool (such as an angle meter) to ensure that the horizontal angle between the lens and the patient's eyes is kept within 10°. Use a ruler to ensure that the patient is always in the center of the lens, and the adjustment can be assisted by marking the center point of the lens and the patient's position. In each preset scene, the patient's conversation video is recorded separately. During the recording process, the video screen is monitored in real time to ensure that the patient is always in the center of the lens. After the recording is completed, save the video file and ensure that the file format is MP4 or AVI for subsequent processing. Use a video analysis tool (such as Video-Analyzer) to perform a preliminary screening of the recorded video. Through the frame stability analysis function of Video-Analyzer, check the stability of the video frame by frame, and remove the fragments that are blurred for more than 5 seconds due to lens shaking. At the same time, use the expression recognition function of Video-Analyzer to mark the clips in which the patient's expression changes in the video. Use body movement detection tools (such as OpenPose) to mark the clips in which the patient's body movements appear in the video, and retain the marked expression changes and body movement clips. Use video editing software (such as Adobe Premiere Pro) to crop the retained video clips, and according to the video timestamp, accurately crop the first 3 seconds and last 2 seconds of the video frame in which the patient does not enter the conversation state. Splice the cropped video clips according to the timestamp of the patient's conversation, use video editing software to ensure that the total length of the spliced video is not less than 1 minute, and save the final spliced video file, ensuring that the file format is MP4 or AVI.
[0031] Preferably, the eye gaze frequency in the continuous autistic patient conversation video in step S1 includes: Analyze the conversation videos of autistic patients frame by frame to determine the patient's eye direction in each frame; In each frame, determine whether the patient's eye direction forms an angle less than 30° with the camera direction. If the eye direction forms an angle less than 30° with the camera direction, record the frame as a valid gaze. Take every 10 seconds as a statistical unit, count the number of frames with effective gaze in each statistical unit, and get the gaze frequency of each statistical unit; Divide the entire video into the above statistical units, calculate the gaze frequency of each statistical unit, and record it as the frequency value of the statistical unit; The frequency values of all statistical units are summarized and the average eye gaze frequency of the entire video is calculated as the eye gaze frequency in the conversation video of the autistic patient.
[0032] In an embodiment of the present invention, the OpenCV library is used to load a video file and read the video content frame by frame. The frame reading interval is set to 1 frame to ensure that the video is analyzed frame by frame. Each frame of the image is grayed and the color image is converted into a gray image to reduce the interference of color information and improve the efficiency of subsequent processing. The gray image is Gaussian blurred to remove noise in the image and smooth the image. The kernel size of the Gaussian blur is set to 5×5 and the standard deviation is set to 1.5 to ensure the smoothing effect of the image. The Haar feature cascade classifier in OpenCV is used for eye detection. The pre-trained Haar cascade file (such as haarcascade_eye.xml) is loaded to locate the eye area in each frame of the image. The detected eye area is geometrically analyzed to determine the center point of the eye. The specific method is to take the average value of the coordinates of the upper left corner and the lower right corner of the eye rectangular box to calculate the center coordinates of the eye. Assuming that the lens direction is horizontal (i.e., parallel to the horizontal axis of the video), the angle between the center point of the eye and the horizontal axis of the video is calculated. The specific method is to calculate the horizontal and vertical distances between the center point of the eye and the center point of the video, and then calculate the angle by the inverse tangent function. Convert the calculated angle to degrees for subsequent judgment. Determine whether the calculated angle is less than 30 degrees. If the angle is less than 30 degrees, mark the frame as a valid gaze frame. To ensure accuracy, calculate the angle for each detected eye region separately, and take the average of the angles of the two eyes as the final judgment basis. If the average of the angles of the two eyes is less than 30 degrees, the frame is marked as a valid gaze frame. According to the frame rate of the video (assuming 30 frames / second), every 10 seconds contains 300 frames. Divide the total number of video frames by 300 to get the total number of statistical units. Traverse the frames in each statistical unit and count the number of frames marked as valid gaze. Calculate the gaze frequency of each statistical unit, that is, the number of valid gaze frames divided by the total number of frames in the statistical unit. For example, if 150 frames in a statistical unit are marked as valid gazes and the total number of frames is 300, then the gaze frequency of the statistical unit is 150 divided by 300, which is 0.5, or 50%. Sum up the gaze frequencies of all statistical units and calculate the average gaze frequency of the entire video. The specific method is to add up the gaze frequencies of all statistical units and then divide by the total number of statistical units. For example, if there are 10 statistical units, and the gaze frequencies of each statistical unit are 0.4, 0.5, 0.6, etc., add these frequency values and divide by 10, and the result is the average gaze frequency of the entire video.
[0033] It is particularly important that when the gaze frequency is detected to be lower than the gaze frequency threshold, determining it as a gaze reduction state and recording the gaze duration includes: Performing time series analysis on the gaze frequency data, dividing the gaze frequency data into multiple time windows of preset lengths in chronological order, and treating the gaze frequency data in each time window as an independent analysis unit; In each time window, the gaze frequency data is smoothed to obtain a smoothed gaze frequency curve; Based on the smoothed gaze frequency curve, the gaze frequency change rate in each time window is calculated, and the gaze frequency trend is determined by comparing the gaze frequency change rates of adjacent time windows. When the downward trend of the gaze frequency shows a continuous downward trend, the time point when the gaze frequency is lower than the gaze frequency threshold is determined, and the time point is taken as the starting point of the gaze reduction state, and the gaze duration is started to be recorded; In the process of recording the gaze duration, the changes in the gaze frequency are monitored in real time; if the gaze frequency is higher than the gaze frequency threshold again, the recording of the gaze duration is stopped; the entire process time from the start of recording to the stop of recording is taken as the gaze duration.
[0034] In an embodiment of the present invention, the gaze frequency data is divided into a plurality of time windows of preset lengths in chronological order, and the gaze frequency data in each time window is used as an independent analysis unit. For example, assuming that the preset time window length is 10 seconds and the video frame rate is 30 frames / second, each time window contains 300 frames of data. In each time window, the gaze frequency data is smoothed to reduce noise and highlight trends. The exponential smoothing method is used for smoothing, and the specific steps are as follows: a smoothing coefficient α is selected, for example, α=0.3. The closer the α value is to 1, the more sensitive the smoothed data is to the nearest data point. The data points in each time window are subjected to exponential smoothing, and the calculation formula is: current smoothing value=α×current data point+(1-α)×previous smoothing value, so as to obtain the gaze frequency curve after smoothing in each time window. Based on the smoothed gaze frequency curve, the gaze frequency change rate in each time window is calculated. The specific method is to calculate the difference between the gaze frequencies of two adjacent time points and then divide it by the frequency value of the previous time point. The gaze frequency trend is determined by comparing the gaze frequency change rates of adjacent time windows. If the rates of change of multiple consecutive time windows are all negative and the amplitude of change gradually increases, it is judged that the gaze frequency shows a continuous downward trend. When the downward trend of the gaze frequency shows a continuous downward trend, determine the time point when the gaze frequency is lower than the gaze frequency threshold. For example, set the gaze frequency threshold to 0.3. If the gaze frequency in a time window is lower than 0.3, the starting time point of the time window is used as the starting point of the gaze reduction state, and the gaze duration is recorded from this starting point. In the process of recording the gaze duration, the changes in the gaze frequency are monitored in real time. If the gaze frequency is higher than the gaze frequency threshold again, stop recording the gaze duration, and the entire process time from the start of recording to the stop of recording is used as the gaze duration.
[0035] Preferably, the social language analysis of the conversation video of the autistic patient in step S2 includes: Separate audio signals from autistic patients’ conversation videos, convert them into text, and record the pronunciation duration and intonation changes of each word; Perform word segmentation on the text, split the text into individual words, and mark each word with its appearance timestamp in the audio; Extract the pronunciation duration and intonation changes of each word and generate a speech feature vector corresponding to each word; Count the frequency of occurrence of each word in the text, and combine it with the speech feature vector to mark the sentiment tendency of each word; Identify interrogative sentences and declarative sentences in the text according to the sentiment tendency, count the number of interrogative sentences and declarative sentences, and record the average intonation change of each sentence type; The order of word usage in the text is marked based on the average intonation change, and the collocation relationship between each word is recorded to obtain social language data.
[0036] In an embodiment of the present invention, an audio processing software (such as Audacity) is used to separate an audio signal from a video of a conversation of an autistic patient to ensure that the audio is clear and free of noise. The separated audio signal is imported into a speech recognition tool (such as Google Speech-to-Text API), the language model is set to Mandarin or English (selected according to the patient's language), and the transcription function is started to convert the speech into text. During the speech recognition process, the audio analysis function is enabled to record the pronunciation duration and intonation changes of each word. Specifically, the pronunciation duration is recorded in milliseconds, and the intonation changes are quantified by analyzing the frequency fluctuations of the audio, expressed as the frequency change amplitude (Hz). The transcribed text is segmented using a natural language processing tool (such as NLTK or Jieba) to split the text into individual words. Each word is marked with its appearance timestamp in the audio. The timestamp is in seconds, accurate to two decimal places, indicating the start and end time of the word in the audio. The pronunciation duration and intonation changes of each word are extracted to generate a speech feature vector corresponding to each word. The speech feature vector contains two dimensions: pronunciation duration (unit: milliseconds) and intonation variation (unit: Hz). The frequency of occurrence of each word in the text is counted, and the emotional tendency of each word is marked in combination with the speech feature vector. The emotional tendency is determined by analyzing the intonation variation: rising intonation usually indicates positive emotion, falling intonation indicates negative emotion, and stable intonation indicates neutral emotion. According to the emotional tendency, interrogative sentences and declarative sentences in the text are identified. Specifically, interrogative sentences usually start with question words (such as "what" and "why"), and the intonation rises at the end of the sentence; the intonation of declarative sentences is relatively stable. The number of interrogative sentences and declarative sentences is counted, and the average intonation variation of each sentence type is recorded. The average intonation variation is determined by calculating the average intonation variation in all corresponding sentence types. The order of word usage in the text is marked based on the average intonation variation, and the collocation relationship of each word is recorded. Specifically, the intonation variation trend between words is analyzed, which words tend to appear in positive or negative contexts are marked, and their collocation patterns are recorded. The above information is integrated to generate social language data, including the frequency of word occurrence, emotional tendency, intonation change range and collocation relationship, which is stored in a table form for subsequent analysis.
[0037] Preferably, the step S2 of identifying facial reaction features of autistic patients in a conversation video according to social language data and recording facial muscle reaction data includes: Extract each frame of image from the conversation video of autistic patients, perform face detection on each frame of image, and locate the zygomatic major muscle area, corrugator supercilii area and medial frontalis muscle area of the patient's face; The facial contour in each frame of the image is analyzed, and the muscle activity intensity of the zygomatic major muscle area, the corrugator supercilii muscle area and the medial frontalis muscle area is recorded to generate the facial muscle activity intensity; Determine the video frame range corresponding to each word according to the timestamps of word appearance marked in the social language data; The occurrence time and duration of facial muscle activity intensity are determined within the video frame range and summarized into facial muscle response data.
[0038] In an embodiment of the present invention, the cv2.VideoCapture function of OpenCV is used to load a video file. The video content is read frame by frame by calling the cap.read() method in a loop to ensure that each frame is extracted. Each frame of the image is saved as a separate picture file in JPEG format and stored in a specified folder. The pre-trained Haar cascade file haarcascade_frontalface_default.xml is loaded using cv2.CascadeClassifier. Each frame of the image is converted into a grayscale image using the cv2.cvtColor function with the parameter cv2.COLOR_BGR2GRAY. The detectMultiScale method is called for facial detection, and the parameters are set: scaleFactor=1.2, minNeighbors=3, minSize=(32, 32). A rectangular box is drawn on each frame of the image to mark the detected facial area. The pre-trained facial key point detection model shape_predictor_68_face_landmarks.dat is loaded using Dlib's shape_predictor. For the facial area detected in each frame of the image, the predictor method is called to detect 68 facial key points. According to the key point coordinates, two key points above the corners of the mouth (such as key points 48 and 54) are selected as reference points for the zygomaticus major region. Key points below the inner side of the eyebrows (such as key points 21 and 22) are selected as reference points for the corrugator supercilii region. Key points in the central part of the forehead (such as key point 27) are selected as reference points for the medial frontalis region. For each facial region of each frame image, the cv2.findContours method is used to extract the contour. For each region (zygomaticus major, corrugator supercilii, medial frontalis), the area and perimeter of the contour are calculated using the cv2.contourArea and cv2.arcLength functions. The change in contour area is used as a quantitative indicator of muscle activity intensity. For example, the difference between the contour area of the current frame and the previous frame is calculated. The larger the difference, the higher the muscle activity intensity. The timestamps of the words marked in the social language data are converted into video frame numbers. Assuming that the video frame rate is 30 frames / second, the starting frame number corresponding to the words with a timestamp of 2 seconds is 60 (2 seconds × 30 frames / second). According to the pronunciation duration of the word, the corresponding video frame range is determined. If the pronunciation duration of a word is 1 second, the corresponding frame range is 60 to 90 frames. Within the determined video frame range, the muscle activity intensity of the zygomatic major muscle area, the corrugator supercilii area, and the medial frontalis area is analyzed frame by frame. The starting frame number of the muscle activity intensity from low to high is recorded as the occurrence time, and the ending frame number from high to low is recorded as the end point of the duration. The occurrence time, duration, and activity intensity value of the facial muscle activity intensity corresponding to each word are summarized to form a facial muscle reaction data table.The data table contains information such as words, occurrence time (frame number), duration (number of frames), and muscle activity intensity value.
[0039] Preferably, the step S2 of identifying body posture features in a conversation video of an autistic patient according to social language data and recording body posture data includes: Extract each frame of image from the conversation video, perform limb detection on each frame of image, and locate the body contour of the patient; Measure the vertical angle of the trunk of the body contour and divide the vertical angle of the trunk into three intervals: 0-15 degrees is normal posture, 15-30 degrees is mild forward leaning, and more than 30 degrees is severe forward leaning; Measure the knee extension angle of the leg of the torso silhouette, and divide the knee angle into three intervals: 160-180 degrees is fully extended, 120-160 degrees is slightly bent, and less than 120 degrees is severely bent; Determine the video frame range corresponding to each word according to the timestamps of word appearance marked in the social language data; The vertical angle of the trunk and the extension angle of the knee joint are recorded within the video frame to form limb posture data.
[0040] In an embodiment of the present invention, a video file of a conversation of an autistic patient is loaded through a video processing software. The video processing software is set to read the video content frame by frame to ensure that each frame is extracted. Each frame of the image is saved as a separate picture file in JPEG format and stored in a specified folder. The file naming format is frame_0001.jpg, frame_0002.jpg, etc., for subsequent processing. Each frame of the image is processed using a limb detection tool to extract multiple key points of the human body, including parts such as the trunk and limbs. According to the detected key points, the body contour of the patient is determined. The key points include shoulders, waists, and hips, etc. Among the detected limb key points, the key points of the shoulders and hips are selected. Calculate the angle between the line between the shoulder and hip key points and the vertical direction. The specific method is to calculate the horizontal distance and vertical distance of the two points, and then use the inverse tangent function to calculate the angle. The angle is divided into three intervals: 0-15 degrees is a normal posture, 15-30 degrees is a slight forward lean, and more than 30 degrees is a severe forward lean. Record the angle value of each frame and its corresponding posture classification in the data table, including frame number, angle value and posture classification. Select the key points of the knee and ankle joints from the detected limb key points. Calculate the extension angle of the knee joint. The specific method is to calculate the angle between the line between the key points of the knee and ankle joints and the vertical direction. Divide the angle into three intervals: 160-180 degrees is fully extended, 120-160 degrees is slightly bent, and less than 120 degrees is severely bent. Record the angle value of each frame and its corresponding bending classification in the data table, including frame number, angle value and bending classification. Load the social language data file, which contains words and their timestamps. Convert the timestamps of the words marked in the social language data to the corresponding video frame numbers. Assuming that the video frame rate is 30 frames / second, the starting frame number corresponding to the word with a timestamp of 2 seconds is 60 (2 seconds × 30 frames / second). Determine the corresponding video frame range according to the pronunciation duration of the word. For example, if the pronunciation duration of a word is 1 second, the corresponding frame range is 60 to 90 frames. Within the determined video frame range, record the vertical angle of the torso and the knee extension angle of each frame. Summarize the recorded data to form a limb posture data table. The data table contains information such as words, frame range, torso angle, knee angle, etc. Save the summarized limb posture data as a CSV file for subsequent analysis.
[0041] More importantly, the method of identifying body posture features in autistic patients' conversation videos based on social language data and recording body posture data also includes: Extract emotional expression words and corresponding timestamps from social language data, and classify the conversation content into two emotional tendencies: positive and negative; synchronously analyze the body gestures in the video based on the timestamps of social language data; When the conversation is positive, analyze whether the patient's body posture shows open characteristics; calculate the vertical distance between the center point of the shoulder and the center point of the waist. If the vertical distance gradually decreases within the time window, it is judged that the body is leaning forward; calculate the vertical distance between the key points of the wrists of both hands and the horizontal line of the shoulders. If the key points of the wrists of both hands are below the horizontal line of the shoulders and remain stable within the time window, it is judged that the hands are placed naturally; When the conversation is negative, analyze whether the target's body posture shows closed characteristics; calculate the horizontal distance between the key points of the wrists of both hands. If the horizontal distance gradually decreases within the time window and the key points of the wrists of both hands are close to the midline of the body, it is judged that the hands are crossed; calculate the angle between the waist and the back. If the angle gradually increases within the time window and exceeds 15 degrees, it is judged that the body is leaning back; Integrate body leaning forward, body leaning back, hands crossed and hands placed naturally to obtain body posture data.
[0042] In an embodiment of the present invention, social language data is read from a file, and the data includes each word and the timestamp of its occurrence. The text is segmented using a natural language processing tool to extract emotional expression vocabulary. According to a predefined emotional dictionary (such as a positive vocabulary list and a negative vocabulary list), the vocabulary is divided into positive and negative categories. The extracted emotional vocabulary and its corresponding timestamp are recorded to ensure the accuracy of the timestamp so that it can be synchronously analyzed with the video frame later. The video file is loaded using a video processing tool (such as OpenCV). The video content is read frame by frame, and each frame of the image is saved as a separate image file. Each frame of the image is processed using a limb detection tool to extract the key points of the human body, including shoulders, waist, hips, wrists, etc. According to the timestamp of the social language data, the corresponding video frame is extracted for limb posture analysis. The coordinates of the key points of the shoulder and waist are obtained by the limb detection tool. The vertical distance between the center point of the shoulder and the center point of the waist is calculated. Within a set time window (such as 5 seconds), the change in vertical distance is compared. If the vertical distance gradually decreases, it is judged that the body is leaning forward. The coordinates of the key points of the wrists of both hands are obtained by the limb detection tool. Calculate the vertical distance between the key points of the wrists of both hands and the horizontal line of the shoulders. Within the set time window (such as 5 seconds), if the key points of the wrists of both hands are below the horizontal line of the shoulders and remain stable, it is judged that the hands are placed naturally. Obtain the coordinates of the key points of the wrists of both hands through the limb detection tool, calculate the horizontal distance between the key points of the wrists of both hands, and within the set time window (such as 5 seconds), if the horizontal distance gradually decreases and the key points of the wrists of both hands are close to the midline of the body, it is judged that the hands are crossed. Obtain the coordinates of the key points of the waist and back through the limb detection tool, calculate the angle between the waist and the back, and within the set time window (such as 5 seconds), if the angle gradually increases and exceeds 15 degrees, it is judged that the body is leaning back. Integrate the forward leaning, backward leaning, crossed hands and natural placement of the hands obtained from the above analysis; record the integrated limb posture data, including emotional tendencies, limb posture characteristics and corresponding timestamps, for subsequent analysis. The recorded data is stored as a CSV file in the following format: Column 1: timestamp; Column 2: emotional tendency (positive / negative); Column 3: body posture features (body leaning forward, body leaning back, hands crossed, hands placed naturally); Column 4: corresponding video frame number.
[0043] Preferably, the communication disorder assessment based on gaze duration and facial muscle reaction data in step S3 includes: Determine the start time and end time of the gaze duration; determine the start time and end time of the facial muscle reaction data; The duration of gaze fixation was compared with the start and end time of the facial muscle reaction data; if the difference between the start time of gaze fixation and the start time of facial muscle reaction was less than 1 second, it was marked as a synchronous reaction; if the difference between the start time of gaze fixation and the start time of facial muscle reaction was more than 1 second, it was marked as an asynchronous reaction; The ratio of gaze duration to facial muscle reaction duration was calculated; if the ratio was less than 0.5, it was considered a low ratio; if the ratio was between 0.5 and 1.5, it was considered a medium ratio; if the ratio was greater than 1.5, it was considered a high ratio; Communication disorders are assessed based on the synchronization of eye gaze and facial muscle response and the ratio range; if the response is synchronous and the ratio is low, it is marked as a slow communication mode; if the response is synchronous and the ratio is high, it is marked as a rapid communication mode; if the response is asynchronous and the ratio is low, it is marked as an incoordinated communication mode; if the response is asynchronous and the ratio is high, it is marked as a disconnected communication mode; The above communication disorder assessment results are summarized into a communication disorder model.
[0044] In an embodiment of the present invention, gaze data is read from a file, and the data includes the gaze state of each frame and the corresponding frame number. Traverse the data, find the frame number where the gaze state changes from "not gazing" to "gazing", and record it as the start time of gaze gazing. Continue to traverse the data, find the frame number where the gaze state changes from "gazing" to "not gazing", and record it as the end time of gaze gazing. Gaze duration=end time-start time. Facial muscle reaction data is read from a file, and the data includes the facial muscle activity state of each frame and the corresponding frame number. Traverse the data, find the frame number where the facial muscle activity state changes from "not active" to "active", and record it as the start time of facial muscle reaction. Continue to traverse the data, find the frame number where the facial muscle activity state changes from "active" to "not active", and record it as the end time of facial muscle reaction. Facial muscle reaction duration=end time-start time. Align the gaze data and the facial muscle reaction data according to the frame number to ensure that the timestamps of the two groups of data are consistent. For each set of gaze and facial muscle reaction data, calculate the difference between the start time of gaze and the start time of facial muscle reaction. If the time difference does not exceed 1 second, it is marked as "synchronous reaction"; if the time difference exceeds 1 second, it is marked as "asynchronous reaction". For each set of gaze and facial muscle reaction data, calculate the ratio of gaze duration to facial muscle reaction duration. Determine the ratio interval: If the ratio is less than 0.5, it is judged as "low ratio". If the ratio is 0.5 to 1.5, it is judged as "medium ratio". If the ratio is greater than 1.5, it is judged as "high ratio". Evaluation mode: If synchronous reaction and low ratio: mark as "slow communication mode". If synchronous reaction and high ratio: mark as "rapid communication mode". If asynchronous reaction and low ratio: mark as "incoordinated communication mode". If asynchronous reaction and high ratio: mark as "broken communication mode". The above assessment results were summarized into a communication disorder pattern and recorded in a data table, including timestamp, gaze duration, facial muscle response duration, synchronization marker, ratio interval, and communication disorder pattern.
[0045] Preferably, the social interaction disorder assessment based on gaze duration and body posture data in step S3 includes: Determine the start time and end time of the gaze duration; determine the start time and end time of the facial muscle reaction data; set the gaze duration of more than 3 seconds as long gaze, and the gaze duration of less than 1 second as short gaze; Extract the angle between the trunk and the vertical direction and the extension angle of the knee joint from the limb posture data; set the change of the vertical angle of the trunk to 15-30 degrees as a small change of the trunk angle, and the change of the vertical angle of the trunk to more than 30 degrees as a large change of the trunk angle; The duration of gaze fixation was compared with the changes in the trunk angle and knee extension angle. If the duration of gaze fixation was long and the trunk angle had a small change, it was marked as a coordinated social interaction mode; if the duration of gaze fixation was short and the trunk angle had a large change, it was marked as a discontinuous social interaction mode. By integrating the social interaction coordination model and the social interaction disruption model, we obtain the social interaction disorder model.
[0046] In an embodiment of the present invention, an eye tracker is used to collect the gaze data of the subject, and the eye tracker records the eye movement trajectory of the subject during the video playback of a specific social scene at a high sampling rate (such as 1000Hz). The collected eye movement data is filtered to remove noise interference caused by blinking, head micro-movement, etc. The Kalman filter is used to smooth the eye movement data to reduce the random fluctuation of the data. The start time and end time of the gaze are determined by the change of the gaze point coordinates in the eye movement data. When the change of the gaze point coordinates within a certain time interval (such as 100ms) is less than a set threshold (such as 1° visual angle), it is determined to be the beginning of a gaze, until the change of the gaze point coordinates exceeds the threshold, it is determined to be the end of the gaze. According to the duration of gaze, the gaze is classified, and the gaze duration of more than 3 seconds is marked as long gaze, and the gaze duration of less than 1 second is marked as short gaze. The facial action coding system (FACS) is used in combination with a high-precision camera to collect the facial expression video data of the subject, and the frame rate of the camera is 30fps. The collected video data is subjected to inter-frame difference processing to highlight the slight movement changes of facial muscles. At the same time, the facial area in each frame is located and cropped to ensure the accuracy of subsequent analysis. By analyzing the cropped facial area image sequence, the pixel motion vector between each frame is calculated using the optical flow method. When the average value of the facial muscle motion vector exceeds the set threshold (such as 0.1 pixel / frame) within several consecutive frames (such as 5 frames), it is determined to be the beginning of the facial muscle reaction, and until the average value of the motion vector is lower than the threshold for several consecutive frames, it is determined to be the end of the facial muscle reaction. The limb posture data of the subject is collected with the help of inertial measurement unit (IMU) sensors. The IMU sensors are installed near the trunk and knee joints of the subject, respectively, with a sampling rate of 100Hz. The collected limb posture data is time-aligned to ensure that the angle data between the trunk and the vertical direction and the extension angle data of the knee joint are consistent in time. The data is smoothed with a window size of 10 sampling points. The acceleration and angular velocity data collected by the IMU sensor are used to calculate the angle between the trunk and the vertical direction and the extension angle of the knee joint using the quaternion algorithm. The change of the vertical angle of the trunk is set to 15-30 degrees as a small change of the trunk angle, and the change of the vertical angle of the trunk is set to more than 30 degrees as a large change of the trunk angle. The processed time series data are synchronized, and the gaze duration data is compared and analyzed with the change data of the trunk angle and the knee extension angle based on the timestamp. When the gaze duration is long gaze and the trunk angle is a small change, it is marked as a social interaction coordination mode; when the gaze duration is short gaze and the trunk angle is a large change, it is marked as a social interaction disruption mode. The above-mentioned marked social interaction coordination mode and social interaction disruption mode are integrated, and the occurrence frequency and duration of the two modes in a specific time period are counted to obtain the social interaction disorder mode.
[0047] Preferably, the step S4 of matching the social interaction disorder pattern and the communication disorder pattern with the autism spectrum disorder history library includes: Extract the pattern occurrence time, pattern duration and pattern frequency of the social interaction disorder pattern and record them as the social interaction disorder characteristic data; Extract the pattern occurrence time, pattern duration and pattern frequency of the communication disorder pattern and record them as communication disorder characteristic data; The social interaction disorder data and communication disorder feature data are matched with the autism spectrum disorder history library respectively, and the similarity between each pattern and the known patterns in the autism spectrum disorder history library is determined and quantified. If the quantified similarity value exceeds 0.8, the match is considered successful. Interaction barrier matching data and communication barrier matching data are recorded based on similarity quantification values.
[0048] In the embodiment of the present invention, the marked social interaction disorder pattern data are sorted by timestamp to ensure data continuity; the starting position left and the ending position right of the window are initialized, and the window size is set to 1 second and the step length is 0.5 seconds. The window is slid on the data sequence, and the right pointer is moved each time until the window covers 1 second of data. For each window, the timestamp of the window starting point is recorded as the pattern occurrence time. The difference between the timestamp of the window end point and the timestamp of the starting point is calculated as the pattern duration. Within a specific time interval (such as 1 minute), the number of patterns covered by the window is counted as the pattern frequency, and the above-extracted feature data is stored in a structured data table, and the fields include "pattern occurrence time", "pattern duration" and "pattern frequency". The communication disorder pattern data is sorted by timestamp, and the same sliding window parameters as the social interaction disorder pattern are used (window size 1 second, step length 0.5 seconds). According to the steps of the above sliding window algorithm, the pattern occurrence time, the calculated pattern duration and the statistical pattern frequency are recorded respectively. The extracted feature data is stored in a structured data table, and the fields are consistent with the social interaction disorder feature data table. The social interaction disorder feature data and communication disorder feature data are normalized, and the minimum-maximum normalization method is used to scale all feature values to the range of 0 to 1. The feature data of known patterns are extracted from the autism spectrum disorder history library and normalized in the same way. For each pattern to be matched, the Euclidean distance between it and the known patterns in the history library in the three feature dimensions of "pattern occurrence time", "pattern duration" and "pattern frequency" is calculated respectively; the distance between two feature vectors is calculated using the Euclidean distance, and the smaller the distance, the higher the similarity. The Euclidean distances of the three feature dimensions are weighted and summed to obtain the comprehensive similarity quantization value. The weights are assigned according to the importance of each feature in the identification of autism spectrum disorders. If the similarity quantization value exceeds 0.8, the match is considered successful. It is stored in the data table named "Interaction Disorder Matching Results", and the fields include "Matching Pattern", "Similarity Quantization Value" and "Matching Results". It is stored in the data table named "Communication Disorder Matching Results", and the fields are consistent with the "Interaction Disorder Matching Results" data table.
[0049] Preferably, the step S4 of constructing an autism spectrum disorder recognition model based on the interaction disorder matching data and the communication disorder matching data includes: Extracting high-dimensional feature vectors of interaction barrier matching data and communication barrier matching data; The high-dimensional feature vector is input into the deep learning model, and the MLP technology is used for feature fusion. In the deep learning model, the attention mechanism is introduced to perform weighted processing on different feature vectors to generate a comprehensive feature representation; Use the autoencoder to reduce the dimension of the comprehensive feature representation to obtain the feature dimension reduction representation data; Input the feature dimension reduction representation data into the classifier and use the decision tree for classification to generate preliminary recognition results; Use sliding window technology to smooth the preliminary recognition results to obtain processed recognition results; Perform secondary matching on the processed recognition results and the known patterns in the history library, and align the time series to obtain the secondary matching results; An autism spectrum disorder recognition model was constructed based on the secondary matching results.
[0050] In an embodiment of the present invention, feature extraction is performed on the interaction barrier matching data and the communication barrier matching data, and the convolutional neural network (CNN) structure in deep learning is used. A CNN model comprising multiple convolutional layers and pooling layers is constructed, the input is the feature matrix of the matching data, and the output is a high-dimensional feature vector. For example, the dimension of the input feature matrix is 100×50 (assuming 100 feature points, each feature point is 50 dimensions), and after being processed by the CNN model, the dimension of the output high-dimensional feature vector is 128 dimensions. A multi-layer perceptron (MLP) model is constructed, comprising multiple hidden layers, and each layer uses a ReLU activation function. The input is a high-dimensional feature vector, the number of neurons in the hidden layer is 256, 128, and 64, respectively, and the number of neurons in the output layer is 32. An attention mechanism is introduced into the MLP model, which is specifically implemented by adding an attention layer. The weights of the attention layer are automatically learned through training, and different feature vectors are weighted to generate a comprehensive feature representation. For example, the dimension of the input high-dimensional feature vector is 128 dimensions, and after being processed by the MLP and attention mechanism, the dimension of the generated comprehensive feature representation is 32 dimensions. Construct an autoencoder model. The encoder part compresses the input comprehensive feature representation into a low-dimensional representation, and the decoder part reconstructs the low-dimensional representation to the original size. The encoder structure is: 32-dimensional input layer, 16-dimensional hidden layer, and 8-dimensional output layer. The decoder structure is: 8-dimensional input layer, 16-dimensional hidden layer, and 32-dimensional output layer. Use mean square error (MSE) as the loss function and Adam optimizer for training. Train the autoencoder model with the comprehensive feature representation as input and the feature dimension reduction representation data as output, with a dimension of 8. Construct a decision tree classifier with the feature dimension reduction representation data as input and the preliminary recognition result as output. Use the trained decision tree model to classify the feature dimension reduction representation data and generate preliminary recognition results. For example, the dimension of the feature dimension reduction representation data is 8 dimensions, and the decision tree model classifies according to these features and outputs preliminary recognition results. Use the sliding window technology to smooth the preliminary recognition results, with the window size set to 5 and the step size set to 1. Slide the window on the preliminary recognition result sequence, count the recognition results in the window, and take the recognition result with the most occurrences as the final recognition result of the window. For example, the sequence of preliminary recognition results is [1,2,1,1,2,1,1,1]. After sliding window smoothing, the processed recognition result sequence is [1,1,1,1]. The processed recognition results are matched with the known patterns in the history library for a second time, and the dynamic time warping (DTW) algorithm is used to align the time series. The similarity between the processed recognition results and the known patterns in the history library is calculated. If the similarity exceeds the set threshold (such as 0.8), the match is considered successful. For example, the processed recognition result sequence is [1,1,1,1], and the known pattern sequence in the history library is [1,1,1,1]. After alignment by the DTW algorithm, the similarity is 0.9, and the match is successful.Based on the secondary matching results, an autism spectrum disorder recognition model is constructed. The model input is the processed recognition result, and the output is the recognition result of autism spectrum disorder. Use the trained model to recognize new data and generate the final recognition result. For example, the recognition result sequence after input processing is [1,1,1,1], and the model outputs the recognition result of autism spectrum disorder as "yes".
[0051] Therefore, the embodiments should be regarded as illustrative and non-restrictive from all points, and the scope of the present invention is limited by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and range of equivalent elements of the application documents are included in the present invention.
[0052] The above description is only a specific embodiment of the present invention, so that those skilled in the art can understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but should conform to the widest scope consistent with the principles and novel features invented herein.
Claims
1. A method for constructing an autism spectrum disorder recognition model based on deep learning, characterized in that: The following steps are involved: Step S1: Obtaining a video of a conversation between an autistic patient; The frequency of eye gaze in the video of the conversation of the autistic patient; when the eye gaze frequency is detected to be lower than the eye gaze frequency threshold, it is determined as the state of eye gaze reduction and the eye gaze duration is recorded; Step S2: Performing social language analysis on the conversation video of the autistic patient to obtain social language data; identifying the facial reaction features of the patient in the conversation video of the autistic patient based on the social language data and recording the facial muscle reaction data; identifying the body posture features in the conversation video of the autistic patient based on the social language data and recording the body posture method data; Step S3: evaluating communication disorder based on the gaze duration and facial muscle reaction data to obtain a communication disorder pattern; evaluating social interaction disorder based on the gaze duration and body posture data to obtain a social interaction disorder pattern; Step S4: Acquire the autism spectrum disorder history library; Perform feature matching of social interaction disorder patterns and communication disorder patterns with the autism spectrum disorder history database to generate interaction disorder matching data and communication disorder matching data; An autism spectrum disorder identification model is constructed based on interaction disorder matching data and communication disorder matching data.
2. The method for constructing an autism spectrum disorder recognition model based on deep learning according to claim 1, characterized in that: The step S1 of obtaining the conversation video of the autistic patient includes: In multiple preset scenarios, the conversation videos of autistic patients were recorded. During the recording process, the patient was at the center of the camera, and the horizontal angle between the camera and the patient's eyes was kept within 10°. Perform a preliminary screening of the recorded videos to remove the clips with blurry images for more than 5 seconds due to camera shaking, while retaining the clips in which the patient's facial expressions and body movements are shown; The retained video clips were cropped, and the video frames of the first 3 seconds and the last 2 seconds before the patient entered the conversation state were cropped; The cropped video segments are spliced together according to the timestamps of the patient's conversation. The total duration of the spliced video should be no less than 1 minute to obtain a conversation video of an autistic patient.
3. The method for constructing an autism spectrum disorder recognition model based on deep learning according to claim 1, characterized in that: The eye gaze frequency in the continuous autistic patient conversation video in step S1 includes: Analyze the conversation videos of autistic patients frame by frame to determine the patient's eye direction in each frame; In each frame, determine whether the patient's eye direction forms an angle less than 30° with the camera direction. If the eye direction forms an angle less than 30° with the camera direction, record the frame as a valid gaze. Take every 10 seconds as a statistical unit, count the number of frames with effective gaze in each statistical unit, and get the gaze frequency of each statistical unit; Divide the entire video into the above statistical units, calculate the gaze frequency of each statistical unit, and record it as the frequency value of the statistical unit; The frequency values of all statistical units are summarized and the average eye gaze frequency of the entire video is calculated as the eye gaze frequency in the conversation video of the autistic patient.
4. The method for constructing an autism spectrum disorder recognition model based on deep learning according to claim 1, characterized in that: The social language analysis of the conversation video of the autistic patient described in step S2 includes: Separate audio signals from autistic patients’ conversation videos, convert them into text, and record the pronunciation duration and intonation changes of each word; Perform word segmentation on the text, split the text into individual words, and mark each word with its appearance timestamp in the audio; Extract the pronunciation duration and intonation changes of each word and generate a speech feature vector corresponding to each word; Count the frequency of occurrence of each word in the text, and combine it with the speech feature vector to mark the sentiment tendency of each word; Identify interrogative sentences and declarative sentences in the text according to the sentiment tendency, count the number of interrogative sentences and declarative sentences, and record the average intonation change of each sentence type; The order of word usage in the text is marked based on the average intonation change, and the collocation relationship between each word is recorded to obtain social language data.
5. The method for constructing an autism spectrum disorder recognition model based on deep learning according to claim 1, characterized in that: The step S2 of identifying the facial reaction features of the autistic patient in the conversation video according to the social language data and recording the facial muscle reaction data includes: Extract each frame of image from the conversation video of autistic patients, perform face detection on each frame of image, and locate the zygomatic major muscle area, corrugator supercilii area and medial frontalis muscle area of the patient's face; The facial contour in each frame of the image is analyzed, and the muscle activity intensity of the zygomatic major muscle area, the corrugator supercilii muscle area and the medial frontalis muscle area is recorded to generate the facial muscle activity intensity; Determine the video frame range corresponding to each word according to the timestamps of word appearance marked in the social language data; The occurrence time and duration of facial muscle activity intensity are determined within the video frame range and summarized into facial muscle response data.
6. The method for constructing an autism spectrum disorder recognition model based on deep learning according to claim 1, characterized in that: The step S2 of identifying body posture features in a conversation video of an autistic patient based on social language data and recording body posture data includes: Extract each frame of image from the conversation video, perform limb detection on each frame of image, and locate the body contour of the patient; Measure the vertical angle of the trunk of the body contour and divide the vertical angle of the trunk into three intervals: 0-15 degrees is normal posture, 15-30 degrees is mild forward leaning, and more than 30 degrees is severe forward leaning; Measure the knee extension angle of the leg of the torso silhouette, and divide the knee angle into three intervals: 160-180 degrees is fully extended, 120-160 degrees is slightly bent, and less than 120 degrees is severely bent; Determine the video frame range corresponding to each word according to the timestamps of word appearance marked in the social language data; The vertical angle of the trunk and the extension angle of the knee joint are recorded within the video frame to form limb posture data.
7. The method for constructing an autism spectrum disorder recognition model based on deep learning according to claim 1, characterized in that: The communication disorder assessment based on the gaze duration and facial muscle reaction data in step S3 includes: Determine the start time and end time of the gaze duration; determine the start time and end time of the facial muscle reaction data; The duration of gaze fixation was compared with the start and end time of the facial muscle reaction data; if the difference between the start time of gaze fixation and the start time of facial muscle reaction was less than 1 second, it was marked as a synchronous reaction; if the difference between the start time of gaze fixation and the start time of facial muscle reaction was more than 1 second, it was marked as an asynchronous reaction; The ratio of gaze duration to facial muscle reaction duration was calculated; if the ratio was less than 0.5, it was considered a low ratio; if the ratio was between 0.5 and 1.5, it was considered a medium ratio; if the ratio was greater than 1.5, it was considered a high ratio; Communication disorders are assessed based on the synchronization of eye gaze and facial muscle response and the ratio range; if the response is synchronous and the ratio is low, it is marked as a slow communication mode; if the response is synchronous and the ratio is high, it is marked as a rapid communication mode; if the response is asynchronous and the ratio is low, it is marked as an incoordinated communication mode; if the response is asynchronous and the ratio is high, it is marked as a disconnected communication mode; The above communication disorder assessment results are summarized into a communication disorder model.
8. The method for constructing an autism spectrum disorder recognition model based on deep learning according to claim 1, characterized in that: The social interaction disorder assessment based on gaze duration and body posture data in step S3 includes: Determine the start time and end time of the gaze duration; determine the start time and end time of the facial muscle reaction data; set the gaze duration of more than 3 seconds as long gaze, and the gaze duration of less than 1 second as short gaze; Extract the angle between the trunk and the vertical direction and the extension angle of the knee joint from the limb posture data; set the change of the vertical angle of the trunk to 15-30 degrees as a small change of the trunk angle, and the change of the vertical angle of the trunk to more than 30 degrees as a large change of the trunk angle; The duration of gaze fixation was compared with the changes in the trunk angle and knee extension angle. If the duration of gaze fixation was long and the trunk angle had a small change, it was marked as a coordinated social interaction mode; if the duration of gaze fixation was short and the trunk angle had a large change, it was marked as a discontinuous social interaction mode. By integrating the social interaction coordination model and the social interaction disruption model, we obtain the social interaction disorder model.
9. The method for constructing an autism spectrum disorder recognition model based on deep learning according to claim 1, characterized in that: The step S4 of matching the social interaction disorder pattern and the communication disorder pattern with the autism spectrum disorder history library includes: Extract the pattern occurrence time, pattern duration and pattern frequency of the social interaction disorder pattern and record them as the social interaction disorder characteristic data; Extract the pattern occurrence time, pattern duration and pattern frequency of the communication disorder pattern and record them as communication disorder characteristic data; The social interaction disorder data and communication disorder feature data are matched with the autism spectrum disorder history library respectively, and the similarity between each pattern and the known patterns in the autism spectrum disorder history library is determined and quantified. If the quantified similarity value exceeds 0.8, the match is considered successful. Interaction barrier matching data and communication barrier matching data are recorded based on similarity quantification values.
10. The method for constructing an autism spectrum disorder recognition model based on deep learning according to claim 1, characterized in that: The step S4 of constructing an autism spectrum disorder recognition model based on the interaction disorder matching data and the communication disorder matching data includes: Extracting high-dimensional feature vectors of interaction barrier matching data and communication barrier matching data; The high-dimensional feature vector is input into the deep learning model, and the MLP technology is used for feature fusion. In the deep learning model, the attention mechanism is introduced to perform weighted processing on different feature vectors to generate a comprehensive feature representation; Use the autoencoder to reduce the dimension of the comprehensive feature representation to obtain the feature dimension reduction representation data; Input the feature dimension reduction representation data into the classifier and use the decision tree for classification to generate preliminary recognition results; Use sliding window technology to smooth the preliminary recognition results to obtain processed recognition results; Perform secondary matching on the processed recognition results and the known patterns in the history library, and align the time series to obtain the secondary matching results; An autism spectrum disorder recognition model was constructed based on the secondary matching results.
Citation Information
Patent Citations
Early autism screening system based on human-computer interaction
CN114974572A
Method and device for establishing intellectual disorder diagnosis model for children with autism spectrum disorder
CN115482924A
Machine learning-based autism spectrum disorder early recognition method
CN115565690A
Autism auxiliary diagnosis method based on eyeball tracking technology
CN118866324A
Language evaluation method and device for autistic children and medium
CN119028594A