A Method for Constructing a Deep Learning-Based Autism Spectrum Disorder Identification Model
By monitoring the frequency of eye contact and analyzing social language in videos of conversations with autistic individuals, and combining facial and body posture data, an autism spectrum disorder identification model was constructed. This model solves the problem of accurately identifying reduced eye contact in existing technologies, improves identification efficiency and accuracy, and provides a basis for personalized intervention.
Patent Information
- Application Number
- CN202510472860.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-04-16
AI Technical Summary
Existing autism spectrum disorder identification models lack dynamic monitoring and real-time comparison when analyzing patients' eye gaze, making it difficult to accurately determine the state of reduced eye gaze and unable to accurately record the duration of eye gaze, resulting in insufficient identification efficiency and accuracy.
By acquiring videos of conversations between autistic individuals, continuously monitoring the frequency of eye contact, setting thresholds to determine reduced eye contact and recording the duration, and combining social language analysis to obtain facial and body posture data, a multi-dimensional assessment is conducted to construct an autism spectrum disorder identification model.
It enables accurate identification of social and communication impairments in autistic patients, improves identification efficiency and accuracy, provides a basis for personalized intervention, and avoids the limitations of single-indicator assessment.
Smart Images

Figure CN119993467B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of biomedical information technology, and particularly relates to a construction method of an autism spectrum disorder identification model based on deep learning. BACKGROUND
[0002] One of the core symptoms of autism spectrum disorder is social communication disorder, including difficulties in social interaction, and limitations in language and non-verbal communication abilities. Existing autism spectrum disorder identification model construction methods often involve multiple types of data, including behavior data, brain image data, psychological emotional state, external behavior pattern, and achievement development trajectory, and other multi-dimensional information. For example, some studies use data of head, torso and foot movement of children, and head movement features for judgment; some studies combine resting-state functional magnetic resonance imaging (rsfMRI) data and machine learning methods. However, when analyzing the gaze fixation of autism patients, the traditional method often only simply counts the fixation times or fixation time, lacks dynamic monitoring of the fixation rate and real-time comparison with the threshold, and is difficult to accurately judge the state of reduced gaze fixation, and cannot accurately record the gaze fixation duration. SUMMARY
[0003] Therefore, it is necessary to provide a construction method of an autism spectrum disorder identification model based on deep learning to solve at least one of the above technical problems.
[0004] To achieve the above-mentioned purpose, a construction method of an autism spectrum disorder identification model based on deep learning, the method comprising the following steps:
[0005] Step S1: obtaining a conversation video of an autism patient; continuously detecting the gaze fixation rate in the conversation video of the autism patient; when detecting that the gaze fixation rate is lower than the gaze fixation rate threshold, determining a reduced gaze fixation state and recording the gaze fixation duration;
[0006] Step S2: performing social language analysis on the conversation video of the autism patient to obtain social language data; identifying patient facial reaction features in the conversation video of the autism patient according to the social language data and recording facial muscle reaction data; identifying limb posture features in the conversation video of the autism patient according to the social language data and recording limb posture mode data;
[0007] Step S3: performing communication disorder evaluation based on the gaze fixation duration and the facial muscle reaction data to obtain a communication disorder mode; performing social interaction disorder evaluation based on the gaze fixation duration and the limb posture mode data to obtain a social interaction disorder mode;
[0008] Step S4: obtaining an autism spectrum disorder history database; performing feature matching of the social interaction disorder pattern and the communication disorder pattern with the autism spectrum disorder history database respectively to generate interaction disorder matching data and communication disorder matching data; and constructing an autism spectrum disorder recognition model based on the interaction disorder matching data and the communication disorder matching data.
[0009] The application can convert the eye gaze behavior of the patient into specific quantifiable data, i.e., the eye gaze rate and the eye gaze duration, by obtaining the conversation video of the autism patient and continuously monitoring the eye gaze rate in the video. This process provides an objective, accurate and operable data basis for subsequent evaluation, avoiding the deviation caused by subjective judgment; setting the eye gaze rate threshold, when detecting that the eye gaze rate is lower than the threshold, the patient can be timely and accurately determined to be in the eye gaze reduction state, and the eye gaze duration is recorded. This helps to capture the abnormal changes of the eye gaze behavior of the autism patient in social interaction, and provides a key basis for further analyzing the communication disorder. The social language data is obtained by performing social language analysis on the conversation video of the autism patient, and the facial muscle reaction data is recorded by identifying the facial reaction features of the patient based on the data, and the body posture mode data is recorded by identifying the body posture features. This process realizes the comprehensive collection of multi-dimensional data from language communication, facial expression to body movement, providing rich and complete information for subsequent comprehensive evaluation; the social language data can reflect the language expression and communication ability of the patient, the facial muscle reaction data can reflect the emotional response of the patient to the social situation, and the body posture mode data can reveal the body language performance of the patient in social interaction. The comprehensive analysis of the three types of data can comprehensively and objectively reflect the comprehensive ability of the autism patient in the social process, and provide strong support for accurately identifying and evaluating the degree of social disorder. The communication disorder mode is obtained by evaluating the communication disorder based on the eye gaze duration and the facial muscle reaction data, and the social interaction disorder mode is obtained by evaluating the social interaction disorder based on the eye gaze duration and the body posture mode data. This comprehensive evaluation method based on multi-dimensional data can accurately identify the specific disorder mode of the autism patient in communication and social interaction, and provide a clear direction and basis for subsequent personalized intervention; by integrating various data for evaluation, the limitations of single indicator evaluation are avoided. For example, relying only on the evaluation of language communication ability may ignore the patient's disorder in non-verbal social interaction, and combining facial expression and body movement data can more comprehensively reflect the social ability of the patient, and improve the accuracy and reliability of the evaluation. The autism spectrum disorder history library is obtained, and the social interaction disorder mode and the communication disorder mode are respectively matched with the history library to generate interaction disorder matching data and communication disorder matching data, and an autism spectrum disorder recognition model is constructed based on the interaction disorder matching data and the communication disorder matching data. Therefore, the application evaluates the communication disorder and social interaction disorder of the autism patient by data processing technology and deep learning technology, and constructs an autism spectrum disorder recognition model, thereby improving the recognition efficiency and accuracy of autism spectrum disorder.
[0010] Preferably, the obtaining the conversation video of the autism patient in step S1 comprises:
[0011] In the preset multiple different scenes, the autistic patient is recorded in the conversation video respectively, and the patient is in the center position of the lens during the recording process, and the horizontal angle between the lens and the eyes of the patient is kept within 10°;
[0012] The recorded video is preliminarily screened, the segment in which the picture is blurred for more than 5 seconds due to lens shaking is removed, and the segment in which the patient appears expression change and body movement in the video is retained;
[0013] The retained video segment is cut, and the video frames of the first 3 seconds and the last 2 seconds in which the patient does not enter the conversation state in the video are cut off;
[0014] The cut video segment is spliced according to the timestamp of the patient conversation, and the total length of the spliced video is not less than 1 minute, so as to obtain the conversation video of the autistic patient.
[0015] The patient in the lens center position and the lens and the patient's eyes horizontal angle within 10° can ensure that the video data collection angle is consistent in different scenes, reduce the interference caused by the difference in shooting angle on the subsequent analysis, and make the model more accurately capture the key information such as patient facial expression and body movement. The segment in which the picture is blurred for more than 5 seconds due to lens shaking can effectively avoid the interference of blurred picture on the recognition of patient behavior characteristics by the model, and ensure that the clarity of the retained video segment meets the analysis requirement, thereby improving the data quality. The segment in which the patient appears expression change and body movement can accurately focus on the key behavior performance of the patient in the conversation process, and these segments often contain important clues for identifying autism spectrum disorder. The video frames of the first 3 seconds and the last 2 seconds in which the patient does not enter the conversation state are cut off, which can remove the redundant information irrelevant to the conversation, make the video segment more focused on the real behavior performance of the patient in the conversation process, and further improve the pertinence and effectiveness of the data. According to the timestamp of the patient conversation, the effective conversation segments in different scenes can be integrated into a complete video, and the total length of the spliced video is ensured to be not less than 1 minute, so as to provide sufficient continuous conversation samples for the deep learning model.
[0016] Preferably, the step S1 includes the following steps:
[0017] The conversation video of the autistic patient is analyzed frame by frame, and the eye direction of the patient in each frame is determined;
[0018] In each frame, it is judged whether the eye direction of the patient forms an angle less than 30° with the lens direction, if the eye direction forms an angle less than 30° with the lens direction, the frame is recorded as an effective eye fixation;
[0019] With every 10 seconds as a statistical unit, the number of frames of effective eye gaze in each statistical unit is counted to obtain the eye gaze video rate of each statistical unit;
[0020] The entire video is divided according to the above statistical unit, and the eye gaze video rate of each statistical unit is calculated, and the frequency value of the statistical unit is recorded;
[0021] The frequency values of all statistical units are summarized, and the average eye gaze video rate of the entire video is calculated as the eye gaze video rate in the conversation video of the autism patient.
[0022] The application can accurately obtain the eye change information of the patient during the conversation by analyzing the direction of the patient's eyes frame by frame, and provide accurate basis for subsequent judgment of eye gaze; the judgment standard of less than 30° included angle between the direction of the eye and the direction of the lens is used as effective eye gaze, the eye gaze behavior is quantified as the number of frames that can be counted, and specific and operable quantitative indexes are provided for subsequent analysis; every 10 seconds is used as a statistical unit, the number of frames of effective eye gaze in each statistical unit is counted to obtain the eye gaze video rate of each statistical unit, which can more carefully observe the eye gaze of the patient in different time periods, and avoid the whole statistics from covering up the local characteristics; the frequency values of all statistical units are summarized, and the average eye gaze video rate of the entire video is calculated, which provides a comprehensive and quantitative target feature for the autism spectrum disorder identification model.
[0023] Preferably, the communication disorder evaluation based on the eye gaze duration and the facial muscle reaction data in step S3 comprises:
[0024] The start time and the end time of the eye gaze of the eye gaze duration are determined, and the start time and the end time of the facial muscle reaction data are determined;
[0025] The start time and the end time of the eye gaze of the eye gaze duration are compared with the start time and the end time of the facial muscle reaction data; if the start time of the eye gaze and the start time of the facial muscle reaction differ by no more than 1 second, it is marked as a synchronous reaction; if the start time of the eye gaze and the start time of the facial muscle reaction differ by more than 1 second, it is marked as an asynchronous reaction;
[0026] The ratio of the eye gaze duration and the facial muscle reaction duration is calculated; if the ratio is less than 0.5, it is determined as a low ratio; if the ratio is 0.5 to 1.5, it is determined as a medium ratio; if the ratio is greater than 1.5, it is determined as a high ratio;
[0027] Communication impairment is assessed based on the synchronicity and ratio range of eye contact and facial muscle responses; if the responses are synchronous and the ratio is low: it is marked as a slow communication pattern; if the responses are synchronous and the ratio is high: it is marked as a hasty communication pattern; if the responses are asynchronous and the ratio is low: it is marked as a disharmony communication pattern; if the responses are asynchronous and the ratio is high: it is marked as a disconnected communication pattern.
[0028] The above assessment results of communication barriers are summarized into a communication barrier model.
[0029] This invention can accurately quantify the temporal characteristics of eye gaze and facial muscle response by determining the start and end times of eye gaze and the start and end times of facial muscle response data. By comparing the start time of eye gaze with the start time of facial muscle response, if the time difference is no more than 1 second, it is marked as a synchronous response; if it is more than 1 second, it is marked as an asynchronous response. This quantitative method can clearly distinguish patients' response patterns in social interactions, providing a clear classification standard for identifying communication disorders. By calculating the ratio of eye fixation duration to facial muscle response duration and classifying it into low ratio (less than 0.5), medium ratio (0.5 to 1.5), and high ratio (greater than 1.5), it can further refine the patients' communication behavior characteristics. Based on the synchronicity of eye fixation and facial muscle response and the ratio range, the communication disorder assessment results are divided into communication delay mode (synchronous response and low ratio), communication haste mode (synchronous response and high ratio), communication incoordination mode (asynchronous response and low ratio), and communication interruption mode (asynchronous response and high ratio). The above communication disorder assessment results are summarized into communication disorder patterns, providing a comprehensive and quantitative target feature for the autism spectrum disorder identification model.
[0030] Preferably, the social interaction disorder assessment based on eye gaze duration and body posture data in step S3 includes:
[0031] Determine the start and end times of eye fixation duration; determine the start and end times of facial muscle response data; define eye fixation duration exceeding 3 seconds as long fixation and eye fixation duration less than 1 second as short fixation.
[0032] Extract the angle between the torso and the vertical direction and the extension angle of the knee joint from the limb posture data; define the change of the torso's vertical angle as small change of 15-30 degrees, and the change of the torso's vertical angle as large change of more than 30 degrees.
[0033] Compare the duration of eye fixation with changes in trunk angle and knee extension angle. If the duration of eye fixation is long and the trunk angle changes little, it is marked as a social interaction coordination pattern; if the duration of eye fixation is short and the trunk angle changes greatly, it is marked as a social interaction disruption pattern.
[0034] By integrating the social interaction coordination model and the social interaction disruption model, a social interaction disorder model is obtained.
[0035] This invention, by determining the start and end times of eye contact, can accurately calculate the duration of eye contact, providing specific and operable quantitative indicators for subsequent analysis; by determining the start and end times of facial muscle response data, it can capture the patient's social interaction behavior characteristics from multiple dimensions; by defining eye contact duration exceeding 3 seconds as long eye contact and less than 1 second as short eye contact, it categorizes eye contact behavior, providing clear classification criteria for subsequent analysis; by extracting the angle between the torso and the vertical direction and the knee extension angle from limb posture data, it can obtain the patient's limb behavior characteristics in social interaction; and by defining the torso's vertical angle variation as 15-30 degrees as the torso angle... Small changes, exceeding 30 degrees, are considered large changes in trunk angle. These changes in limb movement are quantified and categorized. Comparing the duration of eye fixation with changes in trunk angle and knee extension angle allows for the correlation between eye fixation and limb movement. If the eye fixation duration is long and the trunk angle changes little, it is marked as a social interaction coordination pattern; if the eye fixation duration is short and the trunk angle changes greatly, it is marked as a social interaction disruption pattern, identifying different patterns in the patient's social interactions. Integrating the social interaction coordination and disruption patterns yields a social interaction impairment pattern, providing a comprehensive and quantifiable target feature for the autism spectrum disorder identification model.
[0036] Preferably, step S4, which involves matching the social interaction impairment pattern and the communication impairment pattern with the autism spectrum disorder history database, includes:
[0037] Extract the occurrence time, duration, and frequency of social interaction disorder patterns and record them as social interaction disorder feature data;
[0038] Extract the occurrence time, duration, and frequency of communication barrier patterns and record them as communication barrier characteristic data;
[0039] The data on social interaction impairment and communication impairment features were matched with the historical database of autism spectrum disorder. The similarity between each pattern and the known patterns in the historical database of autism spectrum disorder was determined and the similarity was quantified. If the similarity quantification value exceeded 0.8, the match was considered successful.
[0040] Interaction barrier matching data and communication barrier matching data are recorded based on similarity metric values.
[0041] This invention extracts the occurrence time, duration, and frequency of social interaction disorder patterns and records them as social interaction disorder feature data. This process accurately quantifies the characteristics of social interaction disorders, providing accurate input for subsequent feature matching. It also extracts the occurrence time, duration, and frequency of communication disorder patterns and records them as communication disorder feature data. This process accurately quantifies the characteristics of communication disorders, further enriching the model's input data. The social interaction disorder feature data and communication disorder feature data are then matched against a historical database of autism spectrum disorders to determine the similarity between each pattern and known patterns in the database, and a similarity metric is performed. If the similarity metric value exceeds 0.8, a successful match is considered, effectively assessing the similarity between features and known patterns and improving recognition accuracy. Based on the similarity metric value, the social interaction disorder matching data and communication disorder matching data are recorded. This process provides labeled data for model training, further optimizing the model's recognition capabilities.
[0042] Preferably, the construction of the autism spectrum disorder identification model based on interaction impairment matching data and communication impairment matching data in step S4 includes:
[0043] Extract high-dimensional feature vectors from interaction barrier matching data and communication barrier matching data;
[0044] High-dimensional feature vectors are input into a deep learning model, and feature fusion is performed using MLP technology. An attention mechanism is introduced into the deep learning model to weight different feature vectors in order to generate a comprehensive feature representation.
[0045] An autoencoder is used to reduce the dimensionality of the comprehensive feature representation to obtain the feature dimensionality-reduced representation data.
[0046] The data is represented by reduced-dimensionality features and input into a classifier, and then classified using a decision tree to generate preliminary recognition results.
[0047] The initial recognition results are smoothed using a sliding window technique to obtain the processed recognition results.
[0048] The processed recognition results are then matched a second time with known patterns in the historical database, and time series alignment is performed to obtain the second matching results.
[0049] A model for identifying autism spectrum disorder was constructed based on the results of the second-order matching.
[0050] This invention inputs high-dimensional feature vectors into a deep learning model, uses MLP technology for feature fusion, and introduces an attention mechanism to weight different feature vectors, generating a comprehensive feature representation. This process effectively integrates multi-source feature information and highlights important features. An autoencoder is then used to reduce the dimensionality of the comprehensive feature representation, resulting in a dimensionality-reduced feature representation. The autoencoder removes redundant information, retains key features, improves the model's computational efficiency and generalization ability, and reduces the risk of overfitting. The dimensionality-reduced feature representation is then input into a classifier, and a decision tree is used for classification, generating preliminary recognition results. The decision tree classifier provides clear decision logic, quickly distinguishes different categories, and provides a foundation for subsequent processing. A sliding window technique is used to smooth the preliminary recognition results, resulting in a processed recognition result. The sliding window technique smooths noise and abrupt changes in the recognition results, enhancing the stability and continuity of the results and improving the reliability of the recognition results. The processed recognition result is then matched a second time with known patterns in a historical database and aligned with time series data to obtain a secondary matching result. Secondary matching further verifies the accuracy of the identification results, while time series alignment ensures consistency between the results and known patterns in the time dimension, improving the model's recognition accuracy. An autism spectrum disorder identification model is constructed based on the secondary matching results. Through multi-step feature processing and result optimization, this model can more accurately identify autism spectrum disorder characteristics, improving the accuracy and reliability of identification. Attached Figure Description
[0051] Fig. 1 A flowchart illustrating the steps involved in constructing a deep learning-based autism spectrum disorder identification model.
[0052] Fig. 2 A detailed flowchart illustrating the implementation steps of obtaining the conversational video of an autistic patient in step S1;
[0053] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0054] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0055] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.
[0056] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0057] To achieve the above objectives, please refer to Figs. 1-2 A method for constructing an autism spectrum disorder identification model based on deep learning, the method comprising the following steps:
[0058] Step S1: Acquire a video of conversations between autistic patients; continuously measure the frequency of eye contact in the conversation videos of autistic patients; when the frequency of eye contact is detected to be lower than the threshold, determine it as a state of reduced eye contact and record the duration of eye contact.
[0059] Step S2: Perform social language analysis on the conversation videos of autistic patients to obtain social language data; identify the facial reaction features of autistic patients in the conversation videos based on the social language data and record facial muscle reaction data; identify the body posture features in the conversation videos of autistic patients based on the social language data and record body posture data.
[0060] Step S3: Based on eye gaze duration and facial muscle response data, conduct a communication impairment assessment to obtain a communication impairment pattern; based on eye gaze duration and body posture data, conduct a social interaction impairment assessment to obtain a social interaction impairment pattern.
[0061] Step S4: Obtain the historical database of autism spectrum disorders; perform feature matching between the social interaction disorder pattern and the communication disorder pattern and the historical database of autism spectrum disorders to generate interaction disorder matching data and communication disorder matching data; construct an autism spectrum disorder identification model based on the interaction disorder matching data and the communication disorder matching data.
[0062] This invention acquires videos of conversations between autistic individuals and continuously monitors the frequency of eye contact within them. This transforms the patients' eye contact behavior into specific, quantifiable data, namely, eye contact frequency and duration. This process provides an objective, accurate, and operable data foundation for subsequent assessments, avoiding biases caused by subjective judgment. By setting a threshold for eye contact frequency, when the detected eye contact frequency falls below this threshold, it is possible to promptly and accurately determine that the patient is in a state of reduced eye contact, and record the duration of their eye contact. This helps to capture abnormal changes in the eye contact behavior of autistic individuals during social interactions, providing crucial evidence for further analysis of their communication impairments. Sociolinguistic analysis is performed on the conversation videos of autistic individuals to obtain sociolinguistic data. Based on this data, facial reaction features are identified and recorded as facial muscle reaction data, and limb posture features are identified and recorded as limb posture pattern data. This process achieves comprehensive data collection across multiple dimensions, from verbal communication and facial expressions to body movements, providing rich and complete information for subsequent comprehensive assessments. Socio-language data reflects the patient's verbal expression and communication abilities, facial muscle response data reflects the patient's emotional responses to social situations, and body posture data reveals the patient's body language performance in social interactions. The comprehensive analysis of these three types of data can comprehensively and objectively reflect the overall ability of autistic patients in social processes, providing strong support for accurately identifying and assessing the degree of their social impairments. Communication impairment assessments based on eye gaze duration and facial muscle response data yield communication impairment patterns; similarly, social interaction impairment assessments based on eye gaze duration and body posture data yield social interaction impairment patterns. This comprehensive assessment method based on multi-dimensional data can accurately identify specific impairment patterns in communication and social interaction for autistic patients, providing a clear direction and basis for subsequent personalized interventions. By integrating multiple data sources for assessment, the limitations of single-indicator assessments are avoided. For example, assessments relying solely on verbal communication abilities may overlook nonverbal social impairments, while combining facial expressions and body movements data can more comprehensively reflect the patient's social abilities, improving the accuracy and reliability of the assessment. This invention acquires a historical database of autism spectrum disorders and performs feature matching on social interaction and communication impairment patterns against this database, generating interaction impairment matching data and communication impairment matching data, respectively. An autism spectrum disorder identification model is then constructed based on this data. Therefore, this invention utilizes data processing and deep learning technologies to assess communication and social interaction impairments in autistic individuals and constructs an autism spectrum disorder identification model, thereby improving the efficiency and accuracy of autism spectrum disorder identification.
[0063] In this embodiment of the invention, reference Fig. 1The diagram shown illustrates the steps of constructing a deep learning-based autism spectrum disorder identification model according to the present invention. In this example, the method for constructing the deep learning-based autism spectrum disorder identification model includes the following steps:
[0064] Step S1: Acquire a video of conversations between autistic patients; continuously measure the frequency of eye contact in the conversation videos of autistic patients; when the frequency of eye contact is detected to be lower than the threshold, determine it as a state of reduced eye contact and record the duration of eye contact.
[0065] In this embodiment of the invention, a high-resolution camera is selected, ensuring a resolution of 1920×1080 pixels or higher, with a frame rate set to 30 frames per second to guarantee video clarity and smoothness, accurately capturing the patient's various subtle movements and expressions. The camera is fixed on a tripod to ensure stability. The camera's shooting angle should be directly facing the patient, ensuring the patient's face and upper body occupy the main position in the frame for easy subsequent analysis. The distance between the camera and the patient should be adjusted according to the actual scene, generally maintained at around 1.5-2 meters, to ensure the patient's entire body is included in the image while facial details are clearly discernible. A well-lit and evenly spaced shooting environment is chosen, avoiding direct sunlight or shadows to ensure moderate brightness and contrast in the video image. If natural light is insufficient, auxiliary lighting equipment, such as evenly distributed LED lights, can be used to ensure soft light without significant reflections. Before recording begins, the camera is warmed up and tested to ensure normal operation and a clear and stable image. The image clarity, color reproduction, and any distortion issues are checked using the camera's built-in preview function or video monitoring software connected to a computer. Start the video recording function to record the entire process of the patient's conversation with others. During recording, monitor the video in real time to ensure there are no obstructions, shakiness, or other interference factors affecting video quality. If any problems are found, pause recording immediately and readjust the equipment. Store the recorded video file on a storage device with sufficient capacity, such as a large-capacity hard drive or solid-state drive, to ensure the integrity and security of the video data. The video file format should be a common format, such as MP4 or AVI, for subsequent processing and analysis. Perform preliminary processing on the video, including trimming irrelevant beginnings and endings to ensure the video content only includes the complete process of the patient's conversation. Use video editing software (such as Adobe Premiere Pro) to perform trimming operations, and check whether parameters such as the video frame rate and resolution meet the requirements. Make appropriate adjustments if necessary. Use the TobiiTX300 eye tracker for eye tracking. Install and set up the TobiiExperience software, ensuring that the software version is compatible with the eye tracker. Before calibration, ensure that the relative position of the eye tracker and the screen is correct, aligning the two white lines above the eye tracker with the lines on the screen. Next, select the eye to be monitored and perform the calibration operation according to the software prompts. During calibration, the patient fixates on multiple fixed points on the screen, sequentially following the small circles on the screen until they break, completing the calibration. Throughout the calibration process and subsequent monitoring, the patient should keep their body position as still as possible, and multiple calibration targets should be displayed to cover the participant's entire field of vision. Verification procedures must be performed at the beginning and end of the experiment, as well as after the eye tracker is moved, to ensure the accuracy of the calibration. The eye tracker records the patient's gaze data in real time, including the coordinates of the fixation point and the duration of fixation.The system calculates gaze frequency, or the number of times a patient's eyes focus per unit of time, using a preset algorithm. The specific calculation method is as follows: within a set time window (e.g., 1 second), the number of effective gazes is counted, and then this number is divided by the length of the time window to obtain the gaze frequency. A gaze frequency threshold is set, determined based on the statistical difference between the average gaze frequency of the general population and the gaze frequency of autistic patients. For example, through statistical analysis of a large sample of data, the average gaze frequency of the general population is 5 times per second, while the gaze frequency of autistic patients is usually lower than this value by a certain percentage (e.g., 20%), so the threshold is set to 4 times per second. When the gaze frequency is detected to be lower than the set threshold, the system determines that the patient is in a state of reduced gaze and immediately starts a timing function to record the duration of gaze reduction. The timer accurately records the time period from when the gaze frequency falls below the threshold until it returns to normal, in milliseconds; this time period is the duration of reduced gaze.
[0066] Step S2: Perform social language analysis on the conversation videos of autistic patients to obtain social language data; identify the facial reaction features of autistic patients in the conversation videos based on the social language data and record facial muscle reaction data; identify the body posture features in the conversation videos of autistic patients based on the social language data and record body posture data.
[0067] In this embodiment of the invention, professional audio processing software (such as Audacity) is used to extract the speech track from the video file, ensuring that the extracted speech is clear and free of noise. The extracted speech file is imported into a speech recognition tool (such as IBM Watson Speech to Text or Microsoft Azure Speech Service), the language model is set to Mandarin or English (selected according to the patient's language), and the transcription function is activated to convert the speech into text format and save it as a TXT file. The transcribed text is manually proofread to ensure that the text accuracy is above 95%. A text analysis tool (such as NVivo) is used to encode the transcribed text, marking key social language features, such as greetings, questions, responses, and repetitive language. The frequency of various language features in the text is statistically analyzed, such as the number of times greetings are used and questions are asked, generating a social language data table, including indicators such as lexical richness (total number of different words) and turn count (number of dialogue exchanges). Video analysis software (such as OpenCV) is used to process the video frame by frame, loading a pre-trained face detection model (such as HaarCascadeClassifier) to locate the patient's facial region in each frame. Set detection parameters: minimum neighbor count is 5, scaling factor is 1.1, ensuring face detection accuracy is no less than 95%. Track the detected facial regions and record changes in facial position within the video. Within the detected facial regions, use a facial expression analysis tool (such as AffectivaSDK) to identify facial expressions, recording the occurrence time and duration of each expression (e.g., happiness, sadness, surprise). Generate a facial muscle response data file, recording the expression category and corresponding duration for each frame in CSV format, including timestamp, expression category, and duration fields. Use a pose estimation tool (such as OpenPose) to process the video, extracting keypoints of the patient's limbs in each frame, including the head, shoulders, elbows, wrists, knees, and ankles. Set OpenPose parameters: detection threshold is 0.5, keypoint connectivity distance is 50 pixels, ensuring keypoint detection accuracy is no less than 90%. Export the detected keypoint data in JSON format, containing keypoint coordinate information for each frame. The Python script parses the JSON file to determine the angles and distances between key points and identify body posture features, such as crossed arms, hands behind the head, and leaning forward. It records the occurrence time and duration of each posture, generating a body posture data file in CSV format, containing fields such as timestamp, posture category, and duration.
[0068] Step S3: Based on eye gaze duration and facial muscle response data, conduct a communication impairment assessment to obtain a communication impairment pattern; based on eye gaze duration and body posture data, conduct a social interaction impairment assessment to obtain a social interaction impairment pattern.
[0069] In this embodiment of the invention, eye gaze duration data and facial muscle response data are aligned according to timestamps to ensure that the eye gaze duration at each moment matches the corresponding facial muscle response data. For example, if the eye gaze duration is recorded as from the 5th second to the 10th second, the facial muscle response data extracts the expression category and its duration within the same time period. For each time period, the total duration of eye gaze and the category and duration of facial expressions are statistically analyzed. For example, the proportion of eye gaze duration to the total duration of a given time period, and the number of facial expression transitions are determined. A threshold for eye gaze duration is set; for example, if the eye gaze duration is less than 30% of the time period, it is marked as reduced eye gaze. A threshold for facial expression transition frequency is set; for example, if the number of facial expression transitions is less than 5 times per minute, it is marked as low expression transition frequency. Based on the above two marking results, it is determined whether a communication barrier pattern has occurred. If both reduced eye gaze and low expression transition frequency are met simultaneously, it is determined to be a communication barrier pattern, and the start time and duration of that time period are recorded. The evaluation results are recorded in tabular form, including timestamps, eye gaze duration, facial expression category and its duration, and whether it is a communication barrier pattern. For example, a row in the table might display: timestamp (5-10 seconds), eye gaze duration (3 seconds), facial expression (happy, lasting 2 seconds; sad, lasting 1 second), and communication barrier pattern (yes). Eye gaze duration data is aligned with body posture data according to the timestamp to ensure a match between the eye gaze duration and the corresponding body posture data at each moment. For example, if the eye gaze duration is recorded from the 15th to the 20th second, the body posture data extracts the body posture categories and their durations within the same time period. For each time period, the total duration of eye gaze, as well as the category and duration of body posture, are calculated. For example, the proportion of eye gaze duration within a given time period and the number of body posture transitions are determined. A threshold is set for eye gaze duration; for example, if the eye gaze duration is less than 30% of the time period, it is marked as reduced eye gaze. A threshold is set for body posture transition frequency; for example, if the number of body posture transitions is less than 3 times per minute, it is marked as low body posture transition frequency. Based on these two marking results, it is determined whether a social interaction barrier pattern has occurred. If both reduced eye contact and low frequency of body posture changes are met, it is identified as a social interaction disorder pattern, and the start time and duration of this period are recorded. The assessment results are recorded in tabular form, including timestamps, duration of eye contact, body posture category and its duration, and whether it is a social interaction disorder pattern. For example, one row in the table may show: timestamp (15-20 seconds), duration of eye contact (4 seconds), body posture (arms crossed, lasting 3 seconds; hands behind head, lasting 1 second), social interaction disorder pattern (yes).
[0070] Step S4: Obtain the historical database of autism spectrum disorders; perform feature matching between the social interaction disorder pattern and the communication disorder pattern and the historical database of autism spectrum disorders to generate interaction disorder matching data and communication disorder matching data; construct an autism spectrum disorder identification model based on the interaction disorder matching data and the communication disorder matching data.
[0071] In this embodiment of the invention, historical data is obtained from the open-source Autism Spectrum Disorder Knowledge Base (AsdKB). AsdKB is a Chinese knowledge base for early screening of autism spectrum disorder, covering multiple information sources, including ICD-10 and DSM-5. Feature data related to social interaction and communication disorders are extracted; patterns of social interaction and communication disorders are matched with feature data in the historical database of autism spectrum disorder. Specifically, for social interaction disorder patterns, key features are extracted, such as low frequency of body posture transitions and short duration of eye fixation, and compared with symptom instances related to social interaction disorders recorded in AsdKB. For communication disorder patterns, key features are extracted, such as low frequency of facial expression transitions and reduced eye fixation, and compared with symptom instances related to communication disorders recorded in AsdKB. Cosine similarity is used to quantify the similarity between features. Specifically, the cosine similarity between the current pattern feature vector and the feature vector in the historical database is determined; a similarity higher than a set threshold (e.g., 0.8) is considered a successful match. Record the matching results to generate interaction barrier matching data and communication barrier matching data, including the matched feature names and similarity scores. Based on the interaction barrier matching data and communication barrier matching data, build a recognition model. Select the XGBoost algorithm, which can handle imbalanced data and provide high-accuracy classification results. Preprocess the data, including handling imbalanced data (e.g., by setting weights) and splitting the data into training and validation sets (e.g., in a 7:3 ratio). Train the XGBoost model using the training set to optimize the model parameters. General parameters are set as follows: booster: select gbtree as the booster; silent: set to 0, print the running message; nthread: set to the maximum number of available threads. Tree boosting is set as follows: eta: set to 0.3, control the learning rate; gamma: set to 0, control the minimum loss reduction when splitting nodes; max_depth: set to 6, control the maximum tree depth; min_child_weight: set to 1, control the minimum weight of leaf nodes; subsample: set to 1, control the sample sampling ratio; colsample_bytree: set to 1, control the feature sampling ratio of each tree; lambda: set to 1, control the L2 regularization term. alpha: Set to 0 to control the L1 regularization term. Learning task parameters are set as follows: objective: Set to binary: logistic, for binary classification problems, outputting probabilities. eval_metric: Set to AUC, using AUC as the evaluation metric for the validation set.
[0072] As an example of the present invention, reference is made to Fig. 2 As shown, obtaining the conversation video of the autistic patient in step S1 includes:
[0073] S11: Record conversation videos of autistic patients in several different preset scenarios. During the recording, the patient is in the center of the camera and the horizontal angle between the camera and the patient's eyes is kept within 10°.
[0074] S12: Perform preliminary screening of the recorded videos, removing segments with blurred images for more than 5 seconds due to camera shake, while retaining segments in the videos in which the patient shows changes in facial expressions and body movements.
[0075] S13: Trim the retained video clips, removing the first 3 seconds and the last 2 seconds of video frames in the video where the patient is not in a conversational state;
[0076] S14: Splice the cropped video segments according to the timestamps of the patient's conversation. The total length of the spliced video should be no less than 1 minute to obtain the conversation video of the autistic patient.
[0077] In this embodiment of the invention, multiple different social scenarios are designed, such as home environment, school environment, and social gathering environment. Each scenario should have a clear background and interactive objects to simulate the social behavior of autistic patients in different environments. A high-definition camera is used for recording, ensuring the camera resolution is no less than 1920×1080 pixels and the frame rate is set to 30 frames per second. The position and angle of the camera are adjusted, and an angle measuring tool (such as an angle measuring instrument) is used to ensure that the horizontal angle between the lens and the patient's eyes is kept within 10°. A ruler is used to ensure that the patient is always centered in the lens; this can be assisted by marking the center point of the lens and the patient's position. Conversational videos of the patient are recorded separately in each preset scenario. During recording, the video screen is monitored in real time to ensure that the patient is always centered in the lens. After recording, the video file is saved, ensuring the file format is MP4 or AVI for subsequent processing. A video analysis tool (such as Video-Analyzer) is used to perform preliminary screening of the recorded videos. Using the frame stability analysis function of Video-Analyzer, the stability of the video screen is checked frame by frame, and segments with blurred images exceeding 5 seconds due to camera shake are removed. Simultaneously, the facial expression recognition function of Video-Analyzer was used to mark segments in the video where the patient exhibited facial expression changes. A body movement detection tool (such as OpenPose) was used to mark segments in the video where the patient exhibited body movements, and these marked facial expression changes and body movement segments were preserved. Video editing software (such as Adobe Premiere Pro) was used to trim the preserved video segments, precisely trimming the first 3 seconds and last 2 seconds of video frames where the patient was not engaged in conversation, based on the video's timestamps. The trimmed video segments were then spliced together according to the patient's conversation timestamps, ensuring the total length of the spliced video was at least 1 minute. The final spliced video file was saved, ensuring the file format was MP4 or AVI.
[0078] Preferably, the eye contact frequency in the video of a patient with persistent autism in step S1 includes:
[0079] Frame-by-frame analysis of videos of conversations by autistic patients was performed to determine the direction of the patient's eyes in each frame.
[0080] In each frame, it is determined whether the patient's eye direction forms an angle of less than 30° with the lens direction. If the eye direction forms an angle of less than 30° with the lens direction, then the frame is recorded as a valid gaze fixation.
[0081] The number of frames with effective eye fixation within each statistical unit is counted, with each unit consisting of 10 seconds, to obtain the eye fixation frequency for each statistical unit.
[0082] The entire video is divided into the above statistical units, and the gaze frequency of each statistical unit is calculated and recorded as the frequency value of the statistical unit.
[0083] The frequency values of all statistical units are summarized, and the average eye fixation frequency of the entire video is calculated as the eye fixation frequency in the video of the autistic patient talking.
[0084] In this embodiment of the invention, a video file is loaded using the OpenCV library, and the video content is read frame by frame. The frame reading interval is set to 1 frame to ensure frame-by-frame analysis of the video. Each frame image is converted to grayscale to reduce interference from color information and improve the efficiency of subsequent processing. Gaussian blur is applied to the grayscale image to remove noise and smooth the image. The kernel size of the Gaussian blur is set to 5×5, and the standard deviation is set to 1.5 to ensure a smooth image. Eye detection is performed using the Haar feature cascade classifier in OpenCV. A pre-trained Haar cascade file (such as haarcascade_eye.xml) is loaded, and the eye region is located in each frame image. Geometric analysis is performed on the detected eye region to determine the center point of the eye. Specifically, the average of the coordinates of the upper left and lower right corners of the eye rectangle is taken to calculate the center coordinates of the eye. Assuming the camera direction is horizontal (i.e., parallel to the horizontal axis of the video), the angle between the eye center point and the horizontal axis of the video is calculated. Specifically, the horizontal and vertical distances between the eye center point and the video center point are calculated, and then the angle is calculated using the arctangent function. The calculated angles are converted to degrees for subsequent judgment. It is then determined whether the calculated angle is less than 30 degrees. If the angle is less than 30 degrees, the frame is marked as a valid gaze frame. To ensure accuracy, the angle is calculated separately for each detected eye region, and the average of the angles between the two eyes is used as the final judgment criterion. If the average angle between the two eyes is less than 30 degrees, the frame is marked as a valid gaze frame. Based on the video frame rate (assuming 30 frames / second), each 10 seconds contains 300 frames. The total number of video frames is divided by 300 to obtain the total number of statistical units. The frames within each statistical unit are traversed, and the number of frames marked as valid gazes is counted. The gaze frequency of each statistical unit is calculated, which is the number of valid gaze frames divided by the total number of frames in that statistical unit. For example, if 150 frames in a statistical unit are marked as valid gazes, and the total number of frames is 300, then the gaze frequency of that statistical unit is 150 divided by 300, resulting in 0.5, or 50%. The eye-fixing frequencies of all statistical units are summed to calculate the average eye-fixing frequency for the entire video. Specifically, the eye-fixing frequencies of all statistical units are added together, and then divided by the total number of statistical units. For example, if there are 10 statistical units, and the eye-fixing frequencies of each unit are 0.4, 0.5, 0.6, etc., adding these frequency values together and dividing by 10 gives the average eye-fixing frequency for the entire video.
[0085] Of particular importance, the step of determining a reduced gaze state and recording the duration of gaze when the detected gaze frequency is lower than the gaze frequency threshold includes:
[0086] Time series analysis was performed on the gaze frequency data, dividing the gaze frequency data into multiple time windows of preset length according to time sequence, with the gaze frequency data in each time window serving as an independent analysis unit.
[0087] Within each time window, the gaze frequency data is smoothed to obtain a smoothed gaze frequency curve.
[0088] Based on the smoothed gaze frequency curve, the rate of change of gaze frequency within each time window is calculated, and the trend of gaze frequency is determined by comparing the rate of change of gaze frequency in adjacent time windows.
[0089] When the gaze frequency shows a continuous downward trend, the time point when the gaze frequency is lower than the gaze frequency threshold is determined, and this time point is taken as the starting point of the gaze reduction state, and the gaze duration is recorded.
[0090] During the recording of gaze duration, changes in gaze frequency are monitored in real time; if the gaze frequency exceeds the gaze frequency threshold again, the recording of gaze duration is stopped; the entire time from the start of recording to the end of recording is taken as the gaze duration.
[0091] In this embodiment of the invention, gaze frequency data is divided into multiple time windows of preset length according to time sequence, and the gaze frequency data within each time window is treated as an independent analysis unit. For example, assuming the preset time window length is 10 seconds and the video frame rate is 30 frames / second, each time window contains 300 frames of data. Within each time window, the gaze frequency data is smoothed to reduce noise and highlight trends. Exponential smoothing is used for smoothing, and the specific steps are as follows: Select a smoothing coefficient α, for example, α=0.3. The closer the α value is to 1, the more sensitive the smoothed data is to the most recent data point. Perform exponential smoothing on the data points within each time window, and calculate the formula: current smoothing value = α × current data point + (1-α) × previous smoothing value, to obtain the smoothed gaze frequency curve within each time window. Based on the smoothed gaze frequency curve, calculate the gaze frequency change rate within each time window. Specifically, calculate the difference in gaze frequency between two adjacent time points, and then divide it by the frequency value of the previous time point. By comparing the gaze frequency change rates of adjacent time windows, the gaze frequency trend is determined. If the rate of change is negative for multiple consecutive time windows, and the magnitude of the change gradually increases, it is determined that the gaze frequency is showing a continuous downward trend. When the gaze frequency shows a continuous downward trend, the time point when the gaze frequency falls below the gaze frequency threshold is determined. For example, if the gaze frequency threshold is set to 0.3, and the gaze frequency within a certain time window is lower than 0.3, then the starting time point of that time window is taken as the starting point of the gaze reduction state, and the gaze duration is recorded from that starting point. During the recording of the gaze duration, the change in gaze frequency is monitored in real time. If the gaze frequency rises above the gaze frequency threshold again, the recording of the gaze duration is stopped, and the entire time from the start of recording to the stop of recording is taken as the gaze duration.
[0092] Preferably, the sociolinguistic analysis of the conversation videos of autistic patients described in step S2 includes:
[0093] Audio signals were extracted from videos of conversations between individuals with autism, converted into text, and the pronunciation duration and intonation changes of each word were recorded.
[0094] The text is segmented into words, and each word is labeled with its timestamp in the audio.
[0095] Extract the pronunciation duration and intonation changes of each word to generate a speech feature vector corresponding to each word;
[0096] The frequency of each word in the text is counted, and the sentiment of each word is marked by combining the speech feature vector.
[0097] Identify interrogative and declarative sentences in the text based on sentiment tendency, count the number of interrogative and declarative sentences, and record the average intonation variation of each sentence type;
[0098] The order of word usage in the text is marked based on the average intonation variation, and the collocation relationship of each word is recorded to obtain social language data.
[0099] In this embodiment of the invention, audio processing software (such as Audacity) is used to separate audio signals from videos of conversations by autistic patients, ensuring clear audio without background noise. The separated audio signals are imported into a speech recognition tool (such as the Google Speech-to-Text API), the language model is set to Mandarin or English (selected according to the patient's language), and the transcription function is activated to convert speech into text. During speech recognition, audio analysis is enabled to record the pronunciation duration and intonation changes of each word. Specifically, pronunciation duration is recorded in milliseconds, and intonation changes are quantified by analyzing the frequency fluctuations of the audio, expressed as the amplitude of frequency change (Hz). Natural language processing tools (such as NLTK or Jieba) are used to segment the transcribed text into individual words. Each word is labeled with its timestamp in the audio. The timestamp is in seconds, accurate to two decimal places, indicating the start and end times of the word in the audio. The pronunciation duration and intonation changes of each word are extracted to generate a speech feature vector corresponding to each word. The speech feature vector comprises two dimensions: pronunciation duration (in milliseconds) and intonation variation amplitude (in Hz). The frequency of each word in the text is statistically analyzed, and combined with the speech feature vector, the sentiment tendency of each word is labeled. Sentiment tendency is determined by analyzing intonation variation amplitude: rising intonation generally indicates positive sentiment, falling intonation indicates negative sentiment, and stable intonation indicates neutral sentiment. Based on sentiment tendency, interrogative and declarative sentences in the text are identified. Specifically, interrogative sentences typically begin with interrogative words (such as "what" or "why") and rise in intonation at the end; declarative sentences have a relatively stable intonation. The number of interrogative and declarative sentences is counted, and the average intonation variation amplitude for each sentence type is recorded. The average intonation variation amplitude is determined by calculating the average intonation variation amplitude across all corresponding sentence types. The order of word usage in the text is labeled based on the average intonation variation amplitude, recording the collocation relationships of each word. Specifically, the intonation variation trends between words are analyzed to identify which words tend to appear in positive or negative contexts, and their collocation patterns are recorded. The above information is integrated to generate social language data, including word frequency, sentiment, intonation variation, and collocation, which is stored in tabular form for easy subsequent analysis.
[0100] Preferably, step S2, which involves identifying facial reaction features of autistic patients in conversation videos based on social language data and recording facial muscle reaction data, includes:
[0101] Each frame of the video of autistic patients talking is extracted, and facial detection is performed on each frame to locate the zygomaticus major muscle region, corrugator supercilii muscle region, and medial frontalis muscle region of the patient's face.
[0102] The facial contours in each frame of the image are analyzed, and the muscle activity intensity of the zygomaticus major, corrugator supercilii, and medial frontalis muscle regions is recorded to generate facial muscle activity intensity.
[0103] Based on the timestamps of word occurrences tagged in social language data, determine the range of video frames corresponding to each word;
[0104] Within the video frame range, the occurrence time and duration of facial muscle activity intensity are determined and summarized into facial muscle response data.
[0105] In this embodiment of the invention, the OpenCV function `cv2.VideoCapture` is used to load the video file. The `cap.read()` method is called repeatedly to read the video content frame by frame, ensuring that each frame is extracted. Each frame is saved as a separate image file in JPEG format and stored in a specified folder. The pre-trained Haar cascade file `haarcascade_frontalface_default.xml` is loaded using `cv2.CascadeClassifier`. Each frame is converted to grayscale using the `cv2.cvtColor` function with the parameter `cv2.COLOR_BGR2GRAY`. The `detectMultiScale` method is called for face detection, with parameters set to `scaleFactor=1.2`, `minNeighbors=3`, and `minSize=(32, 32)`. A rectangle is drawn on each frame to mark the detected facial regions. The pre-trained facial landmark detection model `shape_predictor_68_face_landmarks.dat` is loaded using Dlib's `shape_predictor`. For each detected facial region in each frame, the `predictor` method is called to detect 68 facial landmarks. Based on keypoint coordinates, two keypoints above the corners of the mouth (e.g., keypoints 48 and 54) are selected as reference points for the zygomaticus major muscle region. Keypoints below the inner side of the eyebrows (e.g., keypoints 21 and 22) are selected as reference points for the corrugator supercilii muscle region. A keypoint in the central part of the forehead (e.g., keypoint 27) is selected as a reference point for the medial frontalis muscle region. For the facial region of each frame, the cv2.findContours method is used to extract the contour. For each region (zygomaticus major, corrugator supercilii, medial frontalis muscle), the area and perimeter of the contour are calculated using the cv2.contourArea and cv2.arcLength functions. The change in contour area is used as a quantitative indicator of muscle activity intensity. For example, the difference between the contour area of the current frame and the previous frame is calculated; the larger the difference, the higher the muscle activity intensity. The timestamps of words in the social language data are converted into video frame numbers. Assuming a video frame rate of 30 frames / second, the starting frame number corresponding to a word with a timestamp of 2 seconds is 60 (2 seconds × 30 frames / second). Based on the pronunciation duration of words, the corresponding video frame range is determined. If the pronunciation duration of a word is 1 second, the corresponding frame range is 60 to 90 frames. Within the determined video frame range, the muscle activity intensity in the zygomaticus major, corrugator supercilii, and medial frontalis muscle regions is analyzed frame by frame. The starting frame number of the muscle activity intensity from low to high is recorded as the occurrence time, and the ending frame number from high to low is recorded as the end point of the duration. The occurrence time, duration, and activity intensity value of facial muscle activity intensity corresponding to each word are summarized to form a facial muscle response data table.The data table contains information such as words, occurrence time (frame number), duration (number of frames), and muscle activity intensity values.
[0106] Preferably, step S2, which involves identifying body posture features in conversation videos of autistic patients based on social language data and recording body posture data, includes:
[0107] Extract each frame of the conversation video and perform limb detection on each frame to locate the patient's body contour.
[0108] Measure the vertical angle of the torso in the body contour and divide the vertical angle of the torso into three intervals: 0-15 degrees is the normal posture, 15-30 degrees is the slight forward tilt, and more than 30 degrees is the severe forward tilt.
[0109] The knee extension angle of the legs is measured in the body contour. The knee angle is divided into three ranges: 160-180 degrees is full extension, 120-160 degrees is slight flexion, and less than 120 degrees is severe flexion.
[0110] Based on the timestamps of word occurrences tagged in social language data, determine the range of video frames corresponding to each word;
[0111] Record the vertical angle of the torso and the knee extension angle within the video frame range to form limb posture data.
[0112] In this embodiment of the invention, a video file of autistic patients conversing is loaded using video processing software. The video processing software is configured to read the video content frame by frame, ensuring that each frame is extracted. Each frame is saved as a separate image file in JPEG format and stored in a designated folder. The file names are frame_0001.jpg, frame_0002.jpg, etc., for subsequent processing. Each frame is processed using a limb detection tool to extract multiple key points of the human body, including the torso, limbs, etc. Based on the detected key points, the patient's body contour is determined. Key points include the shoulders, waist, and hips. Among the detected limb key points, the shoulder and hip key points are selected. The angle between the line connecting the shoulder and hip key points and the vertical direction is calculated. Specifically, the horizontal and vertical distances between these two points are calculated, and then the arctangent function is used to calculate the angle. The angle is divided into three intervals: 0-15 degrees for normal posture, 15-30 degrees for slight forward tilt, and more than 30 degrees for severe forward tilt. Record the angle value and corresponding posture classification of each frame in a data table, including frame number, angle value, and posture classification. Select the knee and ankle joints from the detected limb key points. Calculate the knee extension angle. Specifically, calculate the angle between the line connecting the knee and ankle key points and the vertical direction. Divide the angle into three intervals: 160-180 degrees for full extension, 120-160 degrees for slight bending, and less than 120 degrees for severe bending. Record the angle value and corresponding bending classification of each frame in a data table, including frame number, angle value, and bending classification. Load the social language data file, which contains words and their timestamps. Convert the timestamps of the words in the social language data into corresponding video frame numbers. Assuming a video frame rate of 30 frames / second, the starting frame number for a word with a timestamp of 2 seconds is 60 (2 seconds × 30 frames / second). Determine the corresponding video frame range based on the pronunciation duration of the word. For example, if a word's pronunciation duration is 1 second, the corresponding frame range is 60 to 90 frames. Within a defined video frame range, record the vertical angle of the torso and the knee extension angle for each frame. Summarize the recorded data to form a limb posture data table. The table includes information such as words, frame range, torso angle, and knee angle. Save the summarized limb posture data as a CSV file for subsequent analysis.
[0113] Of particular importance, the method of identifying body posture features in conversation videos of autistic patients based on social language data and recording body posture data also includes:
[0114] Extract emotional expression words and their corresponding timestamps from social language data, and categorize the dialogue content into positive and negative emotional tendencies; simultaneously analyze body postures in the video based on the timestamps of the social language data;
[0115] When the dialogue content is positive, analyze whether the patient's body posture shows openness; calculate the vertical distance between the center point of the shoulder and the center point of the waist. If the vertical distance gradually decreases within the time window, it is judged as leaning forward; calculate the vertical distance between the key points of the wrists and the horizontal line of the shoulders. If the key points of the wrists are below the horizontal line of the shoulders and remain stable within the time window, it is judged as the hands are placed naturally.
[0116] When the dialogue content is negative, analyze whether the target's body posture shows closed characteristics; calculate the horizontal distance between the key points of the wrists of both hands. If the horizontal distance gradually decreases within the time window and the key points of the wrists of both hands are close to the midline of the body, it is judged that the hands are crossed; calculate the angle between the waist and the back. If the angle gradually increases within the time window and exceeds 15 degrees, it is judged that the body is leaning back.
[0117] By integrating data on forward leaning, backward leaning, crossed hands, and natural hand placement, we obtain data on body posture.
[0118] In this embodiment of the invention, social language data, including each word and its timestamp, is read from a file. Natural language processing tools are used to segment the text and extract sentiment expression words. Based on a predefined sentiment lexicon (e.g., a list of positive and negative words), the words are categorized into positive and negative. The extracted sentiment words and their corresponding timestamps are recorded to ensure accuracy for subsequent synchronous analysis with video frames. A video processing tool (e.g., OpenCV) is used to load the video file. The video content is read frame by frame, and each frame is saved as a separate image file. A limb detection tool is used to process each frame, extracting key points of the human body, including the shoulders, waist, hips, and wrists. Based on the timestamps of the social language data, the corresponding video frames are extracted for limb posture analysis. The coordinates of the key points of the shoulders and waist are obtained using the limb detection tool. The vertical distance between the center point of the shoulder and the center point of the waist is calculated. Within a set time window (e.g., 5 seconds), the change in vertical distance is compared. If the vertical distance gradually decreases, it is determined that the body is leaning forward. The coordinates of the key points of both wrists are obtained using the limb detection tool. Calculate the vertical distance between the key points of both wrists and the horizontal line of the shoulders. Within a set time window (e.g., 5 seconds), if the key points of both wrists are below the horizontal line of the shoulders and remain stable, it is judged as a natural hand placement. Obtain the coordinates of the key points of both wrists using a limb detection tool, calculate the horizontal distance between the key points of both wrists, and within a set time window (e.g., 5 seconds), if the horizontal distance gradually decreases and the key points of both wrists move closer to the midline of the body, it is judged as crossed hands. Obtain the coordinates of the key points of the waist and back using a limb detection tool, calculate the angle between the waist and back, and within a set time window (e.g., 5 seconds), if the angle gradually increases and exceeds 15 degrees, it is judged as leaning backward. Integrate the above analysis of forward leaning, backward leaning, crossed hands, and natural hand placement; record the integrated limb posture data, including emotional tendency, limb posture characteristics, and corresponding timestamps, for subsequent analysis. The recorded data is stored as a CSV file with the following format: Column 1: Timestamp; Column 2: Emotional tendency (positive / negative); Column 3: Body posture characteristics (leaning forward, leaning back, hands crossed, hands placed naturally); Column 4: Corresponding video frame number.
[0119] Preferably, the communication barrier assessment based on eye gaze duration and facial muscle response data in step S3 includes:
[0120] Determine the start and end times of eye fixation duration; determine the start and end times of facial muscle response data;
[0121] The duration of eye fixation is compared with the start and end times of facial muscle response data. If the difference between the start time of eye fixation and the start time of facial muscle response is no more than 1 second, it is marked as a synchronous response. If the difference between the start time of eye fixation and the start time of facial muscle response is more than 1 second, it is marked as an asynchronous response.
[0122] Calculate the ratio of eye fixation duration to facial muscle response duration; if the ratio is less than 0.5, it is considered a low ratio; if the ratio is between 0.5 and 1.5, it is considered a medium ratio; if the ratio is greater than 1.5, it is considered a high ratio.
[0123] Communication impairment is assessed based on the synchronicity and ratio range of eye contact and facial muscle responses; if the responses are synchronous and the ratio is low: it is marked as a slow communication pattern; if the responses are synchronous and the ratio is high: it is marked as a hasty communication pattern; if the responses are asynchronous and the ratio is low: it is marked as a disharmony communication pattern; if the responses are asynchronous and the ratio is high: it is marked as a disconnected communication pattern.
[0124] The above assessment results of communication barriers are summarized into a communication barrier model.
[0125] In this embodiment of the invention, gaze data is read from a file. This data includes the gaze state of each frame and its corresponding frame number. The data is traversed to find the frame number where the gaze state changes from "not gazing" to "gazing," and this frame number is recorded as the start time of the gaze. The data is then traversed again to find the frame number where the gaze state changes from "gazing" to "not gazing," and this frame number is recorded as the end time of the gaze. The gaze duration = end time - start time. Facial muscle response data is also read from a file. This data includes the facial muscle activity state of each frame and its corresponding frame number. The data is traversed to find the frame number where the facial muscle activity state changes from "inactive" to "active," and this frame number is recorded as the start time of the facial muscle response. The data is then traversed again to find the frame number where the facial muscle activity state changes from "active" to "inactive," and this frame number is recorded as the end time of the facial muscle response. The facial muscle response duration = end time - start time. The gaze data and facial muscle response data are aligned according to their frame numbers to ensure that the timestamps of the two sets of data are consistent. For each set of eye fixation and facial muscle response data, calculate the difference between the start time of eye fixation and the start time of facial muscle response. If the time difference is less than 1 second, it is marked as "synchronous response"; if the time difference is more than 1 second, it is marked as "asynchronous response". For each set of eye fixation and facial muscle response data, calculate the ratio of the duration of eye fixation to the duration of facial muscle response. Ratio range: If the ratio is less than 0.5, it is marked as "low ratio". If the ratio is between 0.5 and 1.5, it is marked as "medium ratio". If the ratio is greater than 1.5, it is marked as "high ratio". Assessment pattern: Synchronous response and low ratio: marked as "delayed communication pattern". Synchronous response and high ratio: marked as "rapid communication pattern". Asynchronous response and low ratio: marked as "incoherent communication pattern". Asynchronous response and high ratio: marked as "disconnected communication pattern". The above assessment results were summarized into a communication barrier pattern and recorded in a data table, including timestamps, duration of eye contact, duration of facial muscle responses, synchronicity markers, ratio ranges, and communication barrier patterns.
[0126] Preferably, the social interaction disorder assessment based on eye gaze duration and body posture data in step S3 includes:
[0127] Determine the start and end times of eye fixation duration; determine the start and end times of facial muscle response data; define eye fixation duration exceeding 3 seconds as long fixation and eye fixation duration less than 1 second as short fixation.
[0128] Extract the angle between the torso and the vertical direction and the extension angle of the knee joint from the limb posture data; define the change of the torso's vertical angle as small change of 15-30 degrees, and the change of the torso's vertical angle as large change of more than 30 degrees.
[0129] Compare the duration of eye fixation with changes in trunk angle and knee extension angle. If the duration of eye fixation is long and the trunk angle changes little, it is marked as a social interaction coordination pattern; if the duration of eye fixation is short and the trunk angle changes greatly, it is marked as a social interaction disruption pattern.
[0130] By integrating the social interaction coordination model and the social interaction disruption model, a social interaction disorder model is obtained.
[0131] In this embodiment of the invention, an eye tracker is used to collect gaze data of the subject. The eye tracker records the subject's eye movement trajectory during video playback in a specific social scenario at a high sampling rate (e.g., 1000Hz). The collected eye movement data is filtered to remove noise interference caused by blinking, head micro-movements, etc. A Kalman filter is used to smooth the eye movement data to reduce random fluctuations. The start and end times of gaze are determined by the changes in the coordinates of the gaze point in the eye movement data. When the change in the gaze point coordinates is less than a set threshold (e.g., 1° of visual angle) within a certain time interval (e.g., 100ms), it is determined as the start of a gaze, and the gaze ends when the change in the gaze point coordinates exceeds the threshold. Gaze duration is classified: gaze duration exceeding 3 seconds is marked as long gaze, and gaze duration less than 1 second is marked as short gaze. A facial motion coding system (FACS) combined with a high-precision camera is used to collect facial expression video data of the subject, with the camera's frame rate at 30fps. Inter-frame differencing was performed on the acquired video data to highlight subtle changes in facial muscle movement. Simultaneously, the facial region in each frame was located and cropped to ensure the accuracy of subsequent analysis. By analyzing the cropped facial region image sequence, optical flow was used to calculate the pixel motion vector between each frame. When the average value of the facial muscle motion vector exceeded a set threshold (e.g., 0.1 pixels / frame) for several consecutive frames (e.g., 5 frames), it was considered the start of a facial muscle response; the response ended when the average value of the motion vector fell below the threshold for several consecutive frames. Inertial Measurement Unit (IMU) sensors were used to acquire the subject's limb posture data. The IMU sensors were mounted near the subject's torso and knee joints, with a sampling rate of 100Hz. The acquired limb posture data underwent time alignment to ensure temporal consistency between the torso's angle with the vertical direction and the knee joint's extension angle. The data was also smoothed, with a window size of 10 sampling points. Acceleration and angular velocity data collected by IMU sensors were used to calculate the angle between the torso and the vertical direction, as well as the knee extension angle, using a quaternion algorithm. A change of 15-30 degrees in the torso's vertical angle was defined as a small change, and a change exceeding 30 degrees was defined as a large change. The processed time-series data were synchronized, and the duration of eye contact was compared with the changes in torso angle and knee extension angle using timestamps as a baseline. When the eye contact duration was long and the torso angle showed a small change, it was marked as a social interaction coordination pattern; when the eye contact duration was short and the torso angle showed a large change, it was marked as a social interaction disruption pattern. The marked social interaction coordination and disruption patterns were integrated, and the frequency and duration of both patterns within a specific time period were statistically analyzed to obtain a social interaction disorder pattern.
[0132] Preferably, step S4, which involves matching the social interaction impairment pattern and the communication impairment pattern with the autism spectrum disorder history database, includes:
[0133] Extract the occurrence time, duration, and frequency of social interaction disorder patterns and record them as social interaction disorder feature data;
[0134] Extract the occurrence time, duration, and frequency of communication barrier patterns and record them as communication barrier characteristic data;
[0135] The data on social interaction impairment and communication impairment features were matched with the historical database of autism spectrum disorder. The similarity between each pattern and the known patterns in the historical database of autism spectrum disorder was determined and the similarity was quantified. If the similarity quantification value exceeded 0.8, the match was considered successful.
[0136] Interaction barrier matching data and communication barrier matching data are recorded based on similarity metric values.
[0137] In this embodiment of the invention, the labeled social interaction disorder pattern data is sorted by timestamp to ensure data continuity; the starting position (left) and ending position (right) of the window are initialized, the window size is set to 1 second, and the step size is 0.5 seconds. The window slides across the data sequence, moving the right pointer each time until the window covers 1 second of data. For each window, the timestamp of the window's starting point is recorded as the pattern occurrence time. The difference between the timestamp of the window's ending point and the timestamp of the starting point is calculated as the pattern duration. Within a specific time interval (e.g., 1 minute), the number of patterns covered by the window is counted as the pattern frequency. The extracted feature data is stored in a structured data table, with fields including "pattern occurrence time," "pattern duration," and "pattern frequency." The communication disorder pattern data is sorted by timestamp, using the same sliding window parameters as the social interaction disorder pattern (window size 1 second, step size 0.5 seconds). Following the steps of the sliding window algorithm described above, the pattern occurrence time, pattern duration, and pattern frequency are recorded respectively. The extracted feature data is stored in a structured data table, with fields consistent with the social interaction disorder feature data table. The feature data of social interaction impairment and communication impairment were normalized using a min-max normalization method, scaling all feature values to a range of 0 to 1. Feature data of known patterns were extracted from the autism spectrum disorder history database and similarly normalized. For each pattern to be matched, the Euclidean distance between it and known patterns in the history database was calculated on three feature dimensions: "pattern occurrence time," "pattern duration," and "pattern frequency." The Euclidean distance was used to calculate the distance between two feature vectors; a smaller distance indicates a higher similarity. The Euclidean distances of the three feature dimensions were weighted and summed to obtain a comprehensive similarity metric. The weights were allocated according to the importance of each feature in autism spectrum disorder identification; a match was considered successful if the similarity metric exceeded 0.8. The results were stored in a data table named "Interaction Impairment Matching Results," with fields including "Matching Pattern," "Similarity Metric," and "Matching Result." Data was also stored in a data table named "Communication Impairment Matching Results," with the same fields as the "Interaction Impairment Matching Results" data table.
[0138] Preferably, the construction of the autism spectrum disorder identification model based on interaction impairment matching data and communication impairment matching data in step S4 includes:
[0139] Extract high-dimensional feature vectors from interaction barrier matching data and communication barrier matching data;
[0140] High-dimensional feature vectors are input into a deep learning model, and feature fusion is performed using MLP technology. An attention mechanism is introduced into the deep learning model to weight different feature vectors in order to generate a comprehensive feature representation.
[0141] An autoencoder is used to reduce the dimensionality of the comprehensive feature representation to obtain the feature dimensionality-reduced representation data.
[0142] The data is represented by reduced-dimensionality features and input into a classifier, and then classified using a decision tree to generate preliminary recognition results.
[0143] The initial recognition results are smoothed using a sliding window technique to obtain the processed recognition results.
[0144] The processed recognition results are then matched a second time with known patterns in the historical database, and time series alignment is performed to obtain the second matching results.
[0145] A model for identifying autism spectrum disorder was constructed based on the results of the second-order matching.
[0146] In this embodiment of the invention, feature extraction is performed on interaction barrier matching data and communication barrier matching data using a convolutional neural network (CNN) structure from deep learning. A CNN model containing multiple convolutional and pooling layers is constructed. The input is the feature matrix of the matching data, and the output is a high-dimensional feature vector. For example, if the input feature matrix has a dimension of 100×50 (assuming 100 feature points, each with 50 dimensions), after processing by the CNN model, the output high-dimensional feature vector has a dimension of 128. A multilayer perceptron (MLP) model is also constructed, containing multiple hidden layers, each using the ReLU activation function. The input is a high-dimensional feature vector, with the number of neurons in the hidden layers being 256, 128, and 64, respectively, and the number of neurons in the output layer being 32. An attention mechanism is introduced into the MLP model, specifically implemented by adding an attention layer. The weights of the attention layer are automatically learned through training, weighting different feature vectors to generate a comprehensive feature representation. For example, if the input high-dimensional feature vector has a dimension of 128, after processing by the MLP and attention mechanism, the generated comprehensive feature representation has a dimension of 32. Construct an autoencoder model. The encoder compresses the comprehensive feature representation of the input into a low-dimensional representation, and the decoder reconstructs the low-dimensional representation back to its original size. The encoder structure is: 32-dimensional input layer, 16-dimensional hidden layer, and 8-dimensional output layer. The decoder structure is: 8-dimensional input layer, 16-dimensional hidden layer, and 32-dimensional output layer. Mean squared error (MSE) is used as the loss function, and the Adam optimizer is used for training. Train the autoencoder model with the comprehensive feature representation as input and the reduced-dimensionality feature representation as output (8-dimensional). Construct a decision tree classifier with the reduced-dimensional feature representation as input and the preliminary recognition result as output. Use the trained decision tree model to classify the reduced-dimensional feature representation, generating preliminary recognition results. For example, if the reduced-dimensional feature representation is 8-dimensional, the decision tree model classifies based on these features and outputs preliminary recognition results. Use a sliding window technique to smooth the preliminary recognition results, with a window size of 5 and a stride of 1. Slide the window across the sequence of preliminary recognition results, statistically analyze the recognition results within the window, and take the most frequent recognition result as the final recognition result for that window. For example, the initial recognition result sequence is [1,2,1,1,2,1,1,1]. After sliding window smoothing, the processed recognition result sequence is [1,1,1,1]. The processed recognition result is then matched a second time with known patterns in the historical database using the Dynamic Time Warping (DTW) algorithm for time series alignment. The similarity between the processed recognition result and the known patterns in the historical database is calculated. If the similarity exceeds a set threshold (e.g., 0.8), the match is considered successful. For example, if the processed recognition result sequence is [1,1,1,1] and the known pattern sequence in the historical database is [1,1,1,1], after alignment using the DTW algorithm, the similarity is 0.9, indicating a successful match.Based on the results of the secondary matching, an autism spectrum disorder (ASD) identification model is constructed. The model input is the processed identification result, and the output is the ASD identification result. The trained model is used to identify new data to generate the final identification result. For example, if the input sequence of processed identification results is [1,1,1,1], the model outputs "Yes" as the ASD identification result.
[0147] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.
[0148] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.
Claims
1. A method for constructing an autism spectrum disorder identification model based on deep learning, characterized in that, The method comprises the following steps: Step S1: obtaining a conversation video of an autism patient; continuously monitoring the gaze fixation rate in the conversation video of the autism patient; when the gaze fixation rate is detected to be lower than a gaze fixation rate threshold, determining a gaze fixation reduction state and recording the gaze fixation duration, comprising: performing time series analysis on the gaze fixation rate data, dividing the gaze fixation rate data in chronological order into a plurality of time windows of a preset length, and taking the gaze fixation rate data in each time window as an independent analysis unit; in each time window, performing data smoothing processing on the gaze fixation rate data to obtain a smoothed gaze fixation rate curve; based on the smoothed gaze fixation rate curve, calculating the gaze fixation rate change rate in each time window, and determining the gaze fixation rate trend by comparing the gaze fixation rate change rates of adjacent time windows; when the gaze fixation rate decline trend shows a continuous decline, determining the time point at which the gaze fixation rate is lower than the fixation rate threshold, taking the time point as the starting point of the gaze fixation reduction state, and starting to record the fixation duration; in the process of recording the gaze fixation duration, the change of the gaze fixation rate is monitored in real time; if the gaze fixation rate is higher than the fixation rate threshold again, the recording of the fixation duration is stopped; the whole process time from the start of recording to the stop of recording is taken as the gaze fixation duration; Step S2: performing social language analysis on the conversation video of the autism patient to obtain social language data; identifying patient facial response features in the conversation video of the autism patient according to the social language data and recording facial muscle response data; identifying body posture features in the conversation video of the autism patient according to the social language data and recording body posture mode data, further comprising: extracting emotional expression words in the social language data and corresponding time stamps, and dividing the conversation content into two categories of emotional tendencies, positive and negative; synchronously analyzing the body posture in the video according to the time stamps of the social language data; when the conversation content is positive, analyzing whether the patient's body posture presents an open feature; calculating the vertical distance between the shoulder center point and the waist center point, if the vertical distance gradually decreases in the time window, it is judged as body leaning forward; calculating the vertical distance between the wrist key points of the hands and the shoulder horizontal line, if the wrist key points of the hands are below the shoulder horizontal line and remain stable in the time window, it is judged as natural placement of the hands; when the conversation content is negative, analyzing whether the target object's body posture presents a closed feature; calculating the horizontal distance between the wrist key points of the hands, if the horizontal distance gradually decreases in the time window and the wrist key points of the hands are close to the body midline, it is judged as hand crossing; calculating the angle between the waist and the back, if the angle gradually increases in the time window and exceeds 15 degrees, it is judged as body leaning back; integrating body leaning forward, body leaning back, hand crossing and natural placement of hands to obtain body posture mode data; Step S3: communication disorder evaluation based on the gaze fixation duration and facial muscle response data, comprising: determine the start time and end time of the gaze fixation; determine the start time and end time of the facial muscle reaction data; compare the gaze fixation duration with the start time and end time of the facial muscle reaction data; if the start time of the gaze fixation and the start time of the facial muscle reaction differ by no more than 1 second, mark it as a synchronous reaction; if the start time of the gaze fixation and the start time of the facial muscle reaction differ by more than 1 second, mark it as an asynchronous reaction; calculate the ratio of the gaze fixation duration to the facial muscle reaction duration; if the ratio is less than 0.5, determine it as a low ratio; if the ratio is 0.5 to 1.5, determine it as a medium ratio; if the ratio is greater than 1.5, determine it as a high ratio; evaluate the communication disorder according to the synchronicity of the gaze fixation and the facial muscle reaction and the ratio interval; if the synchronous reaction and the low ratio: mark it as a communication retardation mode; if the synchronous reaction and the high ratio: mark it as a communication rapid mode; if the asynchronous reaction and the low ratio: mark it as a communication incoordination mode; if the asynchronous reaction and the high ratio: mark it as a communication broken mode; summarize the communication disorder evaluation results above as a communication disorder mode; evaluate the social interaction disorder based on the gaze fixation duration and the body posture mode data, and obtain a social interaction disorder mode; Step S4: obtain an autism spectrum disorder history library; perform feature matching of the social interaction disorder mode and the communication disorder mode with the autism spectrum disorder history library respectively to generate interaction disorder matching data and communication disorder matching data; and construct an autism spectrum disorder recognition model based on the interaction disorder matching data and the communication disorder matching data. 2.The method of claim 1, wherein the method further comprises: The obtaining of the autism patient conversation video in step S1 includes: In a plurality of different scenes, the conversation video of the autism patient is recorded respectively, and the patient is in the center of the lens during the recording process, and the horizontal angle between the lens and the eyes of the patient is kept within 10°; The recorded video is preliminarily screened to remove the segments with more than 5 seconds of blurred screen caused by lens shaking, while keeping the segments with expression changes and body movements of the patient in the video; The retained video segments are cropped to crop the video frames of the first 3 seconds and the last 2 seconds of the video in which the patient does not enter the conversation state; The cropped video segments are spliced according to the timestamps of the patient's conversation, and the total duration of the spliced video is not less than 1 minute to obtain the autism patient conversation video.
3. The method for constructing a deep learning-based autism spectrum disorder identification model according to claim 1, characterized in that, The continuous monitoring of the gaze fixation video rate in the autism patient conversation video in step S1 includes: frame-by-frame analysis is performed on the autism patient conversation video to determine the eye direction of the patient in each frame; In each frame, it is judged whether the eye direction of the patient forms an angle less than 30° with the lens direction, and if the eye direction forms an angle less than 30° with the lens direction, the frame is recorded as an effective gaze fixation; The number of frames of effective gaze fixation in each statistical unit is counted, and the gaze fixation video rate of each statistical unit is obtained, with each 10 seconds as a statistical unit; The entire video is divided into statistical units according to the above statistical units, and the gaze fixation video rate of each statistical unit is calculated, and the frequency value of the statistical unit is recorded. The frequency values of all statistical units are summarized to calculate the average gaze fixation rate of the whole video as the gaze fixation rate in the autism patient conversation video. 4.The method of claim 1, wherein the method further comprises: training the deep learning-based autism spectrum disorder recognition model using the plurality of training data. The social language analysis of the autism patient conversation video in step S2 includes: Separating the audio signal from the autism patient conversation video, converting the audio signal into a text, and recording the pronunciation duration and intonation change of each word; Performing word segmentation processing on the text to split the text into individual words and label each word with its appearance timestamp in the audio; Extracting the pronunciation duration and intonation change of each word to generate a speech feature vector corresponding to each word; Statistically counting the frequency of each word in the text and combining the speech feature vector to mark the emotional tendency of each word; According to the emotional tendency, identify the interrogative and declarative sentences in the text, and count the number of interrogative and declarative sentences, while recording the average intonation change amplitude of each sentence type; Based on the average intonation change amplitude, mark the usage order of the words in the text, record the collocation relationship of each word, and obtain the social language data. 5.The method of claim 1, wherein the method further comprises: training the deep learning-based autism spectrum disorder recognition model using the plurality of training data. The patient facial response feature recognition in the autism patient conversation video based on the social language data in step S2 includes: Extracting each frame of image from the autism patient conversation video and performing face detection on each frame of image to locate the zygomaticus major region, corrugator supercilii region and medial frontalis region of the patient's face; Analyzing the facial contour in each frame of image to record the muscle activity intensity of the zygomaticus major region, corrugator supercilii region and medial frontalis region to generate the facial muscle activity intensity; According to the word appearance timestamp marked in the social language data, determine the video frame range corresponding to each word; Determine the occurrence time and duration of facial muscle activity intensity within the video frame range and summarize it into facial muscle response data. 6.The method of claim 1, wherein the method further comprises: The body posture feature recognition in the autism patient conversation video based on the social language data in step S2 includes: Extracting each frame of image from the conversation video and performing body detection on each frame of image to locate the body contour of the patient; Measuring the trunk vertical direction angle of the body contour, dividing the trunk vertical direction angle into three intervals, 0-15 degrees for normal posture, 15-30 degrees for slight forward leaning, and more than 30 degrees for severe forward leaning; Measuring the knee joint extension angle of the leg of the body contour, dividing the knee joint angle into three intervals, 160-180 degrees for complete straightening, 120-160 degrees for slight bending, and less than 120 degrees for severe bending; According to the word appearance timestamp marked in the social language data, determine the video frame range corresponding to each word; Record the trunk vertical direction angle and knee joint extension angle within the video frame range to form the body posture mode data. 7.The method of claim 1, wherein the method further comprises: training the deep learning-based autism spectrum disorder recognition model using the plurality of training data. The social interaction disorder assessment based on the gaze fixation duration and body posture mode data in step S3 includes: The starting time and the ending time of the gaze fixation are determined, and the gaze fixation duration is determined; the gaze fixation duration exceeding 3 seconds is set as a long gaze fixation, and the gaze fixation duration less than 1 second is set as a short gaze fixation; The angle between the trunk and the vertical direction and the extension angle of the knee joint in the body posture mode data are extracted; the change of the angle between the trunk and the vertical direction is set as a small change of the trunk angle when the change is 15-30 degrees, and the change of the angle between the trunk and the vertical direction is set as a large change of the trunk angle when the change exceeds 30 degrees; The gaze fixation duration, the change of the angle between the trunk and the vertical direction, and the extension angle of the knee joint are compared; if the gaze fixation duration is a long gaze fixation and the change of the angle between the trunk and the vertical direction is a small change of the trunk angle, the social interaction coordination mode is marked; if the gaze fixation duration is a short gaze fixation and the change of the angle between the trunk and the vertical direction is a large change of the trunk angle, the social interaction fracture mode is marked; The social interaction coordination mode and the social interaction fracture mode are integrated to obtain the social interaction disorder mode. 8.The method of claim 1, wherein the method further comprises: The feature matching of the social interaction disorder mode and the communication disorder mode with the autism spectrum disorder history library in step S4 includes: The mode occurrence time, the mode duration, and the mode frequency of the social interaction disorder mode are extracted and recorded as social interaction disorder feature data; The mode occurrence time, the mode duration, and the mode frequency of the communication disorder mode are extracted and recorded as communication disorder feature data; The social interaction disorder data and the communication disorder feature data are respectively matched with the autism spectrum disorder history library, the similarity of each mode with the known mode in the autism spectrum disorder history library is determined, and the similarity quantization value is quantized; if the similarity quantization value exceeds 0.8, it is considered that the matching is successful; The interaction disorder matching data and the communication disorder matching data are recorded based on the similarity quantization value. 9.The method of claim 1, wherein the method further comprises: The construction of the autism spectrum disorder recognition model based on the interaction disorder matching data and the communication disorder matching data in step S4 includes: The high-dimensional feature vector of the interaction disorder matching data and the communication disorder matching data is extracted; The high-dimensional feature vector is input into a deep learning model, the MLP technology is used for feature fusion, and in the deep learning model, an attention mechanism is introduced to weight process different feature vectors to generate a comprehensive feature representation; The self-encoder is used to reduce the dimension of the comprehensive feature representation to obtain feature dimension reduction representation data; The feature dimension reduction representation data is input into a classifier, and a decision tree is used for classification to generate a preliminary recognition result; The sliding window technology is used to smooth the preliminary recognition result to obtain a processed recognition result; The processed recognition result is secondarily matched with the known mode in the history library, and time sequence alignment is performed to obtain a secondary matching result; The autism spectrum disorder recognition model is constructed according to the secondary matching result.
Citation Information
Patent Citations
Machine learning-based autism spectrum disorder early recognition method
CN115565690A
Autism auxiliary diagnosis method based on eyeball tracking technology
CN118866324A
Language evaluation method and device for autistic children and medium
CN119028594A