Screening Method for Autism Spectrum Disorder Based on Scenario Video Content

KR103025079B1Active Publication Date: 2026-09-29INSIGHT CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
KR1020250188186
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-09-29
Estimated Expiration
2045-12-02

Smart Images

  • Figure 112025135805159-PAT00002_ABST
    Figure 112025135805159-PAT00002_ABST
Patent Text Reader

Abstract

The present invention relates to a scenario video content-based autism spectrum disorder screening method that diagnoses the risk of ASD (Autism Spectrum Disorder) by providing a scenario-based video content to an examiner, actively inducing a specific response corresponding to a behavior-inducing event, and then analyzing the examiner's gaze and behavior. The method comprises: (a) playing a scenario video content including a behavior-inducing event of the examiner; (b) collecting examiner response data including gaze tracking data and behavior tracking data from a video of the examiner filmed while playing the scenario video content; (c) generating a data set aligned on the same time axis by synchronizing the time stamp of the behavior-inducing event with the time stamp of the examiner response data; and (d) analyzing the examiner's behavior pattern using the data set synchronized in step (c) and calculating an ASD index based thereon.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to a scenario-based video content screening method for autism spectrum disorder (ASD) that diagnoses the risk of ASD by providing a scenario-based video content to an examiner, actively inducing a specific response corresponding to a behavior-inducing event, and then comprehensively analyzing the examiner's gaze and behavior. Background Technology

[0002] Autism Spectrum Disorder (ASD) is a neurodevelopmental disorder characterized by difficulties in social interaction and communication, as well as repetitive behaviors or restricted interests. As the term "spectrum" suggests, ASD manifests in various ways depending on the severity and pattern of symptoms. Since symptoms of this autistic disorder often appear during childhood, typically around the age of three, it is crucial to receive an early diagnosis and appropriate treatment during this period.

[0003] Although the causes of ASD have not yet been fully elucidated, it is known that genetic factors and physiological differences related to brain development play a significant role. In addition, environmental factors are also known to influence the onset of the condition. According to one study, environmental factors during pregnancy or health status before and after birth affect the development of autism.

[0004] Early detection is the most critical factor in treating ASD. Traditionally, clinicians or therapists have analyzed responses on-site by presenting specific pictures, videos, or teaching aids; however, this approach relies on the observer's experience and subjectivity, and results can vary depending on changes in the testing environment, making it difficult to secure consistent diagnostic indicators. In other words, technologies capable of quantitatively evaluating ASD in infants and young children in clinical settings have not yet been sufficiently established.

[0005] Meanwhile, there has recently been an increasing number of technological attempts to quantitatively measure signs of ASD using video-based behavior analysis technology. Furthermore, some of the previously proposed behavior analysis technologies are limited to methods that merely determine the presence of repetitive behaviors by attaching wearable sensors to the examiner or analyzing simple camera footage; consequently, they face the problem of being unable to evaluate responsiveness to social stimuli and struggle to elicit natural responses in the testing environment due to the discomfort of wearing the devices. Prior art literature

[0006] Republic of Korea Published Patent No. 10-2025-0093254 "System for Providing Parent Response Guides through AI-based Analysis of Child Language and Behavior" (Published June 24, 2025) The problem to be solved

[0007] The purpose of the present invention is to provide a novel scenario video content-based autism spectrum disorder screening method capable of quantitatively and objectively analyzing the examiner's behavioral patterns and producing consistent ASD indicators by inducing natural behavior from the examiner based on scenario video content and performing precise time-axis analysis by synchronizing eye-tracking data and behavior tracking data during the process.

[0008] The problems solved by the present invention are not limited to those mentioned above, and other problems not mentioned will be clearly understood by those skilled in the art from the description below. means of solving the problem

[0009] A screening method for autism spectrum disorder based on scenario video content according to an embodiment of the present invention comprises: (a) playing scenario video content including an event that triggers an examiner's behavior; (b) collecting examiner response data including eye-tracking data that tracks the gaze and behavior-tracking data that tracks the behavior from a video of the examiner while playing the scenario video content; (c) generating a data set aligned on the same time axis by synchronizing the time stamp of the behavior-triggering event with the time stamp of the examiner response data; and (d) analyzing the examiner's behavior pattern using the data set synchronized in step (c) and calculating an ASD (Autism Spectrum Disorder) index based thereon.

[0010] A screening method for autism spectrum disorder based on scenario video content according to another embodiment of the present invention is characterized in that step (d) comprises: (d-1) a preprocessing step of performing noise removal, coordinate transformation and normalization, and missing value interpolation on the eye tracking result of the examiner; (d-2) a step of setting a Region of Interest (ROI) at the point where the behavior-inducing event occurs in the scenario video content; (d-3) a step of determining whether the examiner gazes at the ROI; and (d-4) a step of extracting and quantifying the gaze delay time and duration based on the determination result of step (d-3), and applying the quantified indicator to a probabilistic learning model to calculate the ASD indicator as a probability value.

[0011] A screening method for autism spectrum disorder based on scenario video content according to another embodiment of the present invention is characterized in that step (d) comprises: (d-1) displaying an attention contrast screen on the scenario video content; (d-2) outputting a pre-recorded call sound while the attention contrast screen is displayed; (d-3) reading the behavior tracking data and observing the reaction delay time at the point when the examiner's face direction faces the point when the call sound is output relative to the point when the call sound is output; and (d-4) quantifying the reaction delay time to adjust the ASD index.

[0012] A screening method for autism spectrum disorder based on scenario video content according to another embodiment of the present invention is characterized in that step (d) comprises: (d-1) playing an imitation-inducing video designed to have a person in the scenario video content perform a specific movement including time-series coordinates of the head, shoulder, arm, and wrist joints; (d-2) reading the behavior tracking data and extracting the joint keypoint coordinates of the examiner's head, shoulder, arm, and wrist as time stamps synchronized with the imitation-inducing video; (d-3) calculating a similarity score by comparing and determining the frame-unit similarity between the time-series coordinates of the imitation-inducing video and the joint keypoint coordinates; and (d-4) quantifying the examiner's imitation behavior based on the similarity score to calculate the ASD index.

[0013] A screening method for autism spectrum disorder based on scenario video content according to another embodiment of the present invention is characterized in that step (d) comprises: (d-1) an LSTM modeling step of extracting repetitive behavior features based on joint keypoint time series data of the examiner while playing the scenario video content; (d-2) a step of detecting the repetitive behavior of the examiner; (d-3) a step of quantifying the periodicity, frequency, or duration of the repetitive behavior; and (d-4) a step of calculating the ASD index based on the quantified stereotyped behavior index.

[0014] A screening method for autism spectrum disorder based on scenario video content according to another embodiment of the present invention comprises: (e) a step of extracting gaze features and behavior features by inputting the examiner’s gaze tracking data and behavior tracking data into a gaze encoder and a behavior encoder, respectively; (f) a step of generating multimodal features by integrating the gaze features and the behavior features through a feature fusion layer; (g) a step of learning high-dimensional features by inputting the multimodal features into a learning neural network; (h) a step of producing a final feature vector by inputting the high-dimensional features into an MLP or Transformer structure; and (i) a step of classifying the examiner’s ASD grade by computing the final feature vector. Effects of the invention

[0015] According to the scenario video content-based autism spectrum disorder screening method of the present invention, natural behavior of the examiner is induced based on scenario video content, and by synchronizing eye-tracking data and behavior tracking data during the process to perform precise time-axis analysis, the examiner's behavioral patterns can be quantitatively and objectively analyzed and consistent ASD indicators can be calculated.

[0016] The effects of the present invention are not limited to those mentioned above, and effects for solving other problems not mentioned will be clearly understood by those skilled in the art from the description below. Brief explanation of the drawing

[0017] FIG. 1 is a block diagram illustrating an ASD behavior analysis system in which the present invention is implemented. FIG. 2 is a flowchart illustrating a scenario video content-based ASD screening method according to the present invention, FIG. 3 is a flowchart illustrating the gaze-based ASD diagnosis process in the present invention, FIG. 4 is a flowchart illustrating a behavior-based ASD diagnosis process in the present invention, and FIG. 5 is a block diagram illustrating the multimodal fusion ASD diagnostic process in the present invention. Specific details for implementing the invention

[0018] Further objects, features, and advantages of the present invention can be more clearly understood from the following detailed description and the accompanying drawings.

[0019] Before providing a detailed description of the present invention, it should be understood that the present invention is capable of various modifications and may have various embodiments, and that the examples described below and illustrated in the drawings are not intended to limit the present invention to specific embodiments, but rather include all modifications, equivalents, and substitutions that fall within the spirit and scope of the present invention.

[0020] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.

[0021] The terms used in this specification are used merely to describe specific embodiments and are not intended to limit the invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as "comprising" or "having" are intended to indicate the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0022] Furthermore, in the description referring to the attached drawings, identical components are assigned the same reference numeral regardless of drawing symbols, and redundant descriptions thereof are omitted. In describing the present invention, if it is determined that a detailed description of related prior art could unnecessarily obscure the essence of the present invention, such detailed description is omitted.

[0023] FIG. 1 is a block diagram illustrating an ASD behavior analysis system in which the present invention is implemented. Referring to FIG. 1, the ASD (Autism Spectrum Disorder) behavior analysis system comprises an ASD analysis device (100), a scenario video playback device (182, 184), an eye tracking device (186), and a camera (188). Additionally, it may further include an analysis result output device (190) that outputs the analysis results of the ASD analysis device (100).

[0024] First, the scenario video playback device (182) is composed of a display means (182) and a speaker (184) and is a device that outputs a voice signal including a scenario video and a calling sound provided by the ASD analysis device (100). For example, the display means (182) can be implemented as a monitor, tablet, VR device, or mixed reality (MR) display.

[0025] The ASD screening method of the present invention is characterized by diagnosing ASD by utilizing eye tracking data and behavior tracking data in combination. Here, eye tracking and behavior tracking can be extracted by separating data preprocessed by tracking the gaze and data preprocessed by tracking the behavior from a video captured by a single camera (188). For example, behavior tracking data described below can be extracted from a video captured by a camera (188), and eye tracking data can be generated separately by extracting the user's gaze and pupils using an AI model. In addition, as listed below, eye tracking and behavior tracking can be collected separately by utilizing separate image detection means such as an eye tracking device (186) and a camera (188).

[0026] The eye tracking device (186) tracks the examiner's gaze and generates eye tracking data, which is then transmitted to the ASD analysis device (100). The eye tracking device (186) is one of an infrared-based eye tracker, a webcam-based gaze estimation, or an infrared (IR)-based gaze tracking module. The eye tracking data includes at least one of the examiner's 2D gaze point coordinates (gaze-2D), 3D gaze point coordinates (gaze-3D), pupil size, and head-to-screen distance.

[0027] For example, the eye tracking device (186) can irradiate IR (infrared) light and detect a glint reflected from the surface of the examiner's eyeball to calculate a gaze vector based on the relative position between the center of the pupil and the corneal reflection. The detected gaze coordinates can be provided as 2D coordinates (x, y). In the case of an IR stereo-based eye tracking device (186), they can also be provided as 3D coordinates (x, y, z). Additionally, the pupil size and head-to-screen distance can be calculated. Through this, it is possible to precisely calculate whether the examiner gazes at the Region of Interest (ROI) of the action-inducing event within the screen of the scenario video content and the duration of the gaze, and as described below, it is possible to quantitatively evaluate how accurately the examiner recognizes and gazes at the region of interest (ROI) within the video, as well as the gaze delay time and duration.

[0028] The camera (188) is a means for photographing at least the upper body of the examiner, and preferably a means for photographing the entire body. Preferably, the camera (188) is an RGB-depth camera that extracts the examiner's joint keypoints into three-dimensional coordinate values ​​to generate behavior tracking data. In addition, the camera (188) may transmit data such as the examiner's posture changes and reaction movements to the ASD analysis device (100) in addition to the joint keypoint coordinates.

[0029] The behavior tracking data generated by the camera (188) is used for various analyses, such as estimating the examiner's pose, analyzing the call response, analyzing imitation behavior, and analyzing stereotyped behavior. For example, when the examiner performs a specific hand movement in an imitation-inducing video, the camera tracks the three-dimensional time-series coordinates of the head, shoulder, arm, and wrist joints and transmits them to the imitation behavior analysis unit (430) described later, and the imitation behavior analysis unit (430) calculates a similarity score by comparing the similarity of the movement with the person in the video.

[0030] For example, the camera (188) can be implemented as an RGB-D camera such as Intel RealSense, Azure Kinect, Orbbec Astra, etc. An advantage of using an RGB-D camera is that the key points of the examiner's entire body joints can be extracted as three-dimensional coordinates (x,y,z). This allows for a significant improvement in the quantitative analysis of the call response, imitation behavior, and stereotyped behavior described later.

[0031] The RGB-D camera (188) can calculate depth (Time-of-Flight) by irradiating an object with an IR pulse and using the time it returns, calculate depth based on deformation by projecting a pattern, or calculate depth using the left and right camera disparity of Stereo Vision. Additionally, the RGB-D camera (188) may be capable of skeleton extraction, 3D trajectory analysis, and reaction delay time calculation, and the behavior tracking data may include RGB video streams, depth maps, IR images, 3D skeletal joint coordinates, body tracking confidence scores, etc.

[0032] The analysis result output device (190) is a device that outputs the analysis results of the ASD analysis device (100). The analysis result output device (190) provides various analysis result information, such as ASD indicators, the examiner's gaze pattern, name-calling response analysis results, imitation behavior analysis results, stereotyped behavior analysis results, and ASD classification, to an examiner, guardian, or clinician in a visual or auditory form.

[0033] The analysis result output device (190) may include various output means. For example, it may be implemented as a monitor, tablet, touchscreen display, etc., and may include a speaker depending on the case. The analysis result output device (190) is the same device as the scenario video playback device (182, 184) and may be a device that displays the inspection results on-site after the scenario video for the inspection has ended. In addition, the analysis result output device (190) may display the ASD analysis results in a quantified, intuitive graphic form such as a chart, graph, or color display.

[0034] The ASD analysis device (100) is a core processing unit of the system and is configured to include a scenario video output unit (110), a call sound output unit (120), a data collection unit (130), a trigger time synchronization unit (140), and an ASD analysis unit (160, 170, 180). Here, the ASD analysis unit (160, 170, 180) may be composed of a gaze-based ASD analysis unit (150), a behavior-based ASD analysis unit (160), and a multimodal fusion analysis unit (170).

[0035] The scenario video output unit (110) transmits the scenario video content to the display means (182) of the scenario video playback device. Here, the scenario video content is a series of scenarios designed for ASD screening or behavior analysis, comprising videos that include events triggering the examiner's behavior. The scenario video content is configured to induce the examiner's gaze and behavior and to allow the reaction to be observed.

[0036] For example, the scenario video content may feature a specific toy or object appearing on the screen to induce the examiner to focus their gaze on that object. The scenario video content may include an attention contrast video—such as a video where playback is suddenly stopped and a black screen is displayed—and when the attention contrast video is played, the call sound output unit (120) may play the examiner's name, which has been memorized in advance, as a voice signal. Additionally, the scenario video content may be a video in which a person in the video performs a specific action (e.g., waving hands, picking up an object) to observe the examiner's gaze and imitation behavior, or a video that includes an event scene to induce the examiner to perform the action repeatedly.

[0037] The data collection unit (130) collects the examiner's eye tracking data and behavior tracking data from the eye tracking device (186) and the camera (188). The collected data may undergo synchronization and preprocessing so that it can be processed by the ASD analysis unit (150, 160, 170).

[0038] The trigger time synchronization unit (140) synchronizes the timestamp of the behavior-inducing event with the timestamp of the examiner response data (eye tracking data and behavior tracking data) to create a data set aligned on the same time axis. This enables the ASD analysis unit (150, 160, 170) to accurately analyze the gaze and behavior patterns.

[0039] The gaze-based ASD analysis unit (150) processes the collected gaze tracking data and performs preprocessing steps such as noise removal, coordinate transformation and normalization, and missing value interpolation. In addition, it sets a region of interest (ROI) at the action-inducing event point within the scenario video content, determines whether the examiner gazes at the ROI, quantifies the gaze delay time and duration, and applies them to a probabilistic learning model to calculate an ASD index.

[0040] The behavior-based ASD analysis unit (160) performs pose estimation, name-calling response analysis, imitation behavior analysis, and stereotyped behavior analysis based on the examiner's joint keypoint coordinate data. The behavior-based ASD analysis unit (160) will be described in detail later with reference to FIG. 4.

[0041] The multimodal fusion analysis unit (170) integrates the results of the gaze-based ASD analysis unit (150) and the behavior-based ASD analysis unit (160) to generate multimodal features. The generated multimodal features are input into a high-dimensional feature learning neural network to learn high-dimensional features, and after producing a final feature vector, the ASD grade is classified. The multimodal fusion analysis unit (170) will be described in detail later with reference to FIG. 5.

[0042] The ASD analysis device (100) illustrated in FIG. 1 is not limited to a specific hardware device and can be implemented in various forms of computing environments. For example, the ASD analysis device (100) may be any one of a computing device, a workstation, an embedded device, an On-Premise Server, a Cloud Server, an edge device, an AI Box, or an IoT device. Additionally, the eye tracking device (186) and the camera (188) may be directly connected to a local device via USB or Wi-Fi, and AI operations such as eye analysis, pose estimation, call response analysis, mimicry behavior analysis, and stereotyped behavior analysis described later may be performed within the device, and the analysis results may be displayed on the device's screen or a connected monitor. For example, such an implementation method may be applied in a hospital examination room or a public health center where the communication environment is not stable.

[0043] An on-premise server may be a server system installed within a hospital or institution. In this case, the terminal (tablet or PC) plays scenario video content and collects gaze and behavior data in real-time or in batches, transmitting it to the server. The server can perform AI analysis based on high-performance GPUs / CPUs and return ASD risk results to the terminal.

[0044] The ASD analysis device (100) may be implemented on a GPU-based virtual server provided by a cloud service. In this case, the user terminal is responsible only for playing scenario video content and collecting inspector response data, and the data transmitted from the user terminal is delivered to the cloud server to continuously expand high-performance computing resources or to allow multiple institutions to share the same analysis policy. Additionally, when the server model is updated, all users can use the latest version at once, and this embodiment is particularly suitable for a SaaS structure that provides ASD screening services to a large number of users.

[0045] In addition, it may be implemented based on a mobile application (App). In this case, scenario video content is provided in the form of streaming within the app or via a CDN, eye tracking uses a front camera-based virtual eye tracking algorithm (2D landmark-based gaze estimation), and behavior-based analysis can be implemented based on the device camera's skeleton tracking.

[0046] FIG. 2 is a flowchart illustrating a scenario video content-based ASD screening method according to the present invention. Referring to FIG. 2, the step begins by playing scenario video content that includes an event triggering an action of the examiner (ST210). For example, the action triggering event may include stopping the video being played and outputting an attention contrast screen, a character on the screen waving a hand toward the examiner, a specific voice instruction (such as "look here," "wave your hand," etc.), the starting point of an action in a mimicry-inducing video, or the appearance of a specific stimulus element that triggers a repetitive action. The action triggering event may include metadata such as the type of event, the time of occurrence, the duration, related ROI information, and the coordinates of a specific action or pose.

[0047] Next, examiner response data including the aforementioned eye tracking data and behavior tracking data is collected from the eye tracking device (186) and camera (188) (ST220). Here, the eye tracking data includes the examiner's 2D gaze point coordinates (gaze-2D), 3D gaze point coordinates (gaze-3D), pupil size, head distance, and a time stamp (t_gaze) for each data frame. The behavior tracking data includes an RGB video stream, depth map, IR image, 3D skeletal joint coordinates, body tracking confidence score, etc., and a time stamp (t_pose) for each frame.

[0048] Next, trigger time synchronization (ST230) is performed to synchronize the timestamp of the behavior-inducing event with the timestamp (t_gaze, t_pose) of the examiner response data to create a data set aligned on the same time axis. At this time, if there is an internal clock difference between the eye tracking device (186) and the camera (188), an offset correction between devices can be performed to eliminate this difference. Additionally, the difference in sampling rates between frames of each data can be corrected.

[0049] Next, the behavioral patterns of the examiner are analyzed using the data set synchronized in step ST230, and an ASD index is calculated based on this (ST240). In this step, the ASD index can be calculated by a probabilistic learning model (such as Logistic Regression, Random Forest, etc.). For example, at least one of a pre-trained neural network-based ASD scoring model, a time-series-based Transformer or MLP model, and a threshold-based classification rule (Threshold Rule) can be applied to calculate an ASD risk probability value (0 to 1), an ASD score (scored from 0 to 100), and an ASD grade (Level 1 to 3), and the analysis results can be generated in the form of a report. Then, the analysis result report can be output through an analysis result output device (190) (ST250).

[0050] FIG. 3 is a flowchart illustrating the gaze-based ASD diagnosis process in the present invention. With reference to FIG. 3, the ASD diagnosis and analysis process performed in the gaze-based ASD analysis unit (150) will be explained.

[0051] First, a preprocessing step is performed on the eye tracking results of the examiner to remove noise, transform coordinates and normalize, and interpolate missing values ​​(ST310). For example, fine shaking, light reflection, instantaneous sensor errors, etc. occurring in the eye tracking device (186) can be removed using a Kalman Filter, IIR Filter, Savitzky-Golay Filter, etc., the coordinate values ​​of the eye tracking data can be transformed into normalized coordinates in the range of 0 to 1, and if there are missing intervals due to eye blinking or screen deviation, spline interpolation, linear interpolation, etc. can be applied to ensure time series continuity.

[0052] Next, a Region of Interest (ROI) is set at the point where an action-inducing event occurs in the scenario video content (ST320). For example, in a scene where a character speaks to the user, the entire face area can be set as the character's Face ROI. As another example, a Gesture ROI can be set in a scene where a character waves their hand to greet.

[0053] Next, determine whether the examiner is looking at the ROI (ST330). For example, preprocessed gaze coordinates (gaze_x, gaze_y) are input, and based on the ROI boundary values ​​(x_min, x_max, y_min, y_max), determine whether the gaze coordinates (gaze_x, gaze_y) are located within the boundary value range. If the condition that the gaze coordinates (gaze_x, gaze_y) are within the boundary value range is satisfied, it is determined as ROI gaze (True), and if not satisfied, it is determined as ROI non-gaze (False).

[0054] The time delay and duration of the examination are extracted based on the point in time when ROI is determined to be true (ST340). Then, the extracted time information is quantified, and the quantified indicator is applied to a probabilistic learning model to calculate the ASD indicator as a probability value (ST350). After the ASD risk probability value, ASD score, and ASD grade have been calculated by ST250 of Fig. 2, the values ​​calculated in ST350 can be corrected or adjusted.

[0055] FIG. 4 is a flowchart illustrating the behavior-based ASD diagnosis process in the present invention. The behavior-based ASD analysis process will be explained in detail with reference to FIG. 4.

[0056] Referring to FIG. 4, the behavior-based ASD analysis unit is largely composed of a pose estimation unit (410), a name-calling response analysis unit (420), an imitation behavior analysis unit (430), and a stereotyped behavior analysis unit (440).

[0057] The pose estimation unit (410) receives RGB or RGB-depth image data collected from the camera (188) and calculates joint keypoint coordinates for each joint of the examiner's body. For example, the pose estimation unit (410) derives the joint positions of the examiner's head, neck, shoulders, elbows, wrists, chest, pelvis, knees, ankles, etc. by applying algorithms such as OpenPose, MediaPipe, HRNet, and PnP-based 3D pose estimation. When using an RGB-depth camera, 3D joint coordinates can be calculated using depth data. When using a mono-camera, a 3D pose can be reconstructed using a deep learning-based 3D lifting method.

[0058] The pose estimation unit (410) aligns joint coordinates along the time axis and converts them into a time-series data structure, after which it provides them in the same form to the call response analysis unit (420), the imitation behavior analysis unit (430), and the stereotyped behavior analysis unit (440). The time-series data of joint keypoint coordinate values ​​generated by the pose estimation unit (410) is used as base data for quantitatively calculating various behavioral characteristics.

[0059] The call response analysis unit (420) evaluates the directional response of the face or head to a sound stimulus (call sound) to analyze ASD characteristics. Referring to FIG. 4, the call response analysis unit (420) outputs an attention contrast screen to the scenario video content (ST421). For example, it outputs an attention contrast screen such as simplifying the background in the scenario video content or outputting a black screen to prevent the examiner's gaze from being distracted.

[0060] Next, a pre-recorded call sound of the examiner is output (ST422). At this time, a directional speaker can be used to pre-define the output direction (left / right / center) of the call sound.

[0061] Next, the response delay time is observed (ST423). For example, the response delay time can be observed by recognizing the time when the examiner turns their face toward the direction of the call sound (T_face-turn) relative to the time when the call sound occurs (T_trigger), and then calculating the time difference between the two times (T_face-turn - T_trigger).

[0062] Then, the response to the call is analyzed based on the response delay time (ST424). At this time, if the response is excessively delayed (e.g., 2 seconds or more), the ASD risk level is increased, and if there is no response at all (NR: No Response), the ASD risk level can be adjusted to the highest grade.

[0063] The imitation behavior analysis unit (430) analyzes how accurately the examiner imitates specific actions performed by a character or person appearing in the scenario video content. The imitation behavior analysis unit (430) first plays an imitation-inducing video (ST431). For example, the imitation-inducing video may consist of a video of a character performing social interaction actions such as waving arms, raising hands, moving up and down, and turning left and right.

[0064] Next, the examiner's joint keypoint coordinates are extracted (ST432), and similarity is determined by comparing the frame-by-frame joint coordinates of the imitation-inducing video with the examiner's joint keypoint coordinates (ST433). Similarity can be calculated by performing Dynamic Time Warping (DTW), Cosine similarity, Euclidean Distance, and Joint-angle similarity on a frame-by-frame basis to produce an imitation behavior score. Then, based on the similarity score, the examiner's imitation behavior is quantified to calculate an ADS index (ST434). Since imitation behavior is one of the key indicators of social interaction development, it is a very important factor in identifying ASD characteristics.

[0065] The stereotyped behavior analysis unit (440) detects repetitive behaviors that occur while the examiner is watching the scenario video content. The stereotyped behavior analysis unit (440) analyzes joint keypoint coordinate values ​​to detect repetitive patterns based on time series.

[0066] The stereotyped behavior analysis unit (440) models repetitive behaviors that frequently occur in autism spectrum disorder, such as hand waving, repeated left-right shaking of the body, and head tilting, using an LSTM-based time series analysis method (ST441). Then, it detects the examiner's repetitive behavior from the behavior tracking data (ST442).

[0067] Next, periodicity characteristics are quantified by analyzing periodicity, such as changes in joint angles (ST443). At this time, spectrum analysis (FFT) or frequency-based analysis may be applied. Next, an ASD index is calculated based on homologous behavior indicators that exhibit periodicity characteristics (ST444).

[0068] FIG. 5 is a block diagram illustrating the multimodal fusion ASD diagnosis process in the present invention. With reference to FIG. 5, the process of classifying ASD by fusion analysis of multimodal data of eye-tracking data and behavior tracking data is explained.

[0069] The multimodal fusion analysis unit (170) integrates characteristic information generated by the gaze-based ASD analysis unit (150) and the behavior-based ASD analysis unit (160), respectively, to more accurately analyze complex ASD patterns that are difficult to detect with a single modal. In particular, taking into account that gaze information and body movement information have behavioral characteristics that are temporally related, the multimodal fusion analysis unit (170) learns the correlation between them as a high-dimensional feature and classifies the ASD grade.

[0070] Referring to FIG. 5, the multimodal fusion analysis unit (170) is configured to include a gaze encoder (510), a behavior encoder (520), a feature fusion layer (530), a high-dimensional feature learning neural network (540), a feature vector output unit (550), and an ASD classification unit (560).

[0071] The gaze encoder (510) receives gaze tracking data collected from the gaze tracking device (186) and extracts gaze features necessary for ASD analysis. The behavior encoder (520) receives time series data of joint keypoint coordinate values ​​provided by the pose estimation unit (410) or the behavior-based ASD analysis unit (160) and extracts behavior features that represent the examiner's behavior information.

[0072] The feature fusion layer (530) integrates the gaze feature and the behavior feature to generate multimodal features. For example, it may perform concatenation-based fusion by simply connecting the two features to form a single vector, or perform attention-based fusion by using Cross-Attention or Self-Attention to emphasize the correlation between gaze and behavior. As another example, it may perform Tensor Fusion / Bilinear Fusion by combining the inner product (interaction) of the two modals in a tensor form to extract richer interaction features, or perform Gate Mechanism-based fusion by increasing the weight of the more important modal in a specific situation (e.g., increasing the weight of behavior information in a situation of detecting repeated behavior).

[0073] The high-dimensional feature learning neural network (540) receives multimodal features as input and learns a high-level representation sensitive to ASD characteristics. The feature vector output unit (550) inputs the high-dimensional features into an MLP or Transformer structure to produce a final feature vector.

[0074] The ASD classification unit (560) calculates the final feature vector and classifies the ASD grade of the examiner. The aforementioned ASD risk and ASD score are readjusted, and the ASD grade can be determined based on the readjusted ASD risk and ASD score. At this time, the probability of the ASD risk can be calculated using Softmax, Sigmoid, etc. Additionally, the probability of ASD can be output as a probability value based on the ASD indicator.

[0075] The embodiments described in this specification and the accompanying drawings are merely illustrative of a part of the technical concept included in the present invention. Accordingly, since the embodiments disclosed in this specification are intended to explain, not limit, the technical concept of the present invention, it is obvious that the scope of the technical concept of the present invention is not limited by these embodiments. All variations and specific embodiments that can be easily deduced by a person skilled in the art within the scope of the technical concept included in the specification and drawings of the present invention should be interpreted as being included within the scope of the rights of the present invention. Explanation of the symbols

[0076] 100: ASD Analyzer 110: Scenario Video Output Unit 120: Call sound output section 130: Data Collection Unit 140: Trigger Time Synchronization Section 150: Gaze-based ASD Analysis Unit 160: Behavior-based ASD Analysis Department 170: Multimodal Fusion Analysis Unit 182: Display means 184: Speaker 186: Eye tracking device 188: Camera 190: Analysis result output device 410: Pose Estimation Section 420: Name Call Response Analysis Unit 430: Imitative Behavior Analysis Department 440: Stereotyped Behavior Analysis Department 510: Eye-line Encoder 520: Behavior Encoder 530: Feature Fusion Layer 540: Learning Neural Networks 550: MLP / Transformer Feature Vector Calculation Unit 560: ASD Classification Unit

Claims

Claim 1 A method for screening autism spectrum disorder based on scenario video content performed by an ASD analysis device, wherein the ASD analysis device comprises: (a) playing scenario video content including an event that triggers an examiner's behavior; (b) collecting examiner response data including eye-tracking data that tracks gaze and behavior-tracking data that tracks behavior from a video of the examiner while playing the scenario video content; (c) generating a data set aligned on the same time axis by synchronizing the time stamp of the behavior-triggering event with the time stamp of the examiner response data; A method for screening autism spectrum disorder based on scenario video content, characterized by comprising: (d) a step of analyzing the behavioral pattern of the examiner using the data set synchronized in step (c) and calculating an ASD (Autism Spectrum Disorder) index based thereon, wherein step (d) includes: (d-1) a step of stopping the playback of the scenario video content and displaying an attention contrast screen with a simplified background to suppress the examiner's gaze dispersion; (d-2) a step of outputting a pre-recorded call sound while the attention contrast screen is displayed; (d-3) a step of reading the behavior tracking data and observing the reaction delay time at the point where the examiner's face direction faces the point where the call sound is output relative to the point where the call sound is output; and (d-4) a step of quantifying the reaction delay time to adjust the ASD index. Claim 2 A method for screening autism spectrum disorder based on scenario video content according to claim 1, wherein step (d) further comprises: (di) a preprocessing step of performing noise removal, coordinate transformation and normalization, and missing value interpolation on the eye tracking results of the examiner; (d-ii) a step of setting a Region of Interest (ROI) at the point where the behavior-inducing event occurs in the scenario video content; (d-iii) a step of determining whether the examiner gazes at the ROI; (d-iv) a step of extracting and quantifying gaze delay time and duration based on the determination result of step (d-iii), and applying the quantified indicators to a probabilistic learning model to calculate the ASD indicators as probability values. Claim 3 delete Claim 4 A method for screening autism spectrum disorder based on scenario video content according to claim 1, wherein step (d) further comprises: (di) playing an imitation-inducing video designed to have a person within the scenario video content perform a specific movement including time-series coordinates of the head, shoulder, arm, and wrist joints; (d-ii) reading the behavior tracking data and extracting the joint keypoint coordinates of the examiner's head, shoulder, arm, and wrist as time stamps synchronized with the imitation-inducing video; (d-iii) calculating a similarity score by comparing and determining the frame-unit similarity between the time-series coordinates of the imitation-inducing video and the joint keypoint coordinates; and (d-iv) calculating the ASD index by quantifying the examiner's imitation behavior based on the similarity score. Claim 5 A method for screening autism spectrum disorder based on scenario video content according to claim 1, wherein step (d) further comprises: (di) an LSTM modeling step of extracting repetitive behavior features based on joint keypoint time series data of the examiner while playing the scenario video content; (d-ii) a step of detecting the repetitive behavior of the examiner; (d-iii) a step of quantifying the periodicity, frequency, or duration of the repetitive behavior; and (d-iv) a step of calculating the ASD index based on the quantified stereotyped behavior index. Claim 6 A method for screening autism spectrum disorder based on scenario video content, characterized in that, in any one of claims 1 to 5, the method further comprises: (e) a step of inputting the examiner’s eye tracking data and behavior tracking data into an eye encoder and a behavior encoder, respectively, to extract eye features and behavior features; (f) a step of integrating the eye features and behavior features through a Feature Fusion Layer to generate multimodal features; (g) a step of inputting the multimodal features into a learning neural network to learn high-dimensional features; (h) a step of inputting the high-dimensional features into an MLP or Transformer structure to produce a final feature vector; and (i) a step of computing the final feature vector to classify the examiner’s ASD grade.

Citation Information

Patent Citations

  • Autism evaluation system and method based on multi-modal time sequence data fusion

    CN120199486A

  • Autism discrimination apparatus based on user eye monitoring and method thereof

    KR1020250116854A

  • A method for providing content for the treatment of ASD and an electronic device on which such method is implemented

    KR102864047B1

  • A Autism Spectrum Disorder early diagnosis and analysis system

    KR102890486B1