Automatic screening method and system for responsiveness of autistic children based on relative semantic action segmentation, terminal and medium
Through relative semantic action segmentation technology, the time window of the doctor's non-verbal instructions is automatically segmented, and the child's response ability is evaluated by combining multimodal information. This solves the problem of autism screening for children relying on manual methods in existing technologies and achieves efficient and accurate screening results.
Patent Information
- Application Number
- CN202511153304.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-18
AI Technical Summary
In existing technologies, behavioral screening for autistic children relies on manual clinical methods, consumes a large amount of medical resources, is highly subjective, has poor consistency, has a low degree of automation, and has low screening effectiveness and efficiency. It lacks specialized screening technology for responding to non-verbal instructions, has a single modality, and lacks depth and breadth of analysis.
Through a method based on relative semantic action segmentation, we obtain videos of interactions between doctors and children, extract the doctor's skeleton action sequence, segment the non-verbal command time window, use multiple behavioral modalities to evaluate the child's response ability, build an automated screening system, and combine algorithms and rules for screening and evaluation.
It realizes the automated and objective screening of autistic children's ability to respond to non-verbal instructions, reduces manpower input, improves screening efficiency and accuracy, meets early screening needs, adapts to small sample situations, and provides multimodal information support.
Smart Images

Figure CN120636829A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine-assisted screening technology, and in particular to an automatic screening method, system, terminal and medium for the response ability of autistic children based on relative semantic action segmentation. Background Art
[0002] In recent years, breakthroughs in computer technology and artificial intelligence have provided new insights into the diagnosis and treatment of childhood developmental and behavioral disorders. These advances primarily utilize methods such as neuroimaging and visual behavioral analysis to provide systematic, intelligent screening and treatment assistance. Digital phenotyping has the potential to provide an objective and scalable approach to autism screening. Faced with this large patient population, machine-assisted screening and assessment can alleviate the shortage of specialized physicians, reduce the burden of repetitive procedures on medical professionals, and improve the objectivity, standardization, and consistency of clinical diagnosis and treatment.
[0003] However, existing technologies still have the following shortcomings: First, machine-assisted screening for autism is in its early stages, and there is currently no method system for complete behavioral screening of children with autism. Existing technologies are usually used as secondary aids in the screening process or only for cluster research. The actual process relies on clinicians, resulting in strong subjectivity and poor consistency; second, there is a lack of specialized screening technology for the ability to respond to non-verbal instructions. Most technologies only use algorithmic models to learn from different children's behavioral videos to classify and screen out autism symptoms, but lack paradigm scenarios and screening evaluation system methods for specific ability triggers; third, the existing technology has a low degree of automation and cannot achieve automated and integrated screening in screening scenarios, requiring manual video recording. Fourth, the existing technology lacks screening breadth and depth, and only uses a single behavioral modality (expression, behavior, posture, etc.) to conduct behavioral perception analysis of children with autism. It is difficult to provide sufficiently comprehensive information to deeply reveal the complexity and logic of the behavior of autistic patients, which may limit the breadth and depth of the analysis, resulting in limited and misleading analysis results; Fifth, the action segmentation technology used for automated screening performs poorly on autism data with few samples. The existing action segmentation technology uses traditional rule algorithms and machine learning to poor effect, and the use of deep learning relies on the training of a large number of samples, and the effect is poor in the early stage of machine-assisted small samples.
[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0005] The main purpose of this application is to provide an automatic screening method, system, terminal and medium for the response ability of autistic children based on relative semantic action segmentation, aiming to solve the problem in the existing technology that early behavioral auxiliary screening of autistic children requires a large amount of human resources and has poor consistency, resulting in poor screening effect and low efficiency.
[0006] A first aspect of an embodiment of the present application provides an automatic screening method for the responsiveness of autistic children based on relative semantic action segmentation. The automatic screening method for the responsiveness of autistic children based on relative semantic action segmentation includes the following steps: obtaining video information of the interaction process between a doctor and a child; extracting the doctor's skeleton action sequence from the video information; segmenting the skeleton action sequence to obtain a time window in which the doctor issues an instruction; obtaining the child's behavioral modal information based on the time window; and evaluating the child's responsiveness to non-verbal instructions based on the behavioral modal information to obtain screening evaluation information.
[0007] Optionally, in one embodiment of the present application, extracting the doctor's skeleton action sequence from the video information specifically includes: preprocessing the video information to obtain a preprocessed video; separating the doctor's red, green, and blue image streams from the preprocessed video; identifying the red, green, and blue image streams to obtain the doctor's skeleton information in each frame of the image; and arranging the skeleton information of each frame in time sequence to obtain the doctor's skeleton action sequence.
[0008] Optionally, in one embodiment of the present application, the segmentation of the skeleton action sequence to obtain the time window in which the doctor issues the instruction specifically includes: constructing a spatial and temporal model, inputting the skeleton action sequence into the spatial and temporal model, and obtaining a frame-level feature representation of the doctor; obtaining a real frame-level label, and obtaining a frame-level action label based on the frame-level feature representation and the real frame-level label; and obtaining the time window in which the doctor issues the action instruction based on the frame-level action label.
[0009] Optionally, in one embodiment of the present application, the spatial and temporal model includes a joint spatial model and an inter-frame temporal model; the spatial and temporal model is constructed, and the skeleton action sequence is input into the spatial and temporal model to obtain a frame-level feature representation of the doctor, specifically including: obtaining joint text, and inputting the joint text into a large language model to obtain joint text features; obtaining a relative semantic graph of the joint based on the joint text features; establishing a spatial model based on a graph convolutional network, and fusing the relative semantic graph with the spatial model to obtain the joint spatial model; inputting the skeleton action sequence into the joint spatial model to obtain spatial features; establishing an inter-frame temporal model based on a neural network, and inputting the spatial features into the inter-frame temporal model to obtain a frame-level feature representation of the doctor.
[0010] Optionally, in one embodiment of the present application, the frame-level action label is obtained based on the frame-level feature representation and the real frame-level label, which specifically includes: obtaining action text and inputting the action text into a large language model to obtain action text features; obtaining an action semantic graph based on the action text features; supervising the frame-level feature representation based on the action text features and the action semantic graph to obtain a supervised frame-level feature representation; and obtaining a frame-level action label based on the frame-level feature representation and the supervised frame-level feature representation.
[0011] Optionally, in one embodiment of the present application, the behavioral modal information includes skeletal posture, gestures and key object positions; obtaining the behavioral modal information of the child based on the time window specifically includes: cropping the video information according to the time window to obtain a target video; obtaining the child's skeletal posture, gestures and key object positions based on the target video.
[0012] Optionally, in one embodiment of the present application, the child's ability to respond to non-verbal instructions is evaluated based on the behavioral modality information to obtain screening evaluation information, specifically including: constructing a response ability model, and training the response ability model according to preset training data to obtain a trained response ability model; inputting the skeleton posture, the gesture and the key object position into the trained response ability model to obtain the child's screening evaluation information.
[0013] A second aspect of the present application further provides an automatic screening system for the responsiveness of autistic children based on relative semantic action segmentation, wherein the automatic screening system for the responsiveness of autistic children based on relative semantic action segmentation comprises: Video acquisition module, used to obtain video information of the interaction process between doctors and children; A skeleton sequence extraction module, configured to extract the doctor's skeleton action sequence from the video information; An action segmentation module, configured to segment the skeleton action sequence to obtain a time window in which the doctor issues an instruction; a modality perception module, configured to obtain behavioral modality information of the child according to the time window; The screening and evaluation module is used to evaluate the child's ability to respond to non-verbal instructions based on the behavioral modality information to obtain screening and evaluation information.
[0014] The third aspect of an embodiment of the present application also provides a terminal, wherein the terminal includes: a memory, a processor, and an automatic screening program for the response ability of autistic children based on relative semantic action segmentation, which is stored on the memory and can be run on the processor. When the automatic screening program for the response ability of autistic children based on relative semantic action segmentation is executed by the processor, the steps of the automatic screening method for the response ability of autistic children based on relative semantic action segmentation are implemented as described above.
[0015] The fourth aspect of an embodiment of the present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores an automatic screening program for the response ability of autistic children based on relative semantic action segmentation, and when the automatic screening program for the response ability of autistic children based on relative semantic action segmentation is executed by a processor, the steps of the automatic screening method for the response ability of autistic children based on relative semantic action segmentation as described above are implemented.
[0016] Beneficial effects: The present application provides an automatic screening method, system, terminal and medium for the response ability of autistic children based on relative semantic action segmentation. The present application segments the doctor's skeleton action sequence into a time window for initiating non-verbal instructions, extracts the child's multiple behavioral modes within the time window, and thus screens and evaluates the child's response ability under non-verbal instructions. It can realize the automated and objective screening of the response ability of autistic children to non-verbal instructions, realize the automation and integration of the screening process, reduce manpower input, improve screening efficiency, and improve the accuracy and breadth of screening. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1This is a flowchart of a preferred embodiment of the method for automatically screening the response ability of autistic children based on relative semantic action segmentation of the present application; Figure 2 This is a system diagram of a preferred embodiment of the method for automatically screening the response ability of autistic children based on relative semantic action segmentation of the present application; Figure 3 This is a skeleton action segmentation model diagram based on relative semantic enhancement in a preferred embodiment of the automatic screening method for the response ability of autistic children based on relative semantic action segmentation of the present application; Figure 4 This is a structural diagram of a preferred embodiment of the automatic screening system for the response ability of autistic children based on relative semantic action segmentation of the present application; Figure 5 This is a structural diagram of a preferred embodiment of the terminal of this application.
[0019] Description of reference numerals: 100. Video acquisition module; 200. Skeleton sequence extraction module; 300. Action segmentation module; 400. Modal perception module; 500. Screening and evaluation module. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solutions and effects of this application clearer and more specific, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. The described embodiments are only possible technical implementations of this application and are not all possible implementations. Based on the embodiments in this application, those skilled in the art can fully combine the embodiments of this application to obtain other embodiments without creative work, and these embodiments are also within the scope of protection of this application.
[0021] Autism Spectrum Disorder (ASD) is a neurodevelopmental disorder that develops in early childhood. Its clinical symptoms primarily include social communication impairments, stereotyped behaviors, and restricted interests. Autism is currently the most common behavioral developmental disorder in children, with the highest prevalence and largest number of patients. The cause of autism is unknown, and there is no specific treatment available. Children with autism cannot recover in adulthood, resulting in a high disability rate. This profoundly impacts patients' daily quality of life and long-term social outcomes, placing a significant burden on families and society. Research has found that "early screening, early diagnosis, and early intervention" are the consensus among experts in the field of childhood developmental and behavioral disorders. Only through early behavioral screening can symptoms and potential risks of autism be diagnosed, allowing behavioral interventions to address the many behavioral, functional, social, and cognitive deficits of children with ASD. The earlier the identification and intervention, the more effective the outcome. Therefore, early behavioral identification, screening, and analysis of autism are crucial for alleviating symptoms and improving the quality of life of children with autism.
[0022] Existing screening and diagnosis of developmental behavioral disorders in children with autism mainly rely on behavioral observation and assessment by manual clinicians, that is, they rely on professionals to make medical diagnoses through behavioral observations and formulate corresponding intervention plans. Because this process relies on experienced clinicians, it consumes a large amount of medical human resources. The long training period and the limited number of rehabilitation professionals make it difficult to meet the growing demand for intervention and assessment services. At the same time, the results of manual assessments are based to a certain extent on the subjective feelings of professionals and lack systematic quantitative assessments. They are highly subjective and have poor consistency. Therefore, the problems of scarce medical human resources, highly subjective assessments, and lack of systematic plans in manual clinical methods have hindered the diagnosis and treatment process of children with autism to a certain extent.
[0023] First, the nouns involved in the embodiments of this application are introduced: ASD (Autism Spectrum Disorder), autism spectrum disorder, also known as autism; STAS (Skeleton-based Temporal Action Segmentation) uses skeleton sequences to segment the actions of a person over a period of time, i.e. skeleton temporal action segmentation; Transformer, a neural network architecture based on attention modeling.
[0024] The following describes the automatic screening method, system, terminal and medium for the response ability of autistic children based on relative semantic action segmentation of the embodiment of the present application with reference to the accompanying drawings. In view of the problem that the early behavioral auxiliary screening of autistic children in the above-mentioned related art requires a large amount of human resources and has poor consistency, resulting in poor screening effect and low efficiency, the present application provides an automatic screening method for the response ability of autistic children based on relative semantic action segmentation. In this method, the time window initiated by non-verbal instructions is segmented from the doctor's skeleton action sequence, and the child's multiple behavioral modes are extracted within the time window, thereby screening and evaluating the child's response ability under non-verbal instructions, which can realize the automation and objective screening of the response ability of autistic children to non-verbal instructions, realize the automation and integration of the screening process, reduce manpower input, improve screening efficiency, and improve the accuracy and breadth of screening. Thus, the technical problem that the early behavioral auxiliary screening of autistic children in the related art requires a large amount of human resources and has poor consistency, resulting in poor screening effect and low efficiency is solved.
[0025] At present, early behavioral screening of children with autism mainly relies on manual clinical methods, which require a large amount of medical human resources, and are highly subjective and inconsistent. These problems urgently need to be solved from a machine-assisted level. However, machine-assisted screening technology is in its early stages, and has technical problems such as poor ability screening specialization, low degree of automation, single modality, and poor effect. In order to solve the above technical problems, this application targets the "no or little response" in the "five no" behavioral standards widely accepted in the field of autism in China. In this application, a screening paradigm for the response ability based on non-verbal instructions of the interaction between doctors and children is used to induce children to make corresponding response behaviors. During the paradigm process, the camera records the interaction video between the child and the doctor for automated screening; in addition, this application is based on an automated response ability screening model based on relative semantic action segmentation. First, the target detection tracking and human posture estimation algorithm are used to obtain the human skeleton points of the doctor in each frame; then, based on the relative semantic action segmentation model, it is used to segment the doctor's key non-verbal instruction actions with high precision; a time window is set according to the starting and ending points of the doctor's key non-verbal instruction actions, and the target detection algorithm is used to obtain the position information of the doctor's hand, the child's hand and the toy, and then the spatial position calculation is used to determine whether the child makes a corresponding response behavior after hearing the instruction, thereby making an ability screening assessment.
[0026] This application can realize the automated screening of autistic children's ability to respond to non-verbal instructions. It only requires an interventionist to accompany the child to give instructions, and no professional physician is required to conduct behavioral assessment screening, thereby saving a lot of medical resources and providing quantitative and objective screening indicators. Compared with other technologies, this application is fully capable of automatically and integratedly realizing the segmentation of instructional behavior and screening and evaluation of response ability. The action segmentation algorithm uses relative semantics to enhance the reliability of small samples, and makes full use of behavioral modalities such as skeletons, objects, and gestures. At the same time, an algorithm + rule screening combination is proposed, which has high robustness and reliability, can be implemented on a large scale, can promote the development of autism auxiliary diagnosis and treatment theory and technology, and is of great significance to promoting the progress of the autism rehabilitation industry.
[0027] The following specific embodiments are used to describe the technical solution of the present application in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0028] The automatic screening method for the response ability of autistic children based on relative semantic action segmentation described in the preferred embodiment of this application is as follows: Figure 1 As shown, the automatic screening method for the response ability of autistic children based on relative semantic action segmentation includes the following steps: In step S101, video information of the interaction process between the doctor and the child is obtained.
[0029] It should be noted that this technical solution is divided into two parts: a screening paradigm and a screening system. The screening paradigm, based on the ability to respond to non-verbal instructions in the interaction between doctors and children, is used to establish a structured interactive scenario in which doctors and children respond to non-verbal instructions, induce children to make expected response behaviors through the doctor's non-verbal instructions, and collect videos of the scene throughout the process for subsequent screening. The screening system is an automated responsiveness screening system based on relative semantic action segmentation. It is used to extract multimodal behavioral information from the video, and uses the designed relative semantic action segmentation algorithm to segment the doctor's instruction window according to skeleton information. The child's responsiveness is evaluated and analyzed based on the child's multi-behavioral modal information in the window, thereby completing the responsiveness screening of children with autism.
[0030] See also Figure 2The screening paradigm for responsiveness to nonverbal cues in physician-child interactions provides a face-to-face scenario for physician-child interaction. The responsiveness to nonverbal cues screening process consists of two rounds. Each round involves attracting the child's attention with a different toy. Once the child is engaged in play, the physician issues a nonverbal cue requesting the toy and observes whether the child's behavior responds to the physician's instructions. The specific process is as follows: The physician sits opposite the child and takes out two small balls and places them on the table. The physician rolls the balls to the child on the table (saying "ball or ball") and allows the child to play freely. After 2-3 minutes of play, the first round of testing begins. The physician extends his hand and prompts, "Give the ball to auntie / uncle." The physician then takes out a wind-up toy (a turtle) and places it on the table. The physician winds the tortoise up twice and places it on the table, allowing the child to touch it. After 2-3 minutes of play, the first round of testing begins. The physician extends his hand and prompts, "Give the ball to auntie / uncle." The process concludes, and the physician puts the toy away.
[0031] Specifically, video acquisition is performed, and a camera fixed in the scene records the face-to-face interaction process between the doctor and the child from the side, recording the entire process at 30FPS or 60FPS for subsequent screening.
[0032] In step S102, the skeleton action sequence of the doctor is extracted from the video information.
[0033] In one possible implementation, the video information is preprocessed to obtain a preprocessed video; the red, green, and blue image streams of the doctor are separated from the preprocessed video; the red, green, and blue image streams are identified to obtain the skeleton information of the doctor in each frame of the image; and the skeleton information of each frame is arranged in time sequence to obtain the skeleton action sequence of the doctor.
[0034] See also Figure 2 The automated responsiveness screening system, based on relative semantic action segmentation, records the entire interaction between doctors and children during the screening process. Using a multi-behavioral modality perception algorithm, it extracts information such as the doctor's and child's skeletons and gestures. This skeletal information is then fed into a relative semantic action segmentation algorithm to detect the timing of the doctor's nonverbal commands during the screening process (i.e., the testing phase of the screening process), thereby segmenting the screening time window. The system then assesses the child's responsiveness to nonverbal commands based on the child's multi-behavioral modality information within the time window, achieving automated screening for the "no or little response" component of the "five no" behavioral criteria.
[0035] Specifically, during the doctor skeleton sequence extraction process, existing object detection and tracking algorithms are used to separate the doctor's RGB stream from the video. Existing human pose estimation algorithms (Openpose) are then used to further extract the doctor's skeleton sequence for subsequent automated evaluation window segmentation. In other words, the goal of this step is to establish a structured interactive scenario in which the doctor and child respond to non-verbal cues, thereby obtaining video data of the doctor-child interaction process.
[0036] Furthermore, the full video obtained from the screening paradigm process is first preprocessed, including adjusting the video frame rate (such as ensuring it is 30FPS or 60FPS) and adjusting the video size to accommodate subsequent algorithm processing requirements. Using existing target detection and tracking algorithms, the doctor's red, green, and blue (RGB) image stream is separated from the preprocessed video. The purpose of this step is to separate the doctor's video information from the child and other background information, so that subsequent analysis can focus on the doctor's movements. The separated doctor's RGB image stream is input into an existing human pose estimation algorithm (such as Openpose), which can identify and extract the doctor's skeleton information in each frame, including the location of joint points and the connection between joints. The skeleton information of each frame output by the algorithm is arranged in time sequence to generate a skeleton sequence of the doctor.
[0037] In step S103, the skeleton action sequence is segmented to obtain the time window in which the doctor issues the instruction.
[0038] In one possible implementation, a spatial and temporal model is constructed, and the skeleton action sequence is input into the spatial and temporal model to obtain a frame-level feature representation of the doctor; a true frame-level label is obtained, and a frame-level action label is obtained based on the frame-level feature representation and the true frame-level label; and a time window in which the doctor issues an action instruction is obtained based on the frame-level action label.
[0039] It should be noted that in the process of automatically segmenting the non-verbal instruction time window, this application uses the doctor's skeleton sequence as input through an action segmentation algorithm based on relative semantics, thereby temporally segmenting all the actions made by the doctor in the entire process (no action, taking out toys, reaching out to initiate a request, and other actions). The action of the doctor reaching out to initiate a request for a toy is set as the time window for evaluating the child's response ability, so as to achieve an automated identification and screening time window. Among them, regarding the time window for automatically segmenting non-verbal instructions, the skeleton temporal action segmentation used is an action segmentation algorithm based on relative semantic enhancement, which can reduce the dependence on sample size during the training model process, enhance its clustering and discrimination between different actions, and be more conducive to feature learning, thereby improving the model's segmentation accuracy for the doctor's actions.
[0040] In one possible implementation, the spatial and temporal model includes a joint space model and an inter-frame temporal model. Joint text is obtained and input into a large language model to obtain joint text features; a relative semantic graph of the joints is obtained based on the joint text features; a spatial model is established based on a graph convolutional network, and the relative semantic graph is fused with the spatial model to obtain the joint space model; the skeleton action sequence is input into the joint space model to obtain spatial features; an inter-frame temporal model is established based on a neural network, and the spatial features are input into the inter-frame temporal model to obtain a frame-level feature representation of the doctor.
[0041] like Figure 3 As shown, the large language model BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model based on the Transformer architecture. A doctor's skeleton sequence is input into the model, which first uses a graph convolutional network to initially establish relationships between joints based on the skeleton's topological connections. The model then uses text descriptions of the joints as input to the large language model, outputting text features for each joint. By calculating the relationships between text features, a relative semantic graph of the joints is generated, which is used for spatial modeling to further enhance the establishment of fine-grained relationships between joints. The model then uses a Transformer network to establish temporal relationships between features across frames, ultimately outputting a feature representation of the doctor's action sequence.
[0042] Furthermore, a graph convolutional network is used to preliminarily establish the relationship between joints based on the topological connection of the skeleton. The purpose of this step is to capture the spatial relationship between joints and provide a basis for subsequent feature extraction and action segmentation; the descriptive text of the joints is input into the large language model to output the text features of each joint. These text features may contain information such as the function, position or relationship of the joint with other joints, which helps to enhance the semantic understanding ability of the model; the relative semantic graph of the joints is obtained by calculating the relationship between text features. This step uses the text features of the joints to capture the relative semantic relationship between them, further enhancing the model's understanding of the fine-grained relationship between joints (that is, the relative semantic graph is integrated with the spatial model to obtain the joint space model); the Transformer network is used to establish the inter-frame temporal relationship of features. The purpose of this step is to capture the temporal continuity of the action, combine the spatial relationship and relative semantic relationship of the joints with the temporal information, and form a complete action feature representation.
[0043] In one possible implementation, an action text is obtained and input into a large language model to obtain action text features; an action semantic graph is obtained based on the action text features; the frame-level feature representation is supervised based on the action text features and the action semantic graph to obtain a supervised frame-level feature representation; and a frame-level action label is obtained based on the frame-level feature representation and the supervised frame-level feature representation.
[0044] like Figure 3 As shown, the model requires supervised training with a certain number of samples to achieve action segmentation. Therefore, at the feature representation level, this model uses action text input into a large language model to output text features for each action category. These action text features are then compared with the features of the corresponding action segments in the feature representation for supervision. The relative relationships between actions are further calculated to generate an action relative semantic graph, which is used to supervise the relationship distances between action features. Finally, the feature representation is input into the classification head to obtain action labels for each frame. This action label is then compared with the real human-labeled labels using a cross-entropy loss for classification supervision. During the training phase, a total loss function is generated through these three supervision methods. This loss function is then used to optimize the network model through backpropagation using an optimizer to achieve the training goal. With a certain number of samples, the final model can achieve higher accuracy than other action segmentation algorithms. Skeleton sequences can be directly input and the model can infer action labels for each frame, effectively segmenting the time and location of each action performed by the doctor, thereby automatically identifying non-verbal instruction time windows.
[0045] Furthermore, at the feature representation level, the action text is input into the large language model to output the features of each type of action text, and compared with the features of the corresponding action segment in the feature representation for supervision. At the same time, the relative relationship between actions is calculated to obtain the action relative semantic graph (i.e. Figure 3 The trained model performs action segmentation on the input skeleton sequence and outputs action labels for each frame. These labels indicate the time and action of the doctor. Among the segmented actions, the doctor's reaching out to request a toy is identified and used as the time window for evaluating the child's responsiveness.
[0046] In step S104, the behavioral modality information of the child is obtained according to the time window.
[0047] In one possible implementation, the behavioral modality information includes skeletal posture, hand gestures, and key object positions. The video information is cropped according to the time window to obtain a target video; and the skeletal posture, hand gestures, and key object positions of the child are obtained based on the target video.
[0048] Specifically, during the process of perceiving children's multi-modal behaviors, the video within the automatically segmented time window is cropped out, resulting in separate videos of the moments when the doctor expresses non-verbal instructions (corresponding to two separate tests in the screening paradigm). High-level modal information, such as the child's skeletal posture, gestures, and key objects, is extracted from these two videos. This involves extracting the child's posture through human posture estimation, extracting gestures through gesture recognition, and extracting the positional information of the ball and clockwork toy through target detection and tracking. This diverse information can inform subsequent assessments of children's responses to non-verbal instructions. Utilizing high-level, easily understood behavioral information such as extracted skeletons, gestures, and key objects is more consistent with human perception and more reliable and interpretable than single-shot evaluations that directly input the video into the model.
[0049] Furthermore, based on the automatically segmented non-verbal instruction time windows, corresponding clips were cropped from the fully recorded video. These clips contain the complete interactive scene when the doctor issued non-verbal instructions. The cropped video was processed using a human posture estimation algorithm to extract the child's skeleton sequence, including joint position and posture information. A gesture recognition algorithm was applied to identify and extract the child's gestures in the video. Gesture information, as one of the important behavioral modalities, can reflect the child's reaction and intention after hearing the instructions. An object detection and tracking algorithm was used to detect and track key objects in the video (such as small balls and clockwork toys), extracting their location information and motion trajectory to analyze the child's interaction with the key objects. The extracted child's skeleton posture, gestures, and key object information complement each other and together constitute a comprehensive description of the child's response behavior.
[0050] In step S105, the child's ability to respond to non-verbal instructions is evaluated based on the behavioral modality information to obtain screening evaluation information.
[0051] In one possible implementation, a responsiveness model is constructed, and the responsiveness model is trained according to preset training data to obtain a trained responsiveness model; the skeleton posture, the gesture, and the position of the key object are input into the trained responsiveness model to obtain screening evaluation information of the child.
[0052] Specifically, based on high-level modal information such as the child's skeletal posture, gestures, and key objects, a screening and assessment is conducted to determine whether the child responds accordingly to non-verbal instructions. In terms of rules, the spatial position is directly used to calculate whether the child's hand is extended and whether the toy reaches the doctor's hand, to preliminarily determine whether the child has handed the toy to the doctor. In terms of algorithms, through a multi-layer convolutional neural network and feature fusion, all posture, gesture, and key object category position information are input into the network for feature fusion and modeling, and the model is trained based on a certain number of samples and labels indicating whether or not to respond, so as to intelligently determine whether the child has responded. In the actual screening process, the child's response ability is comprehensively judged based on screening at two levels: rules and algorithms. The combination of the two makes full use of the intelligence of the algorithm and the reliability, interpretability, and robustness of the rule-based approach, enabling stable screening research and further implementation.
[0053] Furthermore, the extracted information about the child's skeletal posture, gestures, and the categories and locations of key objects (such as toys) are input into a multi-layer convolutional neural network for feature fusion and modeling. This step aims to extract the characteristics of children's behavior from multiple dimensions and construct a model that can reflect the child's response ability. Based on the trained model, an intelligent judgment is made on whether the child responds, that is, by inputting the extracted features into the model's classification head to obtain a prediction result on whether the child responds.
[0054] It is worth noting that the existing technology is in its early stages, and its ability assessment is not fine-grained enough. It usually trains models based on behavioral data collected in specific scenarios to achieve classification screening for autism, but lacks screening and assessment of specific abilities, resulting in poor interpretability. This application is aimed at a screening paradigm and system specialized for response ability, and only targets the "no / little response" ability in the "five no" behavioral standards. A specific behavior induction paradigm is designed to specifically assess the ability of autistic children to respond to non-verbal instructions, which is highly targeted and fine-grained. The existing screening technology has a low degree of automation, and there is no method to directly perform automated full-process evaluation based on collected videos. Manual window segmentation is often required. This application uses action segmentation tools to combine action segmentation algorithms with autism screening, and automatically determines the evaluation time window by segmenting non-verbal behavioral instructions initiated by doctors. This can reduce the manpower input in the process and is efficient and convenient. Current skeleton-based action segmentation algorithms have significant room for improvement in accuracy, but they also require a high sample size, have low similarity between action classes, and have fuzzy boundaries. To address these issues, this application proposes an action segmentation algorithm assisted by relative semantic graphs. This algorithm uses language modalities to assist the model's feature learning, thereby achieving higher segmentation accuracy and significantly reducing sample dependency, adapting to the problem of insufficient samples in initial screening. Current technologies often use a single modality to screen children for autism, such as using only gaze to assess children's attention. However, the lack of breadth in single-modality design makes it difficult to fully establish a representation of children's behavior. This application combines multiple behavioral modalities, using skeleton modalities to segment windows, and utilizing gestures, skeletons, and objects to achieve reliable, interpretable, and high-precision screening. Most existing screening models only use machine learning and deep learning, which have poor reliability and robustness. Similarly, the analysis based on pure rules has a low level of intelligence. This application adopts a two-stage algorithm + rule screening process. The action segmentation algorithm identifies the indicator window, and the algorithm + rule combination model further assesses the child's ability. The two-stage process is highly engineered and reliable, while the algorithm + rule evaluation is highly reliable and has a low error rate.
[0055] Next, an automatic screening system for the response ability of autistic children based on relative semantic action segmentation proposed in an embodiment of the present application will be described with reference to the accompanying drawings.
[0056] Figure 4 This is a structural diagram of an automatic screening system for the response ability of autistic children based on relative semantic action segmentation according to an embodiment of the present application.
[0057] like Figure 4 As shown, the automatic screening system for the response ability of autistic children based on relative semantic action segmentation includes: a video acquisition module 100, a skeleton sequence extraction module 200, an action segmentation module 300, a modality perception module 400 and a screening evaluation module 500.
[0058] Specifically, the video acquisition module 100 is used to acquire video information of the interaction process between the doctor and the child; A skeleton sequence extraction module 200 is used to extract the skeleton action sequence of the doctor from the video information; An action segmentation module 300 is used to segment the skeleton action sequence to obtain a time window for the doctor to issue an instruction; A modality perception module 400 is configured to obtain behavioral modality information of the child according to the time window; The screening and evaluation module 500 is configured to evaluate the child's ability to respond to non-verbal instructions based on the behavioral modality information to obtain screening and evaluation information.
[0059] Figure 5 This is a diagram of the structure of a terminal provided in an embodiment of the present application. The terminal may include: Memory 501 , processor 502 , and computer programs stored in the memory 501 and executable on the processor 502 .
[0060] When the processor 502 executes the program, the automatic screening method for the response ability of autistic children based on relative semantic action segmentation provided in the above embodiment is implemented.
[0061] Furthermore, the terminal further includes: The communication interface 503 is used for communication between the memory 501 and the processor 502 .
[0062] The memory 501 is used to store computer programs that can be run on the processor 502 .
[0063] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0064] If the memory 501, processor 502, and communication interface 503 are implemented independently, the communication interface 503, memory 501, and processor 502 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EIS) bus. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 5 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0065] Optionally, in a specific implementation, if the memory 501, the processor 502 and the communication interface 503 are integrated on a chip, the memory 501, the processor 502 and the communication interface 503 can communicate with each other through an internal interface.
[0066] The processor 502 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0067] This embodiment also provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the above-mentioned method for automatically screening the response ability of autistic children based on relative semantic action segmentation is implemented.
[0068] One embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the Figure 1 Any of the corresponding embodiments provides an automatic screening method for the response ability of autistic children based on relative semantic action segmentation.
[0069] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to user analysis data, user storage data, user display data, etc.) and signals involved in the present invention are all information, data and signals authorized by the user or fully authorized by all parties; and the collection, use and processing of relevant information, data and signals comply with the laws, regulations and standards of relevant countries and regions.
[0070] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0071] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this application, "N" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0072] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing a custom logical function or process step, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed in a different order than shown or discussed, including performing functions in a substantially simultaneous manner or in a reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0073] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable storage medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable storage medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (not exhaustive) of computer-readable storage media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable storage medium may even be paper or other suitable medium on which the program is printed, since the program can be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing it in other suitable ways as necessary, and then storing it in a computer memory.
[0074] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiment, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logical functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.
[0075] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0076] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0077] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
[0078] It should be understood that the application of this application is not limited to the above examples. For ordinary technicians in this field, they can make improvements or changes based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to this application.
[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.
Claims
1. An automatic screening method for the response ability of autistic children based on relative semantic action segmentation, characterized in that: The automatic screening method for the response ability of autistic children based on relative semantic action segmentation includes: Obtain video information of the interaction between doctors and children; extracting a skeleton action sequence of the doctor from the video information; Segmenting the skeleton action sequence to obtain a time window for the doctor to issue an instruction; Obtaining behavioral modality information of the child according to the time window; The child's ability to respond to non-verbal instructions is evaluated based on the behavioral modality information to obtain screening evaluation information.
2. The automatic screening method for the response ability of autistic children based on relative semantic action segmentation according to claim 1, characterized in that: The step of extracting the doctor's skeleton action sequence from the video information specifically includes: Preprocessing the video information to obtain a preprocessed video; Separating the doctor's red, green and blue image streams from the preprocessed video; Identify the red, green, and blue image streams to obtain skeleton information of the doctor in each frame of the image; Arrange the skeleton information of each frame in time sequence to obtain the skeleton action sequence of the doctor.
3. The automatic screening method for the response ability of autistic children based on relative semantic action segmentation according to claim 1, characterized in that: The segmenting of the skeleton action sequence to obtain the time window for the doctor to issue the instruction specifically includes: Constructing a spatial and temporal model, inputting the skeleton action sequence into the spatial and temporal model to obtain a frame-level feature representation of the doctor; Obtaining a true frame-level label, and obtaining a frame-level action label based on the frame-level feature representation and the true frame-level label; According to the frame-level action label, a time window in which the doctor issues an action instruction is obtained.
4. The automatic screening method for the response ability of autistic children based on relative semantic action segmentation according to claim 3, characterized in that: The space and time series model includes a joint space model and an inter-frame time series model; The step of constructing a spatial and temporal model, inputting the skeleton action sequence into the spatial and temporal model, and obtaining a frame-level feature representation of the doctor specifically includes: Acquire joint text, and input the joint text into a large language model to obtain joint text features; According to the joint text features, the relative semantic graph of the joint is obtained; Establishing a spatial model based on a graph convolutional network, and fusing the relative semantic graph with the spatial model to obtain the joint space model; Inputting the skeleton action sequence into the joint space model to obtain spatial features; An inter-frame temporal model is established based on a neural network, and the spatial features are input into the inter-frame temporal model to obtain a frame-level feature representation of the doctor.
5. The automatic screening method for the response ability of autistic children based on relative semantic action segmentation according to claim 3, characterized in that: Obtaining a frame-level action label according to the frame-level feature representation and the true frame-level label specifically includes: Obtaining action text, and inputting the action text into a large language model to obtain action text features; Obtaining an action semantic graph according to the action text features; Supervising the frame-level feature representation according to the action text feature and the action semantic graph to obtain a supervised frame-level feature representation; A frame-level action label is obtained according to the frame-level feature representation and the supervised frame-level feature representation.
6. The automatic screening method for the response ability of autistic children based on relative semantic action segmentation according to claim 1, characterized in that: The behavioral modality information includes skeleton posture, gestures and key object positions; Obtaining the behavioral modality information of the child according to the time window specifically includes: Cropping the video information according to the time window to obtain a target video; According to the target video, the skeleton posture, gestures and key object positions of the child are obtained.
7. The method for automatically screening the response ability of autistic children based on relative semantic action segmentation according to claim 6, characterized in that: The step of evaluating the child's ability to respond to non-verbal instructions based on the behavioral modality information to obtain screening evaluation information specifically includes: Constructing a responsiveness model, and training the responsiveness model according to preset training data to obtain a trained responsiveness model; The skeleton posture, the hand gesture and the key object position are input into the trained response ability model to obtain screening assessment information of the child.
8. An automatic screening system for the response ability of autistic children based on relative semantic action segmentation, characterized by: The automatic screening system for the response ability of autistic children based on relative semantic action segmentation includes: Video acquisition module, used to obtain video information of the interaction process between doctors and children; A skeleton sequence extraction module, configured to extract the doctor's skeleton action sequence from the video information; An action segmentation module, configured to segment the skeleton action sequence to obtain a time window in which the doctor issues an instruction; a modality perception module, configured to obtain behavioral modality information of the child according to the time window; The screening and evaluation module is used to evaluate the child's ability to respond to non-verbal instructions based on the behavioral modality information to obtain screening and evaluation information.
9. A terminal, characterized in that: The terminal includes: a memory, a processor, and an automatic screening program for the response ability of autistic children based on relative semantic action segmentation, which is stored in the memory and can be run on the processor. When the automatic screening program for the response ability of autistic children based on relative semantic action segmentation is executed by the processor, the steps of the automatic screening method for the response ability of autistic children based on relative semantic action segmentation as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores an automatic screening program for the response ability of autistic children based on relative semantic action segmentation. When the automatic screening program for the response ability of autistic children based on relative semantic action segmentation is executed by the processor, the steps of the automatic screening method for the response ability of autistic children based on relative semantic action segmentation as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Primary screening device for autism based on non-social sound stimulus behavior paradigm
CN109431523A
Abnormal behavior detection method for autistic children based on finger object identification and sight tracking
CN118658604A
Hand action mode training method for rehabilitation training of autistic children
CN119361077A