Autism children response ability automatic screening method, system, terminal and medium based on relative semantic action segmentation
By using a method based on relative semantic action segmentation, the non-verbal instruction response ability of children with autism is automatically screened, which solves the problems of high resource consumption and strong subjectivity in the manual screening of existing technologies, and achieves efficient and accurate screening results.
Patent Information
- Application Number
- CN202511153304.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-08-18
AI Technical Summary
Current technologies for screening the behavior of children with autism rely on manual clinical methods, which consume a lot of medical human resources. They are highly subjective, inconsistent, have low automation, poor screening results and low efficiency. They also lack specialized screening technologies that can respond to non-verbal instructions, have a single modality, and lack breadth and depth of analysis.
By using a relative semantic action segmentation method, we acquire videos of doctor-child interactions, extract the doctor's skeletal action sequences, segment the time windows of non-verbal instructions, assess children's responsiveness using multi-behavioral modalities, construct an automated screening system, and combine algorithms and rules for screening and assessment.
It enables automated screening of nonverbal responsiveness in children with autism, reducing manpower input, improving screening efficiency and accuracy, providing quantitative and objective screening indicators, and meeting the needs of early behavioral screening.
Smart Images

Figure CN120636829B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of machine-assisted screening, in particular to an autism child response ability automatic screening method, system, terminal and medium based on relative semantic action segmentation. BACKGROUND
[0002] In recent years, breakthroughs in computer technology and artificial intelligence have provided new ideas for the diagnosis and treatment of children's developmental behavior disorders, mainly through methods based on neuroimaging or visual behavior analysis to provide systematic and intelligent auxiliary screening and diagnosis. Digital phenotype analysis can provide objective and scalable methods for autism screening. In the face of a large patient population, machine-assisted screening and evaluation can alleviate the shortage of professional physicians, reduce the burden of repetitive operations by medical professionals, and improve the objectivity, standardization and consistency of clinical diagnosis and treatment.
[0003] However, the prior art still has the following disadvantages: first, machine-assisted screening for autism is in its early stages, and there is currently no complete method system for screening the behavior of autistic children. The existing technology is usually used as a secondary aid in the screening process or only for clustering research, and the actual process relies on clinicians, resulting in strong subjectivity and poor consistency; second, there is a lack of specialized screening technology for non-verbal response ability. Most technologies only use algorithm models to learn different child behavior videos to classify and screen autism symptoms, but lack of paradigm scenes triggered by specific abilities and screening evaluation system methods; third, the existing technology has low automation, cannot achieve automatic and integrated screening in the screening scene, and needs manual recording and segmentation of videos, lacks automatic behavior segmentation analysis tools, and has low automation; fourth, the screening breadth and depth of the existing technology is insufficient, only a single behavior modality (expression, behavior, posture, etc.) is used for behavior perception analysis of autistic children, which makes it difficult to provide comprehensive information to reveal the complexity and logic of autistic patient behavior, which may limit the breadth and depth of analysis, resulting in limited and misleading analysis results; fifth, the action segmentation technology used for automatic screening performs poorly on small sample data of autism. The existing action segmentation technology uses traditional rule algorithms and machine learning, which performs poorly, while deep learning relies on a large number of sample training, which performs poorly in the early stage of machine assistance with small samples.
[0004] Therefore, the prior art still needs to be improved and developed. SUMMARY
[0005] The main purpose of the present application is to provide an autism child response ability automatic screening method, system, terminal and medium based on relative semantic action segmentation, aiming at solving the problem that the existing technology needs a large amount of human resources for early behavior auxiliary screening of autism children, and the consistency is poor, resulting in poor screening effect and low efficiency.
[0006] The first aspect of the embodiment of the present application provides an autism child response ability automatic screening method based on relative semantic action segmentation, which comprises the following steps: acquiring video information of a doctor and a child interaction process; extracting a skeleton action sequence of the doctor from the video information; segmenting the skeleton action sequence to obtain a time window in which the doctor issues an instruction; obtaining behavior modal information of the child according to the time window; and evaluating the response ability of the child to non-verbal instructions according to the behavior modal information to obtain screening evaluation information.
[0007] Optionally, in an embodiment of the present application, the step of extracting the skeleton action sequence of the doctor from the video information specifically comprises: pre-processing the video information to obtain a pre-processed video; separating a red-green-blue image stream of the doctor from the pre-processed video; identifying the red-green-blue image stream to obtain skeleton information of the doctor in each frame of image; and arranging each frame of skeleton information in time sequence to obtain the skeleton action sequence of the doctor.
[0008] Optionally, in an embodiment of the present application, the step of segmenting the skeleton action sequence to obtain the time window in which the doctor issues an instruction specifically comprises: constructing a space and time sequence model, inputting the skeleton action sequence into the space and time sequence model to obtain a frame-level feature representation of the doctor; acquiring a real frame-level label, and obtaining a frame-level action label according to the frame-level feature representation and the real frame-level label; and obtaining the time window in which the doctor issues an action instruction according to the frame-level action label.
[0009] Optionally, in an embodiment of the present application, the space-time model comprises a joint space model and an inter-frame time sequence model; the constructing the space-time model, inputting the skeleton action sequence into the space-time model to obtain the frame-level feature representation of the doctor specifically comprises: obtaining joint text and inputting the joint text into a large language model to obtain joint text features; obtaining a relative semantic graph of joints according to the joint text features; establishing a space model based on a graph convolution network and fusing the relative semantic graph and the space model to obtain the joint space model; inputting the skeleton action sequence into the joint space model to obtain a space feature; and establishing an inter-frame time sequence model based on a neural network and inputting the space feature into the inter-frame time sequence model to obtain the frame-level feature representation of the doctor.
[0010] Optionally, in an embodiment of the present application, the obtaining the frame-level action label according to the frame-level feature representation and the real frame-level label specifically comprises: obtaining action text and inputting the action text into a large language model to obtain action text features; obtaining an action semantic graph according to the action text features; supervising the frame-level feature representation according to the action text features and the action semantic graph to obtain a supervised frame-level feature representation; and obtaining a frame-level action label according to the frame-level feature representation and the supervised frame-level feature representation.
[0011] Optionally, in an embodiment of the present application, the behavior modal information comprises a skeleton posture, a gesture and a key object position; and the obtaining the behavior modal information of the child according to the time window specifically comprises: cutting the video information according to the time window to obtain a target video; and obtaining the skeleton posture, the gesture and the key object position of the child according to the target video.
[0012] Optionally, in an embodiment of the present application, the evaluating the response ability of the child to the non-verbal instruction according to the behavior modal information to obtain screening evaluation information specifically comprises: constructing a response ability model, training the response ability model according to preset training data to obtain a trained response ability model; and inputting the skeleton posture, the gesture and the key object position into the trained response ability model to obtain the screening evaluation information of the child.
[0013] The second aspect of the embodiments of the present application further provides an automatic screening system for the response ability of an autistic child based on relative semantic action segmentation, wherein the automatic screening system for the response ability of the autistic child based on relative semantic action segmentation comprises:
[0014] a video acquisition module, configured to acquire video information of a process of interaction between a doctor and a child;
[0015] a skeleton sequence extraction module, configured to extract a skeleton action sequence of the doctor from the video information;
[0016] an action segmentation module, configured to segment the skeleton action sequence to obtain a time window in which the doctor issues an instruction;
[0017] a modality perception module, configured to obtain behavior modality information of the child according to the time window;
[0018] a screening evaluation module, configured to evaluate a response ability of the child to the non-verbal instruction according to the behavior modality information, and obtain screening evaluation information.
[0019] The third aspect of the embodiments of the present application further provides a terminal, wherein the terminal comprises a memory, a processor, and an automatic screening program for response ability of an autistic child based on relative semantic action segmentation stored in the memory and capable of running on the processor, and the automatic screening program for response ability of an autistic child based on relative semantic action segmentation, when executed by the processor, implements the steps of the automatic screening method for response ability of an autistic child based on relative semantic action segmentation.
[0020] The fourth aspect of the embodiments of the present application further provides a computer readable storage medium, wherein the computer readable storage medium stores an automatic screening program for response ability of an autistic child based on relative semantic action segmentation, and the automatic screening program for response ability of an autistic child based on relative semantic action segmentation, when executed by a processor, implements the steps of the automatic screening method for response ability of an autistic child based on relative semantic action segmentation.
[0021] Beneficial effects: the present application provides an automatic screening method, system, terminal and medium for response ability of an autistic child based on relative semantic action segmentation, the present application segments a time window in which a non-verbal instruction is initiated from a skeleton action sequence of a doctor, extracts multiple behavior modalities of a child in the time window, thereby screening and evaluating the response ability of the child to the non-verbal instruction, and realizes the automatic and objective screening of the response ability of an autistic child to a non-verbal instruction, realizes the automation and integration of the screening process, reduces the human input, improves the screening efficiency, and improves the accuracy and breadth of the screening. BRIEF DESCRIPTION OF DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments described in the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.
[0023] Figure 1 is a flow chart of a preferred embodiment of the automatic screening method of the response ability of autistic children based on relative semantic action segmentation of the present application;
[0024] Figure 2 is a system diagram of a preferred embodiment of the automatic screening method of the response ability of autistic children based on relative semantic action segmentation of the present application;
[0025] Figure 3 is a skeleton action segmentation model diagram based on relative semantic enhancement in a preferred embodiment of the automatic screening method of the response ability of autistic children based on relative semantic action segmentation of the present application;
[0026] Figure 4 is a structure diagram of a preferred embodiment of the automatic screening system of the response ability of autistic children based on relative semantic action segmentation of the present application;
[0027] Figure 5 is a structure diagram of a preferred embodiment of the terminal of the present application.
[0028] Explanation of reference signs:
[0029] 100, video acquisition module; 200, skeleton sequence extraction module; 300, action segmentation module; 400, modal perception module; 500, screening evaluation module. DETAILED DESCRIPTION
[0030] In order to make the purpose, technical solutions and effects of the present application clearer and more explicit, the technical solutions in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. The described embodiments are only possible technical implementations of the present application, and are not all possible implementations. Based on the embodiments in the present application, those skilled in the art can certainly combine the embodiments of the present application to obtain other embodiments without creative labor, and these embodiments are also within the protection scope of the present application.
[0031] In the related art, autism spectrum disorder (ASD) is a type of neurodevelopmental disorder that occurs in early childhood, and its clinical symptoms mainly include social communication disorders, stereotyped behaviors, and narrow interests. Autism is the most common and largest number of behavioral development disorder among children with developmental diseases. The cause of autism is unknown, and there is no specific drug in clinical practice. Children with autism cannot recover after adulthood, have a high rate of disability, and have a profound impact on the quality of life and long-term social prognosis of patients, causing a huge burden on families and society. Research has found that "early screening, early diagnosis, and early intervention" is the expert consensus in the field of children's developmental behavior disorders. Only through early behavioral screening of children can the symptoms and potential risks of autism in children be diagnosed, so that many behaviors, functions, social and cognitive deficits of ASD children can be improved through behavioral intervention. The earlier the discovery, the earlier the intervention, and the better the effect. Therefore, early behavioral identification and screening analysis of autism is crucial to alleviate the symptoms of children with autism and improve their quality of life.
[0032] Existing screening and diagnosis of developmental behavior disorders in children with autism mainly rely on artificial clinical observation and evaluation of behavior, that is, relying on professional personnel to make medical diagnosis through behavioral observation and to develop corresponding intervention programs. Since this process relies on experienced clinicians, it requires a large amount of medical human resources, a long training period, and the limited number of rehabilitation professionals, making it difficult to meet the growing demand for intervention and evaluation services. At the same time, the results of artificial evaluation are to some extent based on the subjective feelings of professionals, lacking systematic quantitative evaluation, and being highly subjective and inconsistent. Therefore, the medical human resource shortage, strong subjectivity, and lack of systematic program of the artificial clinical method to some extent hinder the diagnosis and treatment process of children with autism.
[0033] Firstly, the terms involved in the embodiments of the present application are introduced:
[0034] ASD (Autism Spectrum Disorder), autism spectrum disorder, i.e., autism;
[0035] STAS (Skeleton-based Temporal Action Segmentation), which divides the actions made by a person in a time sequence through a skeleton sequence, i.e., skeleton temporal action segmentation;
[0036] Transformer, a neural network structure based on attention modeling.
[0037] An automatic screening method, system, terminal and medium for response ability of autistic children based on relative semantic action segmentation are described below with reference to the accompanying drawings. In view of the problem in the prior art that early behavior auxiliary screening of autistic children requires a large amount of human resources and has poor consistency, resulting in poor screening effect and low efficiency, the present application provides an automatic screening method for response ability of autistic children based on relative semantic action segmentation. In the method, the time window of non-verbal instruction initiation is segmented from the doctor's skeleton action sequence, and the multi-behavior modalities of the child are extracted within the time window, so as to screen and evaluate the response ability of the child under non-verbal instruction. The automatic and objective screening of the response ability of autistic children to non-verbal instruction is realized, the automation and integration of the screening process are realized, the human input is reduced, the screening efficiency is improved, and the accuracy and breadth of screening are improved. Thus, the technical problem in the prior art that early behavior auxiliary screening of autistic children requires a large amount of human resources and has poor consistency, resulting in poor screening effect and low efficiency, is solved.
[0038] Currently, early behavior screening of autistic children mainly relies on artificial clinical methods, which consume a large amount of medical human resources, have strong subjectivity and poor consistency, and urgently need to be solved from the machine assistance level. Machine-assisted screening technology is in its infancy, and has the technical problems of poor specialization, low automation, single modal, and poor effect. In order to solve the above technical problems, the present application is based on the "not or less response" in the widely accepted "five not" behavior standard in the domestic autism field. The screening paradigm of the response ability of the doctor and the child to the non-verbal instruction is used to induce the child to make a corresponding response behavior. During the paradigm process, the camera records the interaction video of the child and the doctor for automatic screening. In addition, the automatic response ability screening model based on relative semantic action segmentation is used. First, the target detection tracking and human pose estimation algorithm is used to obtain the human skeleton points of the doctor in each frame. Then, based on the relative semantic action segmentation model, the key non-verbal instruction action of the doctor is segmented with high precision. According to the start and end points of the key non-verbal instruction action of the doctor, a time window is set, the position information of the doctor's hand, the child's hand and the toy is obtained by using the target detection algorithm, and then it is judged whether the child makes a corresponding response behavior after hearing the instruction by spatial position calculation, so as to make ability screening evaluation.
[0039] The application can realize automatic screening of the response ability of autistic children to non-verbal instructions, only needs an interventionist to accompany the child to give instructions, does not need a professional physician to conduct behavior assessment screening, thereby greatly saving medical resources, and can provide quantitative objective screening indicators. Compared with other technologies, the application can completely automatically and integrally realize instruction behavior segmentation and response ability screening and evaluation, wherein the action segmentation algorithm utilizes the relative semantic enhancement to enhance the reliability of small samples, fully utilizes the behavior modalities such as skeleton, object and gesture, and simultaneously proposes an algorithm + rule screening combination, has high robustness and reliability, can be large-scale landed, can promote the development of autism auxiliary diagnosis and treatment theory and technology, and has important significance for promoting the progress of the autism rehabilitation industry.
[0040] The technical solutions of the application will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes can not be described in detail in some embodiments.
[0041] The autism child response ability automatic screening method based on relative semantic action segmentation described in the preferred embodiment of the application, as shown in Figure 1 The autism child response ability automatic screening method based on relative semantic action segmentation includes the following steps:
[0042] In step S101, video information of the interaction process between the doctor and the child is acquired.
[0043] It should be noted that the technical solution is divided into two parts of screening paradigm and screening system. The screening paradigm of the response ability of the child to the non-verbal instruction based on the interaction of the doctor and the child is used to establish a structured interaction scene of the response of the doctor and the child to the non-verbal instruction, to induce the child to make the expected response behavior through the non-verbal instruction of the doctor, and to collect the video in the scene for subsequent screening. The screening system is an automatic response ability screening system based on relative semantic action segmentation, which is used to extract the multi-modal behavior information in the video, and to segment the doctor instruction window according to the skeleton information by using the designed relative semantic action segmentation algorithm, to evaluate and analyze the response ability of the child according to the multi-modal behavior information of the child in the window, so as to complete the screening of the response ability of the autistic child.
[0044] Referring to Figure 2, the response ability screening paradigm based on non-verbal indications of interaction between doctors and children provides a face-to-face scene of interaction between doctors and children in the non-verbal indication response ability screening process, and the screening paradigm is divided into two rounds, different toys are used to attract the attention of children, and when the children are playing the toys attentively, the doctor issues a non-verbal indication of demand for the toy, and observes whether the behavior of the child will respond to the instruction of the doctor. The specific process is as follows: the doctor sits opposite the child, takes out two small balls, and places them on the table; the doctor rolls the ball on the table to the child (the doctor says "ball or ball ball"), and the child freely plays; after playing for 2-3 minutes, enter the first round of testing, the doctor stretches out his hand and says "give the ball to the aunt / uncle"; the doctor takes out the toy with a spring (turtle), and places it on the table; the doctor takes the turtle toy, winds it for two turns, and places it on the table so that the child can touch the toy; after playing for 2-3 minutes, enter the first round of testing, the doctor stretches out his hand and says "give the ball to the aunt / uncle"; the process ends, and the doctor puts away the toys.
[0045] Specifically, video acquisition is performed, and a camera fixed in the scene records the face-to-face interaction process of the doctor and the child from the side, records the whole process video at 30 FPS or 60 FPS, and is used for subsequent screening.
[0046] In step S102, the skeleton action sequence of the doctor is extracted from the video information.
[0047] In one possible implementation, the video information is preprocessed to obtain a preprocessed video; the red-green-blue image stream of the doctor is separated from the preprocessed video; the red-green-blue image stream is identified to obtain the skeleton information of the doctor in each frame of image; and each frame of the skeleton information is arranged in time sequence to obtain the skeleton action sequence of the doctor.
[0048] Referring to Figure 2 , the automatic response ability screening system based on relative semantic action segmentation can record the interaction video of the doctor and the child throughout the screening paradigm process, and extract the skeleton and gesture information of the doctor and the child through a multi-behavior modal perception algorithm, input the skeleton information of the doctor into a relative semantic action segmentation algorithm, thereby detecting the timing position of the non-verbal instruction initiation (i.e. the test of the screening paradigm process) of the doctor in the screening process, and segmenting the screening time window according to the multi-behavior modal information of the child in the time window. According to the multi-behavior modal information of the child in the time window, the response ability of the child to the non-verbal instruction is evaluated, so as to achieve the automatic screening of "not or less response" in the "five not" behavior standard.
[0049] Specifically, in the process of extracting the doctor skeleton sequence, the existing target detection and tracking algorithm is used to separate the doctor RGB stream in the video, and the existing human pose estimation algorithm (Openpose) is used to further extract the doctor skeleton sequence for the subsequent automatic evaluation window segmentation. That is, the purpose of this step is to establish a structured interaction scenario of the doctor and the child responding to non-verbal instructions, so as to obtain video data of the interaction process between the doctor and the child.
[0050] Further, from the whole-process video obtained from the screening paradigm process, first, pre-processing is performed, including adjusting the video frame rate (such as ensuring 30FPS or 60FPS), adjusting the video size to adapt to the subsequent algorithm processing requirements, etc.; the existing target detection and tracking algorithm is used to separate the doctor red-green-blue (RGB) image stream from the pre-processed video. The purpose of this step is to separate the video information of the doctor from the child and other background information, so as to focus on the action analysis of the doctor in the subsequent step; the separated doctor RGB image stream is input into the existing human pose estimation algorithm (such as Openpose), which can identify and extract the skeleton information of the doctor in each frame of image, including the joint position, the connection between joints, etc.; the algorithm output of each frame of skeleton information is arranged in time sequence to generate the skeleton sequence of the doctor.
[0051] In step S103, the skeleton action sequence is segmented to obtain a time window in which the doctor issues an instruction.
[0052] In one possible implementation, a spatial and temporal model is constructed, the skeleton action sequence is input into the spatial and temporal model to obtain a frame-level feature representation of the doctor; a real frame-level label is obtained, and a frame-level action label is obtained according to the frame-level feature representation and the real frame-level label; and the time window in which the doctor issues an action instruction is obtained according to the frame-level action label.
[0053] It should be noted that, in the process of automatically segmenting the non-verbal instruction time window, the skeleton sequence of the doctor is taken as input by using the action segmentation algorithm based on relative semantics, so as to segment all actions (no action, take out a toy, stretch out a hand to initiate a demand, other actions) made by the doctor in the whole process in time sequence, and set the action of the doctor stretching out a hand to initiate a demand toy instruction as the time window for evaluating the response ability of the child, so as to realize the automatic identification and screening of the time window. In this regard, the skeleton time sequence action segmentation for automatically segmenting the non-verbal instruction time window is an action segmentation algorithm based on relative semantic enhancement, which can reduce the dependence on sample size in the process of training the model, enhance the clustering and discrimination between different actions, be more helpful for feature learning, and thus improve the segmentation accuracy of the model for the actions of the doctor.
[0054] In a possible implementation, the space-time model comprises a joint space model and an inter-frame time sequence model. Joint text is obtained, and the joint text is input into a large language model to obtain joint text features; relative semantic graphs of joints are obtained according to the joint text features; a space model is established based on a graph convolution network, and the relative semantic graphs are fused with the space model to obtain the joint space model; the skeleton action sequence is input into the joint space model to obtain space features; an inter-frame time sequence model is established based on a neural network, and the space features are input into the inter-frame time sequence model to obtain frame-level feature representations of the doctor.
[0055] As shown in Figure 3 BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model based on a Transformer architecture; the skeleton sequence of the doctor is input into the model, the model first uses a graph convolution network to preliminarily establish the relationship between joints according to the topological connection of the skeleton; then the description text of the joints is input into the large language model to output the text features of each joint, and the relative semantic graphs of the joints are obtained through the relationship calculation between the text features, which are used for space modeling to further enhance the establishment of the fine-grained relationship between joints. Then, the model uses a Transformer network to establish the inter-frame time sequence relationship of the features, and finally outputs the feature representations of the action sequence of the doctor.
[0056] Further, the graph convolution network is used to preliminarily establish the relationship between joints according to the topological connection of the skeleton, and the purpose of this step is to capture the spatial relationship between joints to provide a basis for subsequent feature extraction and action segmentation; the description text of the joints is input into the large language model to output the text features of each joint, which may contain information such as the function, position or relationship with other joints of the joint, which helps to enhance the semantic understanding ability of the model; the relative semantic graphs of the joints are obtained through the relationship calculation between the text features, which uses the text features of the joints to capture the relative semantic relationship between them, and further enhances the understanding of the fine-grained relationship between joints by the model (that is, the relative semantic graphs are fused with the space model to obtain the joint space model); the Transformer network is used to establish the inter-frame time sequence relationship of the features, and the purpose of this step is to capture the continuity of the action in time, and combine the spatial relationship and the relative semantic relationship of the joints with the time information to form a complete action feature representation.
[0057] In a possible implementation, action text is obtained, and the action text is input into a large language model to obtain action text features; an action semantic graph is obtained according to the action text features; the frame-level feature representation is supervised according to the action text features and the action semantic graph to obtain a supervised frame-level feature representation; and a frame-level action label is obtained according to the frame-level feature representation and the supervised frame-level feature representation.
[0058] As shown in Figure 3 , the model needs to be supervised by a certain sample to realize the segmentation of actions, therefore, in the feature representation level, the model uses the action text to input a large language model to output the action text features of each type, and compares and supervises the features of the corresponding action segments in the action text features and the feature representation; further, the relative relationship between actions is calculated to obtain an action relative semantic graph, and the action semantic graph is used to supervise the relationship distance between the action features; finally, the feature representation is input into a classification head to obtain the action label of each frame, and the action label is calculated with the real artificial labeled label by using cross entropy loss to perform classification supervision. In the training stage, a total loss function is obtained through the three kinds of supervision modes, and the network model is optimized by using the optimizer to perform back propagation, so as to achieve the purpose of training. Under the training of a certain sample, the final model can achieve higher accuracy compared with other action segmentation algorithms, can directly input the skeleton sequence, and can infer the action label of each frame of the model, that is, can segment the action of the doctor at all times, so as to be used for automatically identifying the time window of the non-language instruction.
[0059] Further, in the feature representation level, the action text is input into a large language model to output the action text features of each type, and the features of the corresponding action segments in the feature representation are compared and supervised, and the relative relationship between actions is calculated to obtain an action relative semantic graph (that is, the action text graph of Figure 3 . The action relative semantic graph is used to supervise the relationship distance between the action features. Finally, the feature representation is input into a classification head to obtain the action label of each frame, and the action label is calculated with the real artificial labeled label by using cross entropy loss to perform classification supervision. Through the trained model, the input skeleton sequence is segmented, and the action label of each frame is output. These labels represent the action of the doctor at all times. In the segmented action, the action of the doctor extending his hand to initiate the demand toy instruction is identified, and is set as the time window for evaluating the response ability of the child.
[0060] In step S104, behavior modal information of the child is obtained according to the time window.
[0061] In a possible implementation, the behavior modal information includes skeleton posture, gesture, and key object position. According to the time window, the video information is cropped to obtain target video; and according to the target video, the skeleton posture, gesture, and key object position of the child are obtained.
[0062] Specifically, in the process of multi-behavior modal perception of the child, the video in the time window segmented automatically is cropped, that is, the video when the doctor expresses the non-verbal demand instruction (corresponding to the two test segments of the screening paradigm process respectively) is obtained separately. The skeleton posture, gesture, and key object of the child in the two videos are extracted, that is, the child posture is extracted through human pose estimation, the gesture is extracted through gesture recognition, and the position information of the ball and the clockwork toy is extracted through target detection and tracking. The above various information can provide information basis for the evaluation of the response behavior of the child when the non-verbal instruction is given, and the high-level and easy-to-understand behavior information of the extracted skeleton, gesture, and key object is more consistent with human perception, and is more reliable and interpretable than the one-segment evaluation directly using the video input model.
[0063] Further, according to the non-verbal instruction time window segmented automatically, the corresponding segments are cropped from the video recorded throughout the process, and the segments contain the complete interactive scene when the doctor gives the non-verbal demand instruction; the human pose estimation algorithm is used to process the cropped video to extract the skeleton sequence of the child in the video, including joint position and posture information; the gesture recognition algorithm is applied to recognize and extract the gesture of the child in the video, and the gesture information, as one of the important behavior modalities, can reflect the reaction and intention of the child after hearing the instruction; the target detection and tracking algorithm is used to detect and track the key objects (such as the ball and the clockwork toy) in the video to extract the position information, motion trajectory, etc. of the key objects to analyze the interaction between the child and the key objects. The extracted skeleton posture, gesture, and key object information of the child complement each other and jointly constitute a comprehensive description of the response behavior of the child.
[0064] In step S105, according to the behavior modal information, the response ability of the child to the non-verbal instruction is evaluated to obtain screening evaluation information.
[0065] In a possible implementation, a response ability model is constructed, the response ability model is trained according to preset training data to obtain a trained response ability model, and the skeleton posture, the gesture, and the key object position are input into the trained response ability model to obtain the screening evaluation information of the child.
[0066] Specifically, based on the high-level modal information of the child's skeleton posture, gesture and key objects, the child's response to non-verbal instructions is screened and evaluated. In terms of rules, the child's hand is calculated by spatial position to determine whether the child's hand is extended and whether the toy reaches the doctor's hand, to preliminarily judge whether the child makes the behavior of handing the toy to the doctor. In terms of algorithm, the posture, gesture and key object category position information are all input into the network for feature fusion and modeling through a multi-layer convolutional neural network and feature fusion, and the model is trained according to a certain number of samples and the label of whether to respond, so as to intelligently judge whether the child responds. In the actual screening process, the child's response ability is comprehensively judged according to the screening of the two rules and the algorithm. The combination of the two makes full use of the intelligence of the algorithm and the reliability, interpretability and robustness based on the rules, and can stably carry out screening research and further landing.
[0067] Further, the extracted child skeleton posture, gesture and key object (such as toys) category and position information are input into a multi-layer convolutional neural network for feature fusion and modeling. This step aims to extract the features of the child's behavior from multiple dimensions and build a model that can reflect the child's response ability. Based on the trained model, the child's response is intelligently judged, that is, by inputting the extracted features into the classification head of the model, the prediction result of whether the child responds is obtained.
[0068] It is worth noting that the prior art is in its early stages, and its ability assessment is not granular enough. Typically, model training is based on behavior data collected in specific scenarios to achieve classification screening of autism, but there is a lack of specific ability screening evaluation, resulting in poor interpretability. The present application is aimed at a response ability specialized screening paradigm and system, which only targets the "not / less response" ability in the "five no" behavior standard, designs a specific behavior induction paradigm to assess the response ability of autistic children to non-verbal instructions, which is highly targeted and has high granularity. The existing screening technology has low automation level, and there is no method for automatic full-process evaluation directly based on collected videos, and manual window segmentation is usually required. The present application uses a motion segmentation tool to combine motion segmentation algorithms and autism screening to automatically determine the evaluation time window by segmenting the non-verbal behavior instructions initiated by the doctor, which can reduce the human input in the process and be efficient and convenient. The current skeleton-based motion segmentation algorithm has a lot of room for improvement in accuracy, and it also has high sample size requirements, as well as the problems of small similarity between action classes and fuzzy boundaries. To solve the above problems, the present application proposes a motion segmentation algorithm based on relative semantic graph assistance, which uses the language modality to assist the feature learning of the model to achieve higher segmentation accuracy and greatly reduce the sample dependency, adapting to the problem of insufficient early screening samples. Most of the current technologies use a single modality to screen autistic children, such as using only eye gaze to evaluate children's attention to screen autism. However, the scope of a single modality design is insufficient, and it is difficult to fully represent the behavior of children; the present application combines multiple behavior modalities, uses skeleton modality to segment the window, and uses gestures, skeletons, and objects to achieve reliable, interpretable, and high-precision screening. Most of the existing screening models only use machine learning and deep learning, which have poor reliability and robustness, and the intelligent level of rule-based analysis is also low; the present application adopts a two-stage algorithm + rule screening process, which identifies the indication window through the motion segmentation algorithm and further evaluates the child's ability through the algorithm + rule combination mode. The two-stage process is strong in engineering and has high reliability, while the algorithm + rule evaluation is reliable and has low error rate.
[0069] Next, the automatic screening system for the response ability of autistic children based on relative semantic motion segmentation according to the embodiments of the present application is described with reference to the accompanying drawings.
[0070] Figure 4 is the structure diagram of the automatic screening system for the response ability of autistic children based on relative semantic motion segmentation according to the embodiments of the present application.
[0071] As Figure 4 shown, the automatic screening system for the response ability of autistic children based on relative semantic motion segmentation includes a video acquisition module 100, a skeleton sequence extraction module 200, a motion segmentation module 300, a modality perception module 400, and a screening evaluation module 500.
[0072] Specifically, the video acquisition module 100 is configured to acquire video information of an interaction process between a doctor and a child;
[0073] The skeleton sequence extraction module 200 is configured to extract a skeleton action sequence of the doctor from the video information.
[0074] The action segmentation module 300 is configured to segment the skeleton action sequence to obtain a time window in which the doctor issues an instruction.
[0075] The modality perception module 400 is configured to obtain behavior modality information of the child according to the time window.
[0076] The screening evaluation module 500 is configured to evaluate a response ability of the child to non-verbal instructions according to the behavior modality information, and obtain screening evaluation information.
[0077] Figure 5 A structure diagram of a terminal is provided in the embodiments of the present application. The terminal can include:
[0078] The memory 501, the processor 502, and a computer program stored in the memory 501 and executable on the processor 502.
[0079] The processor 502 implements the automatic screening method for the response ability of an autistic child based on relative semantic action segmentation provided in the above embodiments when executing the program.
[0080] Further, the terminal further includes:
[0081] The communication interface 503 is configured to communicate between the memory 501 and the processor 502.
[0082] The memory 501 is configured to store the computer program executable on the processor 502.
[0083] The memory 501 can include a high-speed RAM memory, and can also include a non-volatile memory, for example, at least one disk memory.
[0084] If the memory 501, the processor 502 and the communication interface 503 are implemented independently, the communication interface 503, the memory 501 and the processor 502 can be connected with each other through a bus and complete the communication between each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 5 Only one thick line is used to represent the bus in the figure, but it does not mean that there is only one bus or only one type of bus.
[0085] Optionally, in a specific implementation, if the memory 501, the processor 502 and the communication interface 503 are integrated on a chip, the memory 501, the processor 502 and the communication interface 503 can complete the communication between each other through an internal interface.
[0086] The processor 502 can be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement one or more embodiments of the present application.
[0087] The embodiment also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the automatic screening method for the response ability of autistic children based on relative semantic action segmentation.
[0088] An embodiment of the present application provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the automatic screening method for the response ability of autistic children based on relative semantic action segmentation. Figure 1 The embodiment of the present application corresponds to any embodiment of the automatic screening method for the response ability of autistic children based on relative semantic action segmentation.
[0089] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to user analysis data, user stored data, user displayed data, etc.) and signals involved in the present application are all information, data and signals authorized by the user or authorized by all parties; and the collection, use and processing of the related information, data and signals comply with the relevant national and regional laws, regulations and standards.
[0090] In the description of the application, reference to "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" means that a particular feature, structure, material, or characteristic being described is included in at least one embodiment or example of the application. The appearances of the phrase in various places in the specification are not necessarily all referring to the same embodiment or example. Furthermore, the described specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples. In addition, the usage of "first", "second" and the like does not imply any importance, unless the context clearly indicates so. The usage of these terms is used as identifiers in this patent application. Thus, a device of a first charge can imply that there is at least one device of a second charge, and vice versa.
[0091] Furthermore, the terms "first", "second", etc. are used herein only to describe various steps in a method, process, and the like and are not intended to refer to relative importance. Thus, a "first" feature can imply a "second" feature, and vice versa, unless otherwise specifically noted. Any embodiment or example of the application can be combined with any other embodiment or example of the application, unless otherwise specifically noted. In addition, the term "N" means at least two, for example two, three, etc., unless specifically noted otherwise.
[0092] Any process or method descriptions or blocks in flow charts or otherwise described herein represent embodiments of modules, segments, or portions of code which include one or more executable instructions for implementing specific logical functions or steps, or combinations of tasks or steps. Alternate implementations are included within the scope of the preferred embodiments of the application in which additional functionality can be added or where functions can be implemented out of the order described or not implemented at all. The preferred embodiments of this application have been described herein with the understanding that these embodiments are implemented as software, hardware, firmware, middleware or a combination of these means.
[0093] The logic and / or steps represented in flow diagrams or otherwise described herein, for example, can be considered as a sequence of instructions to implement logic functions, and can be embodied in any computer-readable storage medium for use by an instruction execution system, apparatus, or device, such as a computer-based system, processor- containing system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions. In the context of this specification, a "computer-readable storage medium" can be any means that can contain, store, communicate, propagate or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer-readable storage medium can be a computer- readable storage medium that can be any media that can be used to store the desired program instructions in a form readable by a computer. More specific examples (a non-exhaustive list) of the computer-readable storage medium include the following: an electrical connection having one or more wires (electrical, optical, and / or other communications medium), a portable computer diskette, a hard disk, a RAM, a ROM, an erasable programmable ROM (EPROM or Flash memory), an optical fiber, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable storage medium can even be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, via, for example, optical scanning of the paper or other medium, then compiled, interpreted, or otherwise processed in a suitable manner, if necessary, and then stored in a computer memory.
[0094] It should be understood that aspects of the application can be implemented in hardware, software, firmware or combinations thereof. In the above embodiments, the N steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. As such, if implemented in hardware, and in another embodiment, any of the following technologies, known in the art, or their combinations can be used: discrete logic circuitry having logic gates for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and the like.
[0095] Those of skill in the art will understand that the steps carried out by the above-mentioned embodiments of the method can be implemented by programs instructing the relevant hardware, and the programs can be stored in a computer-readable storage medium. When the programs are executed, they include one or a combination of the steps of the method embodiments.
[0096] In addition, each of the function units in each of the embodiments of the present application can be integrated in one processing module, or each unit can be physically present separately, or two or more units can be integrated in one module. The integrated module can be realized in the form of hardware or in the form of a software function module. When the integrated module is realized in the form of a software function module and sold or used as an independent product, it can also be stored in a computer readable storage medium.
[0097] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it should be understood that the above embodiments are exemplary and should not be construed as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present application.
[0098] It should be understood that the application of the present application is not limited to the above examples, and those skilled in the art can improve or change the above description, and all these improvements and changes should belong to the protection scope of the claims of the present application.
[0099] Finally, it should be pointed out that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. An automatic screening method for the response ability of autistic children based on relative semantic action segmentation, characterized in that, The automatic screening method for response ability of autistic children based on relative semantic action segmentation comprises: obtaining video information of a doctor and a child interaction process; extracting a skeleton action sequence of the doctor from the video information; segmenting the skeleton action sequence to obtain a time window in which the doctor issues an instruction; obtaining behavior mode information of the child according to the time window; evaluating the response ability of the child to non-verbal instructions according to the behavior mode information to obtain screening evaluation information; the segmentation of the skeleton action sequence to obtain the time window in which the doctor issues an instruction specifically comprises: constructing a space and time sequence model, inputting the skeleton action sequence into the space and time sequence model to obtain a frame-level feature representation of the doctor; obtaining real frame-level labels, and obtaining frame-level action labels according to the frame-level feature representation and the real frame-level labels; obtaining a time window in which the doctor issues an action instruction according to the frame-level action labels; the space and time sequence model comprises a joint space model and an inter-frame time sequence model; the construction of the space and time sequence model, the input of the skeleton action sequence into the space and time sequence model, and the obtaining of the frame-level feature representation of the doctor specifically comprise: obtaining joint text and inputting the joint text into a large language model to obtain joint text features; obtaining a relative semantic graph of joints according to the joint text features; establishing a space model based on a graph convolution network, and fusing the relative semantic graph and the space model to obtain the joint space model; inputting the skeleton action sequence into the joint space model to obtain a space feature; establishing an inter-frame time sequence model based on a neural network, and inputting the space feature into the inter-frame time sequence model to obtain the frame-level feature representation of the doctor.
2. The method for automatic screening of response ability of autistic children based on relative semantic action segmentation according to claim 1, characterized in that, the extraction of the skeleton action sequence of the doctor from the video information specifically comprises: preprocessing the video information to obtain a pretreated video; separating a red-green-blue image stream of the doctor from the pretreated video; identifying the red-green-blue image stream to obtain skeleton information of the doctor in each frame of image; arranging each frame of skeleton information in time sequence to obtain the skeleton action sequence of the doctor. 3.The method of claim 1, wherein the method is characterized by, the obtaining of the frame-level action labels according to the frame-level feature representation and the real frame-level labels specifically comprises: obtaining action text and inputting the action text into a large language model to obtain action text features; obtaining an action semantic graph according to the action text features; supervising the frame-level feature representation according to the action text features and the action semantic graph to obtain a supervised frame-level feature representation; obtaining frame-level action labels according to the frame-level feature representation and the supervised frame-level feature representation. 4.The method of claim 1, wherein the method is characterized by, the behavior mode information comprises skeleton posture, gesture and key object position; the obtaining of the behavior mode information of the child according to the time window specifically comprises: cropping the video information according to the time window to obtain a target video; obtaining the skeleton posture, gesture and key object position of the child according to the target video.
5. The method for automatic screening of response ability of autistic children based on relative semantic action segmentation according to claim 4, characterized in that, The screening evaluation information is obtained according to the response ability of the child to the non-verbal instruction based on the behavior mode information, and specifically includes: A response ability model is constructed, and the response ability model is trained according to preset training data to obtain a trained response ability model; The skeleton posture, the gesture and the key object position are input into the trained response ability model to obtain the screening evaluation information of the child.
6. An automatic screening system for response ability of autistic children based on relative semantic action segmentation, characterized in that, The automatic screening system for the response ability of an autistic child based on relative semantic action segmentation is applied to the automatic screening method for the response ability of an autistic child based on relative semantic action segmentation in any one of claims 1-5. The automatic screening system for the response ability of an autistic child based on relative semantic action segmentation includes: A video acquisition module is configured to acquire video information of a process of interaction between a doctor and a child; A skeleton sequence extraction module is configured to extract a skeleton action sequence of the doctor from the video information; An action segmentation module is configured to segment the skeleton action sequence to obtain a time window in which the doctor issues an instruction; A mode perception module is configured to obtain behavior mode information of the child according to the time window; A screening evaluation module is configured to evaluate the response ability of the child to the non-verbal instruction according to the behavior mode information to obtain screening evaluation information.
7. A terminal, characterized by comprising: The terminal includes a memory, a processor, and an automatic screening program for the response ability of an autistic child based on relative semantic action segmentation stored on the memory and executable on the processor, and the automatic screening program for the response ability of an autistic child based on relative semantic action segmentation, when executed by the processor, implements the steps of the automatic screening method for the response ability of an autistic child based on relative semantic action segmentation in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an automatic screening program for the response ability of an autistic child based on relative semantic action segmentation, and the automatic screening program for the response ability of an autistic child based on relative semantic action segmentation, when executed by the processor, implements the steps of the automatic screening method for the response ability of an autistic child based on relative semantic action segmentation in any one of claims 1-5.
Citation Information
Patent Citations
Primary screening device for autism based on non-social sound stimulus behavior paradigm
CN109431523A
Abnormal behavior detection method for autistic children based on finger object identification and sight tracking
CN118658604A