A method and system for analyzing the neurodevelopmental status of a small infant
By combining cameras and audio modules, the system collects full-body and localized motion data of infants. Utilizing multimodal feature fusion for decision-making, it solves the problems of scarcity of manual assessment and sensor interference in existing technologies, enabling automated and accurate analysis of the neurodevelopmental status of infants.
Patent Information
- Application Number
- CN202411871397.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2044-12-18
AI Technical Summary
Existing methods for analyzing the neurodevelopmental status of infants rely on manual assessment, which is scarce, time-consuming, and subject to observer fatigue, making it difficult to achieve objective and quantitative assessment. Furthermore, the sensors attached to the infants interfere with spontaneous movement.
By combining a camera and an audio module, and collecting data on whole-body active movement and local induced movement, multimodal behavioral features are extracted using human posture estimation and facial behavior analysis models. Logistic regression and support vector machines are then used for decision fusion to achieve automated analysis of the neurodevelopmental status of infants.
It improves the accuracy and convenience of analyzing the neurodevelopmental status of infants, reduces reliance on assessors, lowers the risk of misdiagnosis and missed diagnosis, and realizes automated quantitative assessment of the neurodevelopmental status of infants.
Smart Images

Figure CN119679370B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision, in particular to a small infant neural development state analysis method and system. BACKGROUND
[0002] If a small infant (small infant especially emphasizes the early stage of infant, within 6 months) is harmed by high-risk factors such as premature birth, low body, asphyxia, hypoxic-ischemic encephalopathy, intracranial hemorrhage, etc., it may lead to abnormal or disorder of neural development, and further lead to cerebral palsy, mental retardation, autism, etc. With the improvement of medical level, the survival rate of high-risk small infants is continuously improved, and how to reduce the incidence of high-risk small infant neural development state and reduce the degree of disability is more and more valued. At present, clinicians generally assess small infant brain injury and clinical rehabilitation efficacy according to cranial ultrasound, MRI, CT, brainstem auditory evoked potential, biochemical indicators and various scoring scale examinations. However, imaging and electrophysiological evaluation as a relatively complex technical evaluation requires specific equipment, and the sensitivity, specificity and accuracy of traditional evaluation scales in predicting neurological development outcomes differ greatly, and the diagnosis time is usually after the child is one year old or even two years old. General movements (GMs) are spontaneous movements that exist from early fetal stage to about 20 weeks after full-term birth, which last for several seconds to several minutes, and are divided into normal movements and abnormal movements. Through GMs evaluation, specific neural damage can be sensitively prompted, and reliable early prediction of neurological development disorders such as cerebral palsy and autism can be made, and rehabilitation intervention can be given as soon as possible to improve the prognosis and reverse adverse neurological development outcomes, which has important clinical and social significance.
[0003] At present, GMs can only be visually evaluated by trained clinicians with evaluation license, and these clinicians need a lot of training and years of practical evaluation experience to achieve appropriate accuracy, so evaluators are very scarce; the evaluation method requires clinicians to observe GMs for a long time, which is easily affected by observer fatigue, and it is difficult to perform objective and quantitative evaluation, which affects the accuracy of diagnosis and easily causes misdiagnosis and missed diagnosis problems; the manual and time-consuming nature of GMs evaluation and the scarcity of evaluators make the current GMs test usually only used in cases with medical problems, such as premature birth, hypoxia or congenital heart disease, etc., and not used as a physical examination screening tool for ordinary infants. A large number of infant physical examination screening requires automatic quantitative evaluation of GMs; the existing automatic quantitative evaluation means based on GMs uses position sensors and acceleration sensors attached to the limbs to evaluate the spontaneous movement of the limbs during GMs, analyzes the periodicity of speed and acceleration in spontaneous movement, and diagnoses movement disorders using the extracted features. However, the above research connects sensors or markers to small infants, which will interfere with the spontaneous occurrence of movement, and it is not easy to achieve sensor attachment to infants. Therefore, a small infant neural development state analysis method is proposed. SUMMARY
[0004] The technical problem solved by the present application is how to more accurately and conveniently analyze the neurological development state of small infants, and a small infant neurological development state analysis method is provided.
[0005] The present application solves the above technical problems by the following technical solutions, and the present application comprises the following steps:
[0006] S1: data acquisition preparation
[0007] Place the small infant on the test table, adjust the relative position of the side fixing frame and the test table, so that the normal direction of the camera on the display module is perpendicular to the test table and the field of view covers the whole body of the small infant, and the sound module is fixed on the test table on the left and right sides of the head of the small infant;
[0008] S2: whole body active motion data acquisition
[0009] The camera is used to collect the whole body active motion data of the small infant;
[0010] S3: local induced motion data acquisition
[0011] The display module and / or the sound module cooperate with the camera to collect the local induced motion data of the small infant;
[0012] S4: multi-motion feature extraction and fusion decision
[0013] According to the motion paradigm, the features are extracted, the first classification sub-decision and the second classification sub-decision are formed through feature layer fusion, the final decision result is obtained by using the weighted average of the sub-decision, and then the small infant neurological development state analysis and evaluation work is carried out.
[0014] Further, in the step S2, the specific process is as follows:
[0015] S21: the computer controls the camera to detect the whole body active motion of the small infant in a natural state, adjusts the relative position of the camera and the test table, so that the video contains complete four limbs and head information of the small infant;
[0016] S22: the video is identified by using a human pose estimation model, the key points of the four limbs of the small infant are identified, and the motion feature data of the above four limb key points are extracted.
[0017] Further, in the step S3, the specific process is as follows:
[0018] Step S31: The computer controls the display module to present a light point and move at high speed and large amplitude in the display range, while the camera detects the head turning posture of the visual following process of the small infant, uses a human pose estimation model to identify the head key points of the small infant, and extracts the motion feature data of the head key points;
[0019] Step S32: The computer controls the display module to present a light point and move at low speed and small amplitude in the display range, while the camera detects the eye movement of the visual following process of the small infant, and obtains the gaze point motion trajectory of the small infant under the visual induction of the light point;
[0020] Step S33: The computer controls the sound module to emit continuous sounds on both sides of the head of the small infant, while the camera detects the head turning posture of the auditory following process of the small infant, uses a human pose estimation model to identify the head key points of the small infant, and extracts the motion feature data of the head key points;
[0021] Step S34: The computer controls the display module and the sound module to be combined with synchronous stimulation, the display module presents a light point and moves at high speed and large amplitude in the display range, while the sound module emits a sound in the direction of the motion endpoint of the light point, the camera detects the head turning posture of the audiovisual joint following process of the small infant, uses a human pose estimation model to identify the head key points of the small infant, and extracts the motion feature data of the head key points;
[0022] Step S35: The computer controls the display module and the sound module to be combined with synchronous stimulation, the display module presents a light point and moves at low speed and small amplitude in the display range, while the sound module emits a sound in the direction of the motion endpoint of the light point, the camera detects the eye movement and pupil change of the visual following process of the small infant, and obtains the gaze point motion trajectory and pupil size under the joint induction of visual and auditory of the light point;
[0023] Step S36: The computer controls the display module and the sound module to be combined with asynchronous stimulation, the display module presents a light point and moves at high speed and large amplitude in the display range, while the sound module emits a sound in the direction of the motion starting point of the light point, the camera detects the head turning posture of the audiovisual joint following process of the small infant, uses a human pose estimation model to identify the head key points of the small infant, and extracts the motion feature data of the head key points;
[0024] Step S37: The computer controls the display module and the sound module to be combined with asynchronous stimulation, the display module presents a light point and moves at low speed and small amplitude in the display range, while the sound module emits a sound in the direction of the motion starting point of the light point, the camera detects the eye movement and pupil change of the audiovisual joint following process of the small infant, and obtains the gaze point motion trajectory and pupil size under the joint induction of visual and auditory of the light point.
[0025] Further, in the step S4, the motion paradigm includes whole-body active motion and local evoked motion; the local evoked motion includes local single-modal evoked motion and local multi-modal evoked motion, wherein the single-modal is visual or auditory, and the multi-modal is visual and auditory combined; the local multi-modal evoked motion is divided into synchronous stimulation and asynchronous stimulation according to whether the visual and auditory stimulation directions are consistent.
[0026] Further, in the step S4, the feature extracted from the whole-body active motion data is F G , the feature extracted from the visual modal evoked data is F L1 , the feature extracted from the auditory modal evoked data is F L2 , the feature extracted from the visual-auditory modal synchronous evoked data is F L3 , and the feature extracted from the visual-auditory modal asynchronous evoked data is F L4 , wherein the whole-body active motion data is the four-limb key point motion feature data obtained in the step S22, the visual modal evoked data is the head key point motion feature data and the gaze point motion trajectory obtained in the steps S31 and S32, the auditory modal evoked data is the head key point motion feature data obtained in the step S33, the visual-auditory modal synchronous evoked data is the head key point motion feature data, the gaze point motion trajectory and the pupil size obtained in the steps S34 and S35, and the visual-auditory modal asynchronous evoked data is the head key point motion feature data, the gaze point motion trajectory and the pupil size obtained in the steps S36 and S37.
[0027] Further, the extraction process of the feature F G is as follows: using a human pose estimation model, the four-limb key point coordinate position data of the small infant is obtained, and the joint rotation angle, the joint rotation angular velocity and the joint rotation angular acceleration are further calculated, and then the feature F G is extracted as follows:
[0028] F G =[θ ij ω ij a ij ],i∈[1,12],j∈[1,6]
[0029] , wherein θ, ω and a respectively represent the joint rotation angle, the joint rotation angular velocity and the joint rotation angular acceleration; i represents left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle and right ankle; and j represents the maximum value, the minimum value, the mean value, the peak-to-peak value, the root mean square value and the standard deviation of each joint motion parameter, including the joint rotation angle, the joint rotation angular velocity or the joint rotation angular acceleration.
[0030] The feature F G is used to train a logistic regression machine learning method.G The process of obtaining the first classifier decision is as follows: For feature F G Standardization is performed; a logistic regression model is used to train the fusion decision; the probability value output by the logistic regression model is used to obtain the first classification sub-decision. If the probability is greater than 0.5, the sub-decision is considered to be category 1, i.e., abnormal neurodevelopment; otherwise, it is considered to be category 0, i.e., normal neurodevelopment.
[0031] Furthermore, feature F L1 The extraction process is as follows: Using a human pose estimation model, the coordinates of key points on the infant's head are obtained, and the head rotation angle, angular velocity, and angular acceleration are further calculated. Simultaneously, a facial behavior analysis model is used to obtain the infant's gaze trajectory. The gaze trajectory error is obtained by calculating the distance to the stimulus light point's trajectory, and then feature F is extracted. L1 as follows:
[0032] F L1 =[θ j ω j a j s j ],j∈[1,6]
[0033] Where θ, ω, a, and s represent the head rotation angle, head rotation angular velocity, head rotation angular acceleration, and fixation point trajectory error, respectively; j represents the maximum, minimum, mean, peak-to-peak, root mean square, and standard deviation of the head motion parameters and eye motion parameters. The head motion parameters include the head rotation angle, head rotation angular velocity, and head rotation angular acceleration, while the eye motion parameters are the fixation point trajectory error.
[0034] Feature F L2 The extraction process is as follows: Using a human pose estimation model, the coordinate positions of key points on the infant's head are obtained, and the head rotation angle, head rotation angular velocity, and head rotation angular acceleration are further calculated, thereby extracting feature F. L2 as follows:
[0035] F L2 =[θ j ω j a j ],j∈[1,6]
[0036] Where θ, ω, and a represent the head rotation angle, head rotation angular velocity, and head rotation angular acceleration, respectively; j represents the maximum, minimum, mean, peak-to-peak, root mean square, and standard deviation of the head motion parameters, which include the head rotation angle, head rotation angular velocity, and head rotation angular acceleration.
[0037] Feature FL3 The extraction process is as follows: using a human pose estimation model, obtaining the coordinate position data of the key points of the small infant's head, and further calculating the head rotation angle, head rotation angular velocity and head rotation angular acceleration, while using a facial behavior analysis model to obtain the gaze point motion trajectory of the small infant, calculating the distance between the gaze point motion trajectory and the stimulating light point motion trajectory to obtain the gaze point motion trajectory error, and then extracting the feature F L3 As follows:
[0038] F L3 = [θ j ω j a j s j p j ], j e [1, 6]
[0039] Wherein, θ, ω, a, s, p respectively represent the head rotation angle, the head rotation angular velocity, the head rotation angular acceleration, the gaze point motion trajectory error and the pupil size; j represents the maximum value, the minimum value, the mean value, the peak-to-peak value, the root mean square value, the standard deviation of the head motion parameters, the eye motion parameters and the pupil parameters; the head motion parameters include the head rotation angle, the head rotation angular velocity and the head rotation angular acceleration, the eye motion parameters are the gaze point motion trajectory error, and the pupil parameters are the pupil size;
[0040] The extraction process of the feature F L4 is as follows: using a human pose estimation model, obtaining the coordinate position data of the key points of the small infant's head, and further calculating the head rotation angle, head rotation angular velocity and head rotation angular acceleration, while using a facial behavior analysis model to obtain the gaze point motion trajectory of the small infant, calculating the distance between the gaze point motion trajectory and the stimulating light point motion trajectory to obtain the gaze point motion trajectory error, and then extracting the feature F L4 As follows:
[0041] F L3 = [θ j ω j a j s j p j ], j e [1, 6]
[0042] Wherein, θ, ω, a, s, p respectively represent the head rotation angle, the head rotation angular velocity, the head rotation angular acceleration, the gaze point motion trajectory error and the pupil size; j represents the maximum value, the minimum value, the mean value, the peak-to-peak value, the root mean square value, the standard deviation of the head motion parameters, the eye motion parameters and the pupil parameters; the head motion parameters include the head rotation angle, the head rotation angular velocity and the head rotation angular acceleration, the eye motion parameters are the gaze point motion trajectory error, and the pupil parameters are the pupil size;
[0043] The support vector machine machine learning method is used to obtain the feature F L1 , F L2 , F L3 , F L4 The process of obtaining the first classification sub-decision is as follows: the features F L1 , F L2 , F L3 , F L4 are standardized; the SVM model is used to train the fusion decision; the input features are classified by the hyperplane trained by the SVM to obtain the second classification sub-decision.
[0044] Further, the calculation formula of the gaze point motion trajectory error is as follows:
[0045]
[0046] Wherein, A(x,y) and A'(x',y') are the corresponding points and their coordinates on the motion trajectory of the stimulating light point and the motion trajectory of the gaze point of the small infant.
[0047] Further, in the step S4, each sub-decision is combined into a fusion decision vector D, and the decision vector is further analyzed to obtain the final decision about the task, and the final result of the decision layer fusion calculation is based on the weighted average value of the plurality of sub-decisions, and the process formula is as follows:
[0048]
[0049] Wherein, d final is the final decision result, y is the weighted average value, m is the number of sub-decisions, D is the fusion decision vector combined by the sub-decisions, W=(w1,w2,…,w m ) is the weight vector, and d final is 0, indicating that the neural development of the small infant is normal, and is 1, indicating that the neural development of the small infant is abnormal.
[0050] The application also provides a small infant neural development state analysis system for the small infant neural development state analysis method, comprising a test table, a display module, a sound module, a camera, a computer and a table side fixing frame, wherein the camera is integrated on the display module; the sound module is arranged on the left and right sides of the head of the small infant; the display module is fixed on the table side fixing frame, the relative position between the table side fixing frame and the test table is adjusted, so that the camera can collect the whole body movement video of the small infant; the computer is used for controlling the display module and the sound module and analyzing the video collected by the camera, and the display module is a naked-eye VR device.
[0051] Compared with the prior art, the application has the following advantages:
[0052] The traditional method mainly analyzes abnormal spontaneous movement of the whole body, is single in mode, needs large sample control group data for comparative analysis, and cannot solve clinical problems such as heterogeneity among individuals with neurodevelopmental status diseases; on the basis of whole body spontaneous movement data collection and analysis, the application increases local movement data collection and analysis induced by visual and auditory sensory organs, obtains multi-modal behavior characteristics, and improves clinical classification accuracy through late fusion decision of modal characteristics. BRIEF DESCRIPTION OF DRAWINGS
[0053] Figure 1 is a structural schematic diagram of a small infant neurodevelopmental status analysis system in an embodiment of the application;
[0054] Figure 2 is a test schematic diagram of an embodiment of the application using a naked-eye VR device as a display module;
[0055] Figure 3 is a flowchart of a small infant neurodevelopmental status analysis method in an embodiment of the application;
[0056] Figure 4 is a flowchart of induced movement detection and analysis in an embodiment of the application. DETAILED DESCRIPTION
[0057] The embodiments of the application will be described in detail below, and the embodiments are implemented on the premise of the technical solution of the application, and detailed implementation modes and specific operation processes are given, but the protection scope of the application is not limited to the following embodiments.
[0058] As shown in Figure 1 , the embodiment provides a technical solution: a small infant neurodevelopmental status analysis system for the following small infant neurodevelopmental status analysis method, comprising: a test table 1, a display module 2, an audio module 3, a camera 4, a computer 5 and a table side fixing frame 6, wherein the camera 4 is integrated on the display module 2; the audio module 3 is arranged on the left and right sides of the head of the small infant 7; the display module 2 is fixed on the table side fixing frame 6, the relative position of the table side fixing frame 6 and the test table 1 is adjusted, so that the camera 4 can collect the whole body movement video of the small infant 7; the computer 5 is used for controlling the display module 2 and the audio module 3 and analyzing the video collected by the camera 3.
[0059] It should be noted that, in order to increase the photorealism of the visual stimulation light spot and reduce the potential harm of the electromagnetic radiation of the traditional display screen to the vision of the small infant 7, the display module 2 in the embodiment uses a naked-eye VR device, as shown in Figure 2 The naked-eye VR device is the "suspended display module" in the Chinese invention patent application with the application number 202410664169.0 and the name "a children autism rehabilitation training system and method based on floating interaction"
[0060] As Figure 3 shown, the embodiment provides a small infant neural development state analysis method, including the following steps:
[0061] Step 1: data collection preparation
[0062] Place the small infant on the test bench, adjust the relative position of the bench side fixing frame and the test bench, so that the normal direction of the camera on the display module is perpendicular to the test bench and the field of view covers the whole body of the small infant, and the sound module is fixed on the test bench on the left and right sides of the small infant's head;
[0063] Step 2: whole body active motion data collection
[0064] Use the camera to collect the whole body active motion data of the small infant;
[0065] In this embodiment, the specific process of step 2 is as follows:
[0066] Step S21: the computer controls the camera to detect the small infant's whole body active motion in a natural state, adjusts the relative position of the camera and the test bench, so that the video contains complete four limbs and head information of the small infant;
[0067] Step S22: use the human pose estimation model to identify the video, identify the small infant's four limb key points, extract the four limb key point motion feature data (four limb key point coordinates, joint rotation angle, joint rotation angular velocity, joint rotation angular acceleration), and reflect the small infant's whole body autonomous motion amplitude and rhythm.
[0068] It should be noted that for whole body active motion, the rotation angle, angular velocity, angular acceleration and other parameters of the head are not directly analyzed, but the overall motion pattern of the small infant's limbs (such as twisting motion, restlessness motion, etc.) is observed to evaluate and diagnose.
[0069] In this embodiment, the human pose estimation model is Openpose, Media Pipe Blaze Pose, etc.
[0070] Step 3: local evoked motion data collection
[0071] Use the display module and / or sound module to cooperate with the camera to collect the small infant's local evoked motion data;
[0072] As Figure 4 shown, in this embodiment, the specific process of step 3 is as follows:
[0073] Step S31: The computer controls the display module to present a light point and move at high speed and large amplitude horizontally in the display range, while the camera detects the head rotation posture of the small infant during the visual following process, uses a human pose estimation model to identify the head key points of the small infant, extracts the motion feature data (head key point coordinates, head rotation angle, head rotation angular velocity, head rotation angular acceleration) of the above-mentioned head key points, and reflects the local induced motion amplitude and rhythm of the small infant;
[0074] Step S32: The computer controls the display module to present a light point and move at low speed and small amplitude horizontally in the display range, while the camera detects the eye movement of the small infant during the visual following process, and obtains the gaze point motion trajectory of the small infant under the visual induction of the light point;
[0075] Step S33: The computer controls the sound module to emit continuous sound on both sides of the head of the small infant, while the camera detects the head rotation posture of the small infant during the auditory following process, uses a human pose estimation model to identify the head key points of the small infant, extracts the motion feature data (head key point coordinates, head rotation angle, head rotation angular velocity, head rotation angular acceleration) of the above-mentioned head key points, and reflects the local induced motion amplitude and rhythm and rotation reaction time of the small infant;
[0076] Step S34: The computer controls the display module and the sound module to be combined with synchronous stimulation, the display module presents a light point and moves at high speed and large amplitude horizontally in the display range, while the sound is emitted in the direction of the motion endpoint of the light point, the camera detects the head rotation posture of the small infant during the audiovisual combined following process, uses a human pose estimation model to identify the head key points of the small infant, extracts the motion feature data (head key point coordinates, head rotation angle, head rotation angular velocity, head rotation angular acceleration) of the above-mentioned head key points, and reflects the local induced motion amplitude and rhythm and rotation reaction time of the small infant;
[0077] Step S35: The computer controls the display module and the sound module to be combined with synchronous stimulation, the display module presents a light point and moves at low speed and small amplitude horizontally in the display range, while the sound is emitted in the direction of the motion endpoint of the light point, the camera detects the eye movement and pupil change of the small infant during the visual following process, obtains the gaze point motion trajectory and pupil size under the audiovisual combined induction of the light point, and reflects the local induced motion amplitude and rhythm and induced reaction time of the small infant;
[0078] Step S36: The computer controls the display module and the sound module to jointly perform asynchronous stimulation. The display module presents a light point and performs high-amplitude horizontal motion in the display range. Meanwhile, the sound module emits a sound in the direction of the motion starting point. The camera detects the head rotation posture of the infant during the audiovisual joint following process. The human posture estimation model is used to identify the head key points of the infant. The motion feature data (head key point coordinates, head rotation angle, head rotation angular velocity, and head rotation angular acceleration) of the head key points are extracted, reflecting the local induced motion amplitude and rhythm of the infant and the rotation reaction time.
[0079] Step S37: The computer controls the display module and the sound module to jointly perform asynchronous stimulation. The display module presents a light point and performs low-speed small-amplitude horizontal motion in the display range. Meanwhile, the sound module emits a sound in the direction of the motion starting point. The camera detects the eye movement and pupil changes of the infant during the audiovisual joint following process. The gaze point motion trajectory and pupil size under the visual and auditory joint induction of the light point are obtained, reflecting the local induced motion amplitude and rhythm of the infant and the induced reaction time.
[0080] In this embodiment, the gaze point motion trajectory is calculated by analyzing the eye features and head posture using a facial behavior analysis model such as OpenFace.
[0081] In this embodiment, the amplitude is based on the displayable length in the horizontal direction of the display module. The large amplitude is defined as any value within 70% to 100%, and the small amplitude is defined as any value within 30% to 50%. The speed is based on the horizontal speed on the display module screen converted from the eye movement speed. A speed higher than 40 mm / s is defined as high speed, and vice versa.
[0082] In this embodiment, the pupil size is obtained using a camera through a machine vision method. The specific process is as follows:
[0083] 1) Image acquisition and preprocessing
[0084] The face image is acquired using a camera and preprocessed.
[0085] 2) Eye position detection
[0086] The eye is identified from the face image and the position information is obtained using a cascade classifier of Haar features.
[0087] 3) Pupil positioning and segmentation
[0088] The pupil boundary is determined by thresholding and edge detection on the darkest point of the eye region.
[0089] 4) Pupil feature extraction
[0090] The pupil diameter and area feature parameters are extracted, i.e., the pupil size information is obtained.
[0091] Step 4: Multi-motion feature extraction and fusion decision
[0092] According to the motion paradigm (whole body active motion, local induced motion), the features (aggregated limb motion features, head motion features, eye motion features) are extracted, the classification sub-decision is formed through feature layer fusion, the final decision result is obtained by weighted average of sub-decision, and the doctor's clinical judgment is combined to assist in carrying out the analysis and evaluation of the neurological development state of small infants.
[0093] In this embodiment, the motion paradigm includes whole body active motion and local induced motion; the local induced motion includes local single mode induced motion and local multi-mode induced motion, wherein the single mode is visual or auditory, and the multi-mode is visual and auditory joint; the local multi-mode induced motion is divided into synchronous stimulation and asynchronous stimulation according to whether the visual and auditory stimulation directions are consistent.
[0094] In this embodiment, the features extracted from the whole body active motion data are F G , the features extracted from the visual modal induced data are F L1 , the features extracted from the auditory modal induced data are F L2 , the features extracted from the audiovisual modal synchronous induced data are F L3 , and the features extracted from the audiovisual modal asynchronous induced data are F L4 , wherein the whole body active motion data is the limb key point motion feature data obtained in step S22, the visual modal induced data is the head key point motion feature data and the gaze point motion trajectory obtained in steps S31 and S32, the auditory modal induced data is the head key point motion feature data obtained in step S33, the audiovisual modal synchronous induced data is the head key point motion feature data and the gaze point motion trajectory and pupil size obtained in steps S34 and S35, and the audiovisual modal asynchronous induced data is the head key point motion feature data and the gaze point motion trajectory and pupil size obtained in steps S36 and S37.
[0095] In this embodiment, the feature F G is extracted as follows:
[0096] Using human pose estimation models such as Openpose, Media Pipe Blaze Pose, the limb key point coordinate position data of small infants is obtained, and further calculation of joint rotation angle, joint rotation angular velocity and joint rotation angular acceleration and other spontaneous motion data is carried out, which is as follows:
[0097] 1. Calculate the joint rotation angle:
[0098] The joint angle is obtained by calculating the included angle between the adjacent two bone vectors of the joint, and the bone vector is defined by the corresponding joint position information:
[0099]
[0100] where P A , P B , P C are the coordinates of the joint;
[0101] The joint rotation angle is:
[0102]
[0103] 2. Calculate the joint rotation angular velocity:
[0104] For two consecutive time points t1 and t2, and the corresponding joint rotation angles θ1 and θ2, the rotation angular velocity ω is:
[0105]
[0106] 3. Calculate the joint rotation angular acceleration:
[0107] Similar to the calculation of the joint rotation angular velocity, the angular acceleration can be calculated by the rate of change of the angular velocity:
[0108]
[0109] On the basis of the above-mentioned actively generated spontaneous motion data, further extract the whole body active motion feature F G as follows:
[0110] F G = [θ ij ω ij a ij ], i∈[1,12], j∈[1,6]
[0111] where θ, ω, a represent joint angle, joint angular velocity, and joint angular acceleration, respectively; i represents left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, and right ankle, etc. 12 human joints; j represents the maximum value, minimum value, mean value, peak-to-peak value (difference between maximum and minimum), root mean square value, standard deviation, etc. of each joint motion parameter (angle, angular velocity or angular acceleration).
[0112] In this embodiment, the logical regression (LR) machine learning method is used to obtain the feature layer fusion decision 1 (sub-decision) based on the whole body autonomous motion feature, and the support vector machine (SVM) machine learning method is used to obtain the feature layer fusion decision 2 (sub-decision) based on the local induced motion feature.
[0113] In this embodiment, the feature F GThe specific process of the logistic regression (LR) fusion decision 1 is as follows:
[0114] 1.1 Feature vector preprocessing: The features F G are preprocessed to ensure that the data has the same scale during training. The preprocessing standardization formula is as follows:
[0115]
[0116] where X' is the original data, μ is the mean, σ is the standard deviation, and X is the standardized data.
[0117] 1.2 Model training: The logistic regression model is used to train the fusion decision. The logistic regression obtains the model parameters by solving the maximum likelihood estimation (MLE). The prediction formula of the logistic regression model is:
[0118]
[0119] where w is the feature weight, X is the input feature, b is the bias term, and P(y=1|X) is the probability of class 1 under the input feature X.
[0120] 1.3 Decision output: The fusion decision 1 (sub-decision) can be obtained through the probability value output by the model. Generally, if the probability is greater than 0.5, it is considered as class 1 (neurodevelopmental abnormalities), otherwise as class 0 (neurodevelopmental normal).
[0121] In this embodiment, the features F L1 are extracted as follows:
[0122] Using human pose estimation models such as Openpose, Media Pipe Blaze Pose, the head key point coordinate position data of the small infant is obtained, and further local induced (visual) motion data such as head rotation angle, rotation angular velocity and rotation angular acceleration are calculated. The specific process is similar to the calculation of the joint motion parameters described above.
[0123] Using the OpenFace facial behavior analysis model, the gaze point motion trajectory of the small infant is obtained, and by calculating the distance between the stimulus light point motion trajectory and the gaze point motion trajectory, the gaze point motion trajectory error is obtained:
[0124]
[0125] where A(x, y) and A'(x', y') are the corresponding points and their coordinates on the stimulus light point motion trajectory and the small infant's gaze point motion trajectory, respectively.
[0126] Based on the local induced motion data, further local induced motion features F L1 are extracted as follows:
[0127] F L1 = [θ j ω j a j s j ], j e [1, 6]
[0128] Wherein, θ, ω, a, s represent head rotation angle, head rotation angular velocity, head rotation angular acceleration and gaze point motion trajectory error respectively; j represents the maximum value, minimum value, mean value, peak-to-peak value (difference between maximum value and minimum value), root mean square value, standard deviation and other characteristic values of head motion parameters (angle, angular velocity or angular acceleration) and eye motion parameters (gaze point motion trajectory error).
[0129] In this embodiment, the feature F L2 is extracted as follows:
[0130] Using human pose estimation models such as Openpose, Media Pipe Blaze Pose, the head key point coordinate position data of the small infant is obtained, and local induced (auditory) motion data such as head rotation angle, rotation angular velocity and rotation angular acceleration are further calculated, and the specific process is similar to the aforementioned joint motion parameter calculation.
[0131] On the basis of local induced motion data, the local induced motion feature F L2 is further extracted as follows:
[0132] F L2 = [θ j ω j a j ], j e [1, 6]
[0133] Wherein, θ, ω, a represent head rotation angle, head rotation angular velocity and head rotation angular acceleration; j represents the maximum value, minimum value, mean value, peak-to-peak value (difference between maximum value and minimum value), root mean square value, standard deviation and other characteristic values of head motion parameters (angle, angular velocity or angular acceleration).
[0134] In this embodiment, the feature F L3 is extracted as follows:
[0135] Using human pose estimation models such as Openpose, Media Pipe Blaze Pose, the head key point coordinate position data of the small infant is obtained, and local induced (audiovisual synchronization) motion data such as head rotation angle, rotation angular velocity and rotation angular acceleration are further calculated, and the specific process is similar to the aforementioned joint motion parameter calculation.
[0136] Using the OpenFace facial behavior analysis model, the gaze point motion trajectory of the small infant is obtained, and the gaze point motion trajectory error is obtained by calculating the distance from the motion trajectory of the stimulating light point. The specific process is similar to the foregoing F L1 The extraction process is as follows:
[0137] The pupil size feature calculation method has been described before.
[0138] On the basis of the local evoked motion data, the local evoked motion feature F L3 is further extracted as follows:
[0139] F L3 = [θ j ω j a j s j p j ], j ∈ [1, 6]
[0140] Wherein, θ, ω, a, s, p respectively represent head rotation angle, head rotation angular velocity, head rotation angular acceleration, gaze point motion trajectory error and pupil size; j represents the maximum value, minimum value, mean value, peak-to-peak value (difference between maximum value and minimum value), root mean square value, standard deviation and other characteristic values of head motion parameters (angle, angular velocity or angular acceleration), eye movement parameters (gaze point motion trajectory error) and pupil parameters.
[0141] In this embodiment, the feature F L4 is extracted as follows:
[0142] Using the human pose estimation model such as Openpose, Media Pipe Blaze Pose, the head key point coordinate position data of the small infant is obtained, and the local evoked (audiovisual asynchrony) motion data such as head rotation angle, rotation angular velocity and rotation angular acceleration is further calculated. The specific process is similar to the foregoing joint motion parameter calculation.
[0143] Using the OpenFace facial behavior analysis model, the gaze point motion trajectory of the small infant is obtained, and the gaze point motion trajectory error is obtained by calculating the distance from the motion trajectory of the stimulating light point. The specific process is similar to the foregoing F L1 The extraction process is as follows:
[0144] The pupil size feature calculation method has been described before.
[0145] On the basis of the local evoked motion data, the local evoked motion feature F L4 is further extracted as follows:
[0146] F L4 = [θ j ω j a js j p j ],j∈[1,6]
[0147] where θ, ω, a, s, p represent head rotation angle, head rotation angular velocity, head rotation angular acceleration, gaze point trajectory error and pupil size, respectively; j represents the maximum, minimum, mean, peak-to-peak value (difference between maximum and minimum), root mean square value, standard deviation and other characteristic values of head motion parameters (angle, angular velocity or angular acceleration), eye motion parameters (gaze point trajectory error) and pupil parameters.
[0148] Step 2: Support Vector Machine (SVM) fusion decision 2 based on features F L1 , F L2 , F L3 , F L4
[0149] 2.1 Data preprocessing: Similar to step 1.1, features F L1 , F L2 , F L3 , F L4 are standardized. The standardized feature vectors are concatenated to obtain the composite feature vector F L ={F L1 , F L2 , F L3 , F L4}, which is used as the model input.
[0150] 2.2 Support Vector Machine training: The SVM model is used to train the fusion decision based on local evoked motion features. The goal of SVM is to find the best separating hyperplane by maximizing the classification margin. For a binary classification problem, the optimization objective of SVM is:
[0151]
[0152] where w is the normal vector of the hyperplane, b is the bias term, x i is the input feature, and y i is the corresponding class label.
[0153] 2.3 Decision output: The input features are classified by the hyperplane trained by SVM to obtain fusion decision 2 (sub-decision). The prediction formula of SVM is:
[0154] f(x) = w T x + b
[0155] If the value of f(x) is greater than 0, it is classified as class 1; otherwise, it is classified as class 0.
[0156] In the embodiment, the final decision result is obtained by using weighted average of the sub-decisions, and whether the neural development of the small infant is normal is judged according to the final decision result, if normal, the small infant neural development state analysis work is ended, otherwise, the small infant neural development state analysis work is completed in combination with the doctor's clinical judgment.
[0157] The sub-decisions are combined into a fused decision vector D, and the decision vector is further analyzed to obtain the final decision about the task. The decision layer fusion calculates the final result based on the weighted average of multiple sub-decisions, and the process formula is as follows:
[0158]
[0159] Where d final is the final decision result, y is the weighted average, m is the number of sub-decisions, D is the fused decision vector combined by sub-decisions, W=(w1,w2,…,w m ) is the weight vector, and d final is 0, indicating that the neural development of the small infant is normal, and 1, indicating that it is abnormal, and the small infant neural development state analysis work needs to be completed in combination with the doctor's clinical judgment.
[0160] Decision layer fusion: In order to improve the fusion effect and automatically learn the nonlinear relationship between sub-decisions, the present application introduces a multi-layer perception (MLP) model. The specific process is as follows:
[0161] The MLP network is composed of multiple layers:
[0162] Input layer: the sub-decision D=(d1,d2,…,dm) is taken as the input vector of MLP, and each sub-decision can be binary output (0 or 1) or probability value.
[0163] Hidden layer: contains multiple neuron layers, which learn the complex relationship between sub-decisions through ReLU activation function, and the activation function formula is:
[0164]
[0165] Output layer: generate the final decision y, and the output value is mapped to the interval [0,1] through Sigmoid activation function, which is used to represent the final decision result:
[0166] y=σ(W out ·h+b out )
[0167] Where h is the output of the hidden layer, W out is the weight of the output layer, b out is the bias term of the output layer, and σ is the Sigmoid activation function:
[0168]
[0169] After the MLP is trained, the final decision d final is:
[0170]
[0171] d final is 0, indicating that the infant's neural development is normal, and 1, indicating that it is abnormal, and further analysis of the infant's neural development state needs to be completed in combination with the doctor's clinical judgment.
[0172] In summary, the infant neural development state analysis method in the above embodiment is based on whole-body spontaneous movement data collection and analysis, and increases local movement data collection and analysis induced by visual and auditory sensory organs, obtains multi-modal behavioral characteristics, and improves clinical classification accuracy through late fusion decision of modal characteristics
[0173] Although the embodiments of the present application have been shown and described above, it should be understood that the above embodiments are exemplary and should not be construed as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present application.
Claims
1. A system for analyzing the neurodevelopmental status of a small infant, comprising: The small infant neural development state analysis system comprises a test table, a display module, a sound module, a camera, a computer and a table side fixing frame, wherein the camera is integrated on the display module; the sound module is arranged on the left and right sides of the head of the small infant; the display module is fixed on the table side fixing frame, the relative position between the table side fixing frame and the test table is adjusted, so that the camera can collect the whole body movement video of the small infant; the computer is used for controlling the display module and the sound module and analyzing the video collected by the camera, and the display module is a naked-eye VR device; The small infant neural development state analysis system is used for small infant neural development state analysis, comprising the following steps: S1: data collection preparation The small infant is placed on the test table, the relative position between the table side fixing frame and the test table is adjusted, so that the normal direction of the camera on the display module is perpendicular to the test table and the field of view covers the whole body of the small infant, and the sound module is fixed on the left and right sides of the head of the small infant on the test table; S2: whole body active movement data collection The whole body active movement data of the small infant is collected by the camera; S3: local induced movement data collection The local induced movement data of the small infant is collected by the display module and / or the sound module in cooperation with the camera; S4: multi-movement feature extraction and fusion decision According to the movement paradigm, the features are extracted, the first classification sub-decision and the second classification sub-decision are formed through feature layer fusion, the final decision result is obtained by using the weighted average of the sub-decision, and then the small infant neural development state analysis and evaluation are carried out; In the step S2, the specific process is as follows: S21: the computer controls the camera to detect the whole body active movement of the small infant in a natural state, adjusts the relative position between the camera and the test table, so that the video contains complete four limbs and head information of the small infant; S22: the video is identified by using a human body posture estimation model, the four limb key points of the small infant are identified, and the movement feature data of the four limb key points are extracted; In the step S3, the specific process is as follows: Step S31: the computer controls the display module to present a light point and move at a high speed and a large amplitude in the display range, simultaneously, the camera detects the head rotation posture of the small infant in the visual following process, a human body posture estimation model is used to identify the head key points of the small infant, and the movement feature data of the head key points are extracted; Step S32: the computer controls the display module to present a light point and move at a low speed and a small amplitude in the display range, simultaneously, the camera detects the eye movement of the small infant in the visual following process, and the fixation point movement trajectory of the small infant under the visual induction of the light point is obtained; Step S33: the computer controls the sound module to emit continuous sound on the left and right sides of the head of the small infant, simultaneously, the camera detects the head rotation posture of the small infant in the auditory following process, a human body posture estimation model is used to identify the head key points of the small infant, and the movement feature data of the head key points are extracted; Step S34: The computer controls the display module and the sound module to be stimulated synchronously, the display module presents a light point and moves at a high speed and a large amplitude in the display range, and the sound module emits a sound in the direction of the motion endpoint of the light point. The camera detects the head rotation posture of the infant during the visual-auditory joint following process, uses a human pose estimation model to identify the head key points of the infant, and extracts the motion characteristic data of the head key points. Step S35: The computer controls the display module and the sound module to be stimulated synchronously, the display module presents a light point and moves at a low speed and a small amplitude in the display range, and the sound module emits a sound in the direction of the motion endpoint of the light point. The camera detects the eye movement and pupil change of the infant during the visual following process, and obtains the gaze point motion trajectory and pupil size under the visual-auditory joint induction of the light point. Step S36: The computer controls the display module and the sound module to be stimulated asynchronously, the display module presents a light point and moves at a high speed and a large amplitude in the display range, and the sound module emits a sound in the direction of the motion starting point of the light point. The camera detects the head rotation posture of the infant during the visual-auditory joint following process, uses a human pose estimation model to identify the head key points of the infant, and extracts the motion characteristic data of the head key points. Step S37: The computer controls the display module and the sound module to be stimulated asynchronously, the display module presents a light point and moves at a low speed and a small amplitude in the display range, and the sound module emits a sound in the direction of the motion starting point of the light point. The camera detects the eye movement and pupil change of the infant during the visual-auditory joint following process, and obtains the gaze point motion trajectory and pupil size under the visual-auditory joint induction of the light point. In the step S4, the motion paradigm includes whole-body active motion and local induced motion; the local induced motion includes local single-mode induced motion and local multi-mode induced motion, wherein the single mode is visual or auditory, and the multi mode is visual-auditory joint; the local multi-mode induced motion is divided into synchronous stimulation and asynchronous stimulation according to whether the visual and auditory stimulation directions are consistent. In the step S4, the features extracted from the whole-body voluntary motion data are F G , the features extracted from the visual modality evoked data are F L1 , the features extracted from the auditory modality evoked data are F L2 , the features extracted from the audiovisual modality synchronous evoked data are F L3 , and the features extracted from the audiovisual modality asynchronous evoked data are F L4 , wherein the whole-body voluntary motion data are the four-limb key point motion feature data obtained in the step S22, the visual modality evoked data are the head key point motion feature data and the gaze point motion trajectory obtained in the steps S31 and S32, the auditory modality evoked data are the head key point motion feature data obtained in the step S33, the audiovisual modality synchronous evoked data are the head key point motion feature data, the gaze point motion trajectory and the pupil size obtained in the steps S34 and S35, and the audiovisual modality asynchronous evoked data are the head key point motion feature data, the gaze point motion trajectory and the pupil size obtained in the steps S36 and S37.
2. The system for analyzing the neurodevelopmental state of a small infant according to claim 1, wherein Feature F G The extraction process is as follows: using a human pose estimation model, obtaining the coordinate position data of the four limb key points of the small infant, and further calculating the joint rotation angle, joint rotation angular velocity and joint rotation angular acceleration, and then extracting the feature F G As follows: F G = [θ ij ω ij a ij ], i ∈ [1, 12], j ∈ [1, 6] Wherein, θ, ω, a respectively represent joint rotation angle, joint rotation angular velocity, joint rotation angular acceleration; i represents left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, right ankle, etc. 12 human joints; j represents the maximum value, minimum value, mean value, peak-to-peak value, root mean square value, standard deviation of each joint motion parameter, including joint rotation angle, joint rotation angular velocity or joint rotation angular acceleration; Using a logistic regression machine learning method based on features F G The process of obtaining the first classification sub-decision is as follows: features F G Standardization processing; training the fusion decision using a logistic regression model; obtaining the first classification sub-decision through the probability value output by the logistic regression model.
3. The system for analysis of neurodevelopmental status of a small infant according to claim 2, characterized in that, Feature F L1 The extraction process is as follows: using a human pose estimation model, obtaining the key point coordinate position data of the head of the small infant, and further calculating the head rotation angle, head rotation angular velocity and head rotation angular acceleration, simultaneously using a facial behavior analysis model to obtain the motion trajectory of the fixation point of the small infant, calculating the distance from the motion trajectory of the stimulus light point to obtain the motion trajectory error of the fixation point, and further extracting the feature F L1 as follows: F L1 = [θ j ω j a j s j ],j∈[1,6] Wherein, θ, ω, a, s respectively represent head rotation angle, head rotation angular velocity, head rotation angular acceleration and gaze point motion trajectory error; j represents the maximum value, minimum value, mean value, peak-to-peak value, root mean square value, standard deviation of the head motion parameter and the eye movement parameter, the head motion parameter includes head rotation angle, head rotation angular velocity and head rotation angular acceleration, and the eye movement parameter is the gaze point motion trajectory error. Feature F L2 The extraction process is as follows: using a human pose estimation model, obtaining the key point coordinate position data of the head of the small infant, and further calculating the head rotation angle, head rotation angular velocity and head rotation angular acceleration, and then extracting the feature F L2 As follows: F L2 = [θ j ω j a j ], j e [1, 6] Wherein, θ, ω, a respectively represent head rotation angle, head rotation angular velocity and head rotation angular acceleration; j represents maximum value, minimum value, mean value, peak-to-peak value, root mean square value, standard deviation of head movement parameters, including head rotation angle, head rotation angular velocity and head rotation angular acceleration; Feature F L3 The extraction process is as follows: using a human pose estimation model, obtaining the key point coordinate position data of the head of the small infant, and further calculating the head rotation angle, head rotation angular velocity and head rotation angular acceleration, simultaneously using a facial behavior analysis model to obtain the motion trajectory of the fixation point of the small infant, calculating the distance from the motion trajectory of the stimulus light point to obtain the motion trajectory error of the fixation point, and then extracting the feature F L3 as follows: F L3 = [θ j ω j a j s j p j ],j∈[1,6] Wherein, θ, ω, a, s, p respectively represent head rotation angle, head rotation angular velocity, head rotation angular acceleration, gaze point motion trajectory error and pupil size; j represents maximum value, minimum value, mean value, peak-to-peak value, root mean square value, standard deviation of head movement parameters, eye movement parameters and pupil parameters; head movement parameters include head rotation angle, head rotation angular velocity and head rotation angular acceleration, eye movement parameters are gaze point motion trajectory error, and pupil parameters are pupil size; Feature F L4 The extraction process is as follows: using a human pose estimation model, obtaining the key point coordinate position data of the head of the small infant, and further calculating the head rotation angle, head rotation angular velocity and head rotation angular acceleration, simultaneously using a facial behavior analysis model to obtain the motion trajectory of the fixation point of the small infant, calculating the distance from the motion trajectory of the stimulus light point to obtain the motion trajectory error of the fixation point, and further extracting the feature F L4 as follows: F L4 = [θ j ω j a j s j p j ],j∈[1,6] Wherein, θ, ω, a, s, p respectively represent head rotation angle, head rotation angular velocity, head rotation angular acceleration, gaze point motion trajectory error and pupil size; j represents maximum value, minimum value, mean value, peak-to-peak value, root mean square value, standard deviation of head movement parameters, eye movement parameters and pupil parameters; head movement parameters include head rotation angle, head rotation angular velocity and head rotation angular acceleration, eye movement parameters are gaze point motion trajectory error, and pupil parameters are pupil size; Support vector machine machine learning method based on features F L1 , F L2 , F L3 , F L4 The process of obtaining the first classification sub-decision is as follows: the features F L1 , F L2 , F L3 , F L4 are standardized; the SVM model is used to train the fusion decision; the input features are classified by the hyperplane trained by the SVM to obtain the second classification sub-decision.
4. The system for analysis of neurodevelopmental status of a small infant according to claim 3, characterized in that, The calculation formula of the gaze point motion trajectory error is as follows: Wherein, A(x, y), A'(x', y') are corresponding points and coordinates thereof on the stimulation light point motion trajectory and the small infant gaze point motion trajectory respectively.
5. The system for analysis of neurodevelopmental status of a small infant according to claim 4, characterized in that, In the step S4, the sub-decision groups are combined into a fused decision vector D, and the decision vector is further analyzed to obtain a final decision, and the decision layer fusion calculates the final result based on the weighted average value of multiple sub-decisions, and the process formula is as follows: wherein d final is the final decision result, y is the weighted average value, m is the number of sub-decisions, D is the fusion decision vector of the sub-decision combination, W = (w1, w2, …, w m ) is the weight vector, and d final is 0, indicating that the neurological development of the small infant is normal, and is 1, indicating that the neurological development of the small infant is abnormal.
Citation Information
Patent Citations
Children autism rehabilitation training system and method based on floating interaction
CN118550409A
Method for evaluating a risk of neurodevelopmental disorder with a child
CN112384990A
Intelligent diagnosis equipment for evaluating early-stage neural function of infant
CN116392072A