Video-based 3D cerebral palsy baby action identification method and system
By constructing and training key point recognition models and action recognition models, the problem of low accuracy of infant movement posture recognition in the prior art is solved, and the rapid and accurate recognition of infant movements is achieved, providing support for health monitoring and cerebral palsy assessment.
Patent Information
- Application Number
- CN202510246952.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art has low accuracy in infant movement posture recognition, and relies on manual monitoring, high professional level requirements, making it difficult to effectively evaluate infant neurodevelopment.
By obtaining the baby's video file, pre-processing and frame extraction, annotating the two-dimensional image coordinates and three-dimensional world coordinates of key points, building a key point recognition model and action recognition model, and training to achieve accurate recognition of infant actions.
Improves the accuracy of 3D infant pose estimation, enabling rapid and accurate extraction of infant movement postures to support infant health monitoring and cerebral palsy assessment.
Smart Images

Figure CN120183040A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of infant motion recognition, and more specifically, to a method and system for 3D cerebral palsy infant motion recognition based on video. Background Art
[0002] General Movements (GMs) assessment is a method for evaluating the neuro-motor behavior of newborns and infants, which has the characteristics of safety, reliability, sensitivity, non-invasiveness, etc., and can effectively evaluate the function of the infant nervous system. General Movements (GMs) refer to movements involving the whole body, and the arms, legs, neck, and trunk participate in this whole-body movement in a way of changing the movement sequence. Applying GMs quality assessment within 4 to 5 months after the birth of high-risk newborns can make accurate and effective predictions about the later neurodevelopmental outcomes. The early movement and posture assessment of infants is crucial for timely detection of diseases such as cerebral palsy. However, the current movement assessment methods still rely on manual monitoring and judgment, and the diagnostic results depend on the experience of clinicians, requiring a relatively high professional level of staff; while the automated recognition schemes for action postures are all used for adults, and the recognition accuracy for infant action postures is relatively low. Therefore, how to provide a method and system for 3D cerebral palsy infant motion recognition based on video is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0003] In view of this, the present invention provides a method and system for 3D cerebral palsy infant motion recognition based on video, which improves the accuracy of 3D infant pose estimation in health monitoring videos.
[0004] To achieve the above object, the present invention provides the following technical solutions:
[0005] A method for 3D cerebral palsy infant motion recognition based on video, comprising the following steps:
[0006] S1. Obtain video files of normal infant wriggling movements and abnormal monotonous movements respectively and preprocess them;
[0007] S2. Perform frame extraction on the preprocessed video files to obtain image data of the infant;
[0008] S3. Mark the two-dimensional image coordinates and three-dimensional world coordinates of key points in the image data of the infant;
[0009] S4. Make a data set based on the marked image data of the infant, and the data set is used for training the key point recognition model and the action recognition model;
[0010] S5. Build a key point recognition model, using the images in the dataset as input data and the two-dimensional image coordinates of the corresponding key points as sample labels for training;
[0011] S6. Build an action recognition model, using the two-dimensional image coordinates of the key points as input data and the corresponding three-dimensional world coordinates as sample labels for training;
[0012] S7. Identify the baby's actions through the trained key point recognition model and action recognition model.
[0013] Optionally, S1 is specifically: Collect the baby's video files. The physician classifies the video files into two types: normal twisting actions and abnormal monotonous movements according to the baby's movement conditions, and then uniformly adjusts the resolution of all videos to 720×1280.
[0014] Optionally, S2 is specifically: Set the frame extraction intervals for the video files of normal twisting actions and abnormal monotonous movements respectively, and extract one image every certain number of frames; if the baby has no actions for a period of time, skip the corresponding time period without frame extraction; for the video files of abnormal monotonous movements, when the baby's actions change significantly in a video, reduce the corresponding frame extraction interval; if the image is blocked, select images at a certain adjacent frame interval for replacement; after frame extraction, save the image data and record the seconds and frames corresponding to the images.
[0015] Optionally, S3 is specifically: Set the key points of the baby in the image to be the nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle and right ankle, and label the key points using the labelme annotation tool.
[0016] Optionally, S4 is specifically:
[0017] Convert the JSON file generated by the labelme annotation tool into a JSON file in coco format, and then convert the JSON file into a cocotxt file; save the image data after frame extraction, the labelme annotation standard file, the cocotxt file and the original video data.
[0018] Optionally, the key point recognition model includes a backbone network layer, a feature fusion layer, a detection layer and a prediction layer. The backbone network layer extracts 4 feature maps of different scales based on the input image. The feature fusion layer fuses the feature map with the feature maps of different scales output by the backbone network layer. The detection layer performs baby detection based on the 4 fused feature maps of different scales. The prediction layer outputs the predicted bounding box and the predicted key point positions.
[0019] Optionally, the action recognition model includes a generator network and a discriminator network. The generator network generates the three-dimensional world coordinates of the key points based on the two-dimensional image coordinates of the input key points. The input of the discriminator network is the three-dimensional world coordinates generated by the generator network and the real three-dimensional world coordinates. The discriminator network determines whether the input coordinates are the three-dimensional world coordinates generated by the generator network or the real three-dimensional world coordinates by reconstructing the input coordinates to obtain new three-dimensional world coordinates.
[0020] A video-based 3D cerebral palsy infant action recognition system that executes the above-mentioned video-based 3D cerebral palsy infant action recognition method, including:
[0021] A video data acquisition module that acquires the monitoring video data of the infant through an image acquisition device;
[0022] A video data processing module that processes the monitoring video data and converts the video monitoring data into image data;
[0023] A key point recognition module that recognizes the two-dimensional image coordinates of the key points of the infant in the image based on a key point recognition model;
[0024] An action recognition module that converts the two-dimensional image coordinates of the key points into three-dimensional world coordinates based on an action recognition model.
[0025] From the above technical solutions, it can be seen that compared with the prior art, the present invention provides a video-based 3D cerebral palsy infant action recognition method and system, which has the following beneficial effects: By making a data set containing the normal twisting actions and abnormal monotonous movement images of the infant, the present invention constructs and trains a key point recognition model and an action recognition model for recognizing the actions of the infant, and can quickly and accurately extract the action postures based on the monitoring video data of the infant, which can be applied to the health monitoring of the infant and provide support for the evaluation of cerebral palsy infants. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0027] Figure 1 It is a flow chart of the cerebral palsy infant action recognition method of the present invention;
[0028] Figure 2 It is a schematic diagram of the principle of the key point recognition model of the present invention;
[0029] Figure 3 It is a schematic diagram of the generator network in the embodiment of the present invention;
[0030] Figure 4 This is a schematic diagram of the discriminator network in the embodiments of the present invention. Detailed implementation manners
[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0032] The embodiments of the present invention disclose a method for recognizing 3D cerebral palsy infant movements based on videos, as Figure 1 shown, including the following steps:
[0033] S1. Obtain video files of the normal twisting movements and abnormal monotonous movements of the infant respectively and preprocess them;
[0034] S2. Perform frame extraction on the preprocessed video files to obtain the image data of the infant;
[0035] S3. Mark the two-dimensional image coordinates and three-dimensional world coordinates of the key points in the image data of the infant;
[0036] S4. Make a data set based on the marked image data of the infant, and the data set is used for training the key point recognition model and the action recognition model;
[0037] S5. Construct a key point recognition model, take the images in the data set as input data, and use the two-dimensional image coordinates of the corresponding key points as sample labels for training;
[0038] S6. Construct an action recognition model, take the two-dimensional image coordinates of the key points as input data, and use the corresponding three-dimensional world coordinates as sample labels for training;
[0039] S7. Recognize the movements of the infant through the trained key point recognition model and action recognition model.
[0040] Further, S1 is specifically: collect the video files of the infant, and the physician divides the video files into two types: normal twisting movements and abnormal monotonous movements according to the movement conditions of the infant, and then uniformly adjusts the resolution of all videos to 720×1280.
[0041] In the embodiments of the present invention, the frame rate of the collected video file data is 25 or 30, and there are three resolutions: 720*1280, 1920*1080, 960*544; the video is adjusted to a unified resolution of 720×1280 (width × height).
[0042] Further, S2 is specifically as follows: separately set the frame extraction intervals of the video files of normal writhing movements and abnormal monotonous movements, and extract one image every certain number of frames; if the baby has no movement for a period of time, skip the corresponding time period and do not perform frame extraction; for the video file of abnormal monotonous movement, when the baby's movements change significantly within a video, reduce the corresponding frame extraction interval; if the image is blocked, select images at a certain adjacent frame interval for replacement; after frame extraction is completed, save the image data and record the seconds and frames corresponding to the images.
[0043] In the embodiment of the present invention, one image can be selected every 10 frames (or 15, 20 frames). For the video file of abnormal monotonous movement, when the baby's movements change significantly within a video, 5-frame sampling (or even a smaller interval) can be adopted.
[0044] Further, S3 is specifically as follows: set the key points of the baby in the image as the nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle and right ankle respectively, and label the key points through the labelme annotation tool.
[0045] Further, S4 is specifically as follows:
[0046] Convert the JSON file generated by the labelme annotation tool into a JSON file in coco format, and then convert the JSON file into a cocotxt file; save the image data after frame extraction, the labelme annotation standard file, the cocotxt file and the original video data.
[0047] Further, the key point recognition model includes a backbone network layer, a feature fusion layer, a detection layer and a prediction layer. The principle of the key point recognition model is as Figure 2 shown. The backbone network layer extracts 4 feature maps of different scales based on the input image. The feature fusion layer fuses the feature map with the feature maps of different scales output by the backbone network layer. The detection layer performs baby detection based on the 4 feature maps of different scales after fusion. The prediction layer outputs the predicted bounding box and the predicted positions of the key points.
[0048] In the embodiment of the present invention, the key point recognition model is constructed based on YOLOv5-pose. The backbone network layer in this embodiment includes an input layer, a convolutional layer, a feature extraction module, and an SPP layer connected in sequence. Among them, the feature extraction module consists of 7 feature extraction units and 2 upsampling layers. The upsampling layers are respectively connected to the outputs of the 2nd and 4th feature extraction units and connected to the input of the feature fusion layer. The feature extraction unit includes an extended convolutional layer, a depthwise separable convolutional layer, an SE module, and a skip connection layer. The extended convolutional layer expands the dimension of the input feature. The depthwise separable convolutional layer decomposes the standard convolution into a depth convolution and a pointwise convolution to reduce the computational amount. The SE module weights the features of each channel. The skip connection is used to retain the input feature;
[0049] The loss function of the key point recognition model is as follows:
[0050]
[0051] In the formula, L, L box , L key , L c are respectively the loss of the key point recognition model, the prediction box loss, the key point loss, and the confidence loss. λ box , λ key , λ c are respectively the weights of L box , L key , L c . s represents the size, k represents the prediction box, and i, j represent the horizontal and vertical coordinates of the true bounding box. The prediction box loss uses the CIOU loss function, the key point loss uses the OKS loss function, and the confidence loss uses the binary cross-entropy loss function.
[0052] Furthermore, the action recognition model includes a generator network and a discriminator network. The generator network generates the three-dimensional world coordinates of the key points based on the two-dimensional image coordinates of the input key points. The input of the discriminator network is the three-dimensional world coordinates generated by the generator network and the true three-dimensional world coordinates. The discriminator network determines whether the input coordinates are the three-dimensional world coordinates generated by the generator network or the true three-dimensional world coordinates based on the newly reconstructed three-dimensional world coordinates from the input coordinates.
[0053] In the embodiment of the present invention, as Figure 3 shown, the generator network includes a convolutional unit, a residual module, and a convolutional layer connected in sequence. Among them, the convolutional unit consists of a convolutional layer, a batch normalization layer, and a ReLU activation function. The residual module consists of 4 skip-connected residual units, and each residual unit consists of 2 convolutional units; as Figure 4 shown, the discriminator network consists of 3 convolutional units, 1 convolutional layer, and 4 residual units;
[0054] The loss function L of the generator network G is as follows:
[0055]
[0056] In the formula, N represents the total number of training samples, i ∈ [1, N], and X i represents the predicted value of the i-th sample generated by the generator network, represents X i corresponding true value, represents the output result when the input of the discriminator network is ;
[0057] The loss function of the discriminator network is:
[0058]
[0059] In the formula, k is the balance variable of the discriminator network, and its value range is 0 - 1;
[0060] During the training process of the action recognition model, first determine the parameters of the discriminator network, train the generator network, calculate the loss function of the generator network, update the parameters of the generator network, then determine the parameters of the generator network, calculate the loss function of the discriminator network, update the parameters of the discriminator network, and alternately update the parameters of the generator network and the discriminator network until the training ends.
[0061] Corresponding to Figure 1 the method described above, an embodiment of the present invention also discloses a video-based 3D cerebral palsy infant action recognition system, which executes the above-mentioned video-based 3D cerebral palsy infant action recognition method, including:
[0062] A video data acquisition module, which acquires the monitoring video data of the infant through an image acquisition device;
[0063] A video data processing module, which processes the monitoring video data and converts the video monitoring data into image data;
[0064] A key point recognition module, which recognizes the two-dimensional image coordinates of the key points of the infant in the image based on the key point recognition model;
[0065] An action recognition module, which converts the two-dimensional image coordinates of the key points into three-dimensional world coordinates based on the action recognition model.
[0066] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0067] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A video-based 3D cerebral palsy infant action recognition method, characterized in that: The following steps are involved: S1, respectively obtaining video files of normal twisting movements and abnormal monotonous movements of infants and preprocessing them; S2, extracting frames from the preprocessed video file to obtain image data of the baby; S3, marking the two-dimensional image coordinates and three-dimensional world coordinates of key points in the image data of the infant; S4. Create a data set based on the labeled infant image data, and the data set is used to train the key point recognition model and the action recognition model; S5. Build a key point recognition model, use the images in the data set as input data, and the corresponding two-dimensional image coordinates of the key points as sample labels for training; S6. Build an action recognition model, use the 2D image coordinates of key points as input data, and the corresponding 3D world coordinates as sample labels for training; S7. Recognize the baby's actions through the trained key point recognition model and action recognition model.
2. A video-based 3D cerebral palsy infant action recognition method according to claim 1, characterized in that: S1 specifically includes: collecting video files of infants, and the doctor classifies the video files into two types according to the infants' movements: normal twisting movements and abnormal monotonous movements, and then uniformly adjusts the resolution of all videos to 720×1280.
3. The video-based 3D cerebral palsy infant action recognition method according to claim 1, characterized in that: S2 specifically includes: setting the frame extraction intervals of the video files of normal twisting movements and abnormal monotonous movements respectively, and extracting an image every certain number of frames; if the baby does not move for a period of time, skipping the corresponding period of time and not extracting frames; For video files with abnormally monotonous movements, when the baby's movements change significantly within a video, the corresponding frame extraction interval is reduced; If the image is blocked, select adjacent images with a certain frame interval to replace it; after the frame extraction is completed, save the image data and record the seconds and frames corresponding to the image.
4. The video-based 3D cerebral palsy infant action recognition method according to claim 1, characterized in that: S3 is specifically: setting the key points of the baby in the image to be nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle and right ankle, and annotating the key points using the labelme annotation tool.
5. The video-based 3D cerebral palsy infant action recognition method according to claim 4, characterized in that: S4 is specifically: Convert the JSON file generated by the labelme annotation tool into a JSON file in coco format, and then convert the JSON file into a cocotxt file; save the image data after frame extraction, the labelme annotation standard file, the cocotxt file, and the original video data.
6. The video-based 3D cerebral palsy infant action recognition method according to claim 1, characterized in that: The key point recognition model includes a backbone network layer, a feature fusion layer, a detection layer and a prediction layer. The backbone network layer extracts feature maps of four different scales based on the input image. The feature fusion layer fuses the feature map with the feature maps of different scales output by the backbone network layer. The detection layer performs baby detection based on the fused feature maps of four different scales. The prediction layer outputs the predicted bounding box and the predicted key point positions.
7. The video-based 3D cerebral palsy infant action recognition method according to claim 1, characterized in that: The action recognition model includes a generator network and a discriminator network. The generator network generates the three-dimensional world coordinates of the key points based on the two-dimensional image coordinates of the key points input. The input of the discriminator network is the three-dimensional world coordinates generated by the generator network and the real three-dimensional world coordinates. The discriminator network reconstructs new three-dimensional world coordinates based on the input coordinates to determine whether the input coordinates are the three-dimensional world coordinates generated by the generator network or the real three-dimensional world coordinates.
8. A video-based 3D cerebral palsy infant action recognition system, characterized in that: The method for video-based 3D cerebral palsy infant action recognition according to any one of claims 1 to 7 comprises: A video data acquisition module collects monitoring video data of the baby through an image acquisition device; A video data processing module processes the surveillance video data and converts the video surveillance data into image data; A key point recognition module, which recognizes the two-dimensional image coordinates of the key points of the baby in the image based on the key point recognition model; The action recognition module converts the two-dimensional image coordinates of key points into three-dimensional world coordinates based on the action recognition model.
Citation Information
Cited By
Infant action generative modeling and classification
WO2026025104A1