Deep learning-based multi-view children motion coordination ability evaluation system and method
Through a multi-perspective children's motor coordination ability assessment system based on deep learning, using a dual-camera layout and deep convolutional neural network, an efficient and automated assessment of children's motor coordination ability is achieved, solving the problems of inconsistent assessment and lack of accuracy in existing technologies, and generating personalized reports.
Patent Information
- Application Number
- CN202510916539.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-07-03
AI Technical Summary
Existing technologies make it difficult to accurately capture children's motor coordination, completion time, and number of times through multi-perspective fusion technology, and are unable to achieve efficient and automated multi-person synchronous evaluation, resulting in inconsistent and inaccurate evaluation results.
The multi-perspective children's motor coordination ability assessment system based on deep learning adopts a dual-camera layout combined with deep convolutional neural networks and regression methods. Through the skeleton point detection model and motor ability assessment model, it realizes multi-perspective data association and personalized health assessment.
It realizes automated and accurate dynamic movement assessment of children, eliminates human errors, improves the consistency and efficiency of assessment, supports simultaneous testing of multiple people, and generates personalized health assessment reports.
Smart Images

Figure CN120809152A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of motion evaluation, and particularly relates to a multi-view child motion coordination ability evaluation system and method based on deep learning. BACKGROUND
[0002] Child motion coordination ability is a core observation index of neural development, and its abnormal performance is closely related to Developmental Coordination Disorder (DCD). According to statistics of the World Health Organization, about 5%-6% of school-age children worldwide have different degrees of motion coordination defects, which are manifested as abnormal gait symmetry, delayed motion start, bilateral coordination disorder, etc., and severe cases can lead to decreased learning ability and social avoidance behavior. The internationally recognized Movement Assessment Battery for Children-Second Edition (MABC-2) tool points out that a fine motion control error exceeding 1.5 standard deviations can be determined as abnormal. In a specific motion test scene like a sideways walking test, precise evaluation of the coordination of the shoulder and footstep, motion continuity, and bilateral symmetry, etc. can understand the neural muscle coordination ability of children, such as the parallelism of the shoulder and the motion direction, the parallelism of the shoulder and the ground, which can reflect the trunk stability, and the angle between the shoulder direction and the footstep moving direction, which represents the overall coordination and gait control ability. However, if a child shows obvious shoulder shaking or an inclination angle exceeding 5° during the test, it may mean that the core muscle group control is insufficient or the development of the vestibular function is lagging behind, and uneven sliding amplitude of the legs, short or insufficient air time may be related to limited cerebellum function regulation. However, traditional evaluation methods are difficult to capture millisecond-level coordination deviations in dynamic motion.
[0003] Traditional child motion evaluation mainly relies on manual observation and simple physical tests. Due to the limited experience and attention of the evaluators, it is difficult to synchronously track multiple angles and key details of the child's motion, such as the horizontal deviation of the shoulder, the imbalance of the gait, or the deficiency in the air phase, which leads to difficulty in unifying the evaluation standards. The same child may get different evaluation results under the tests of different evaluators, affecting the credibility of the final evaluation. On the other hand, traditional methods lack high-precision time-series data and trajectory analysis, making it difficult to provide quantifiable and repeatable objective indicators, which makes it difficult to accurately identify early abnormalities such as Developmental Coordination Disorder (DCD) or sensory integration dysfunction, thus easily missing the intervention opportunity.
[0004] Chinese patent CN119533462A discloses a motion estimation method, device and electronic equipment based on multi-modal perception, the method includes obtaining motion evaluation related data of the target object target part region, including IMU sensor data, flexible deformation sensor data, the pose information of the part where the IMU sensor is configured can be determined based on the IMU sensor data, then according to the position relationship between the configured IMU sensor and the flexible deformation sensor and the obtained flexible deformation sensor data, the pose information of the part where the flexible deformation sensor is configured is calculated, so as to obtain the pose information of the position where the IMU sensor is not configured under the condition that the IMU sensor is configured. However, the existing method needs to wear equipment, can only measure the movement of the body part, cannot cover all parts of the body, and the precision is limited, like in some whole body movement coordination evaluation scenarios, the comprehensiveness and accuracy of the evaluation results may be affected due to the inability to comprehensively obtain the movement data of each part of the body.
[0005] Chinese patent CN117457193A discloses a body health monitoring method and system based on human key point detection, the method includes the following steps: based on optical human image acquisition, human key point detection, human detection frame identification, multi-human tracking, human key part posture detection, and bad body posture early warning. At the same time, the method can detect the body health of multiple people appearing in the visual range, and give early warning of bad body posture respectively, and generate a body health report; but the existing method mainly aims at single body posture detection, and then realizes the evaluation of child bone development status, which is difficult to adapt to the flexible and changeable movement characteristics and behavior patterns of children, and cannot evaluate the movement coordination, frequency, time and other dimensions.
[0006] Therefore, for the special needs of children's coordination movement evaluation, the existing technology has significant deficiencies: children have a large range of activities and fast changes in action, making it difficult to perform efficient daily group evaluation, which requires the system to have core capabilities such as multi-angle synchronous acquisition, multi-person real-time tracking, and high-precision action analysis. There is an urgent need to develop a new evaluation scheme that can maintain the naturalness of non-contact measurement while accurately capturing fast and variable movement characteristics through multi-view fusion technology to identify action coordination, completion time, and completion frequency. At the same time, it supports multi-person synchronous evaluation, and truly meets the large-scale and normalized monitoring needs of children's movement ability in kindergarten, school and other actual scenarios. In view of the above problems, we propose a multi-view child movement coordination ability evaluation system and method based on deep learning. SUMMARY
[0007] The purpose of the present application is to solve the problem that the prior art mainly relies on artificial observation and simple physical tests for child movement evaluation, and due to the limited experience and attention of the evaluators, it is difficult to synchronously track multiple angles and key details of the child's action, so that the existing method cannot automatically, accurately and efficiently evaluate the child's dynamic action, especially when it involves complex movement, the traditional method is difficult to provide comprehensive, efficient and accurate standard unified evaluation.
[0008] The present application is implemented in the following way: the deep learning-based multi-view child movement coordination ability evaluation method comprises:
[0009] A special child movement dataset is constructed, and the movement evaluation dataset is captured based on the child movement dataset; the child movement dataset and the movement evaluation dataset are integrated into a modeling dataset; the modeling dataset is used to train a skeleton point detection model and a movement ability evaluation model; and the converged skeleton point detection model and the movement ability evaluation model are outputted;
[0010] Real-time video streams of personnel entering the site are collected by the collection camera, and the real-time video streams are preprocessed to output a preprocessing dataset;
[0011] The preprocessing dataset is loaded, and the skeleton point detection model identifies and analyzes the preprocessing dataset based on the combination of a deep convolutional neural network and a regression method, identifies the skeleton key points in the preprocessing dataset, and obtains a skeleton key point detection result;
[0012] The skeleton key point detection result is obtained, and the skeleton key point detection result is associated with multi-view data to associate the person ID with the corresponding skeleton key points, obtain a personnel matching result containing the person ID and the skeleton key points, and perform multi-person detection on the personnel matching result based on a movement region division method to obtain a multi-person detection result;
[0013] The multi-person detection result is taken as input, and the movement ability evaluation model is executed; the movement ability evaluation model quantitatively analyzes the multi-person detection result to generate an individualized health evaluation report.
[0014] Preferably, the method for constructing the child movement dataset comprises:
[0015] Special video data of children in different scenes are collected by using a multi-angle and full-body shooting method, wherein the video data covers common movement actions such as running, jumping and side sliding, and also covers natural actions of children in play or daily activities;
[0016] loading the special video data, performing frame-by-frame analysis and processing on the special video data, obtaining at least one group of special images, performing child skeleton point marking processing on the special images, and obtaining a child motion data set containing skeleton point marking results, wherein, when performing child skeleton point marking processing on the special images, fixed numbers are set for 17 skeleton points of the child, which are respectively: 1. nose, 2. left eye, 3. right eye, 4. left ear, 5. right ear, 6. left shoulder, 7. right shoulder, 8. left elbow, 9. right elbow, 10. left wrist, 11. right wrist, 12. left hip joint, 13. right hip joint, 14. left knee, 15. right knee, 16. left ankle, and 17. right ankle;
[0017] Based on the child motion data set, the motion evaluation data set is captured, wherein the motion evaluation data set is based on the abnormal motion data collection of the child with motion coordination deviation, and the child motion data set and the motion evaluation data set are integrated into the modeling data set.
[0018] Preferably, the method of training the skeleton point detection model and the motion ability evaluation model by using the modeling data set comprises:
[0019] loading the modeling data set, dividing the modeling data set into a training set and a validation set, wherein the training set is used to train the skeleton point detection model and the motion ability evaluation model, and the validation set is used to verify the generalization ability of the skeleton point detection model and the motion ability evaluation model, the proportion of the training set and the validation set is 4:1;
[0020] training the skeleton point detection model iteratively by using the training set and the validation set, and outputting the converged skeleton point detection model, wherein the training set includes skeleton point labeling data and motion data of normal and abnormal children, and the training set is used to train the skeleton point detection model and the motion ability evaluation model, so that the skeleton point detection model can identify the joint positions of children in different motion actions and distinguish normal and abnormal motion patterns;
[0021] loading the skeleton point detection model, the skeleton point detection model combines the data of the front view and the side view, and stably identifies the skeleton points of the child under different viewing angles, and outputs the skeleton point recognition result of the skeleton point detection model;
[0022] obtaining the skeleton point recognition result, training and optimizing the motion ability evaluation model based on the skeleton point recognition result, and outputting the motion ability evaluation model based on the multi-view sequence skeleton point image data;
[0023] The motion ability evaluation model comprises a feature extraction module and a regression scoring module, the feature extraction module comprises a graph construction layer, a graph convolution network (GCN), a recurrent neural network (RNN), an average pooling layer and a feature fusion layer connected in sequence, and the regression scoring module is used for finally outputting a health evaluation score, the regression scoring module comprises a full connection layer 1, a full connection layer 2 and an output layer connected in sequence, the full connection layer 1 is used for acquiring a feature vector output by the feature fusion layer, the full connection layer 1 and the full connection layer 2 pass data through a ReLu activation function, the full connection layer 2 and the output layer also pass data through the ReLu activation function, and the output layer outputs a score vector through a Sigmoid activation function.
[0024] Preferably, the method for pre-processing the real-time video stream comprises:
[0025] The acquisition camera acquires a video stream in real time at a fixed frame rate of 60 FPS to obtain a real-time video stream;
[0026] The real-time video stream is loaded, and the real-time video stream is processed frame by frame, the real-time images processed frame by frame are numbered and stored in time sequence;
[0027] The real-time images processed in time sequence are subjected to format conversion, grayscale normalization processing and a smoothing filtering method to remove noise in the images to obtain a pre-processed data set.
[0028] Preferably, the method for identifying the skeletal key points in the pre-processed data set comprises:
[0029] The pre-processed data set is loaded, a skeletal point detection model extracts key features of the pre-processed data set through a convolution layer, identifies multiple skeletal point candidate positions in the pre-processed data set, and preliminarily locates the skeletal points;
[0030] The preliminary locating result of the skeletal points is obtained, the preliminary locating result of the skeletal points is classified and accurately classified according to a human body structure through a clustering algorithm combined with regression analysis to obtain a skeletal key point detection result;
[0031] In the classification and accurate classification of the preliminary locating result of the skeletal points according to the human body structure, each class is regarded as a target, and the skeletal point number set of each target is output first:
[0032] P i (t)={p1,p2,...,p n}
[0033] Where p kThe numbers represent the bone point numbers: 1. nose, 2. left eye, 3. right eye, 4. left ear, 5. right ear, 6. left shoulder, 7. right shoulder, 8. left elbow, 9. right elbow, 10. left wrist, 11. right wrist, 12. left hip joint, 13. right hip joint, 14. left knee, 15. right knee, 16. left ankle, 17. right ankle;
[0034] Each number corresponds to a two-dimensional coordinate, representing the position of the bone point in the image:
[0035] C i (t) = {(x1, y1), (x2, y2),..., (x n , y n )}
[0036] Where (x k , y k ) is the coordinate of the bone point p k on the image.
[0037] Preferably, the method for associating multi-view data of the skeleton key point detection result comprises:
[0038] Load the skeleton key point detection result, identify the image-related collection camera in the skeleton key point detection result, use the host timestamp as the unified standard to align the timestamps of the collection video of the collection camera, and output the timestamp-synchronized image data: I front (t), I side (t).
[0039] Pre-calibrate the strict mapping relationship between the actual physical space of the site and the collection camera screen, and associate multi-view data of the skeleton key point detection result according to the pre-calibrated mapping relationship to obtain a personnel matching result containing the person ID and the skeleton key point. The mapping relationship is established through accurate collection camera installation position, measurement and perspective calibration experiment. When the target is located in the left upper corner of the physical area of the site, the matching is realized according to the pre-calibrated mapping relationship through the following specific steps: first, real-time detect all target bone point data appearing in the right side area of the front view screen, and simultaneously detect the target bone point data in the left side area of the side view screen, then perform spatial consistency verification to check whether the relative position relationship of the two sets of bone point data conforms to the pre-calibrated spatial mapping rule, then perform motion state verification to verify whether the motion trends reflected by the two sets of bone point data are consistent, and when all the verifications pass, the same ID number is assigned to the two sets of bone point data, i.e.
[0040]
[0041] Combine the ID number with the bone point data of each person to represent the corresponding relationship between the number and the bone point data:
[0042]
[0043] The multi-person detection result is obtained by performing multi-person detection on the personnel matching result based on a motion region division method.
[0044] Preferably, the method for quantitatively analyzing the multi-person detection result by the sports ability evaluation model comprises:
[0045] The multi-person detection result is loaded, and the multi-person detection result is evaluated based on the sports ability evaluation model to obtain a child health evaluation score corresponding to the health standard;
[0046] The sports ability evaluation model analyzes the health evaluation score based on the health standard, judges whether the health evaluation score is qualified, and obtains a health evaluation result;
[0047] The sports ability evaluation model generates a personalized health evaluation report in combination with the health evaluation result;
[0048] The personalized health evaluation report comprises evaluation index analysis, health score and suggestion, and training suggestion.
[0049] The method for evaluating the multi-person detection result based on the sports ability evaluation model comprises:
[0050] The multi-person detection result is obtained, and the pixel coordinates of the skeleton points in the front view and the side view, the differential values of the pixel coordinates of the skeleton points between multiple frames, and the spatial mapping under the front view and the side view are identified. The skeleton key point detection results in the front view and the side view are respectively graph constructed, and the skeleton key point detection results in the front view and the side view are regarded as a skeleton point graph topology The node is E represents the edge between the corresponding skeleton points The weight of the edge is defined according to the Euclidean distance between the skeleton points;
[0051] When the skeleton key point detection results in the front view and the side view are respectively graph constructed, the pixel coordinates of the skeleton points in the front view and the side view, the differential values of the pixel coordinates of the skeleton points between multiple frames, and the spatial mapping under the two views are taken as input ends, and the mathematical representation is as follows:
[0052]
[0053] Wherein, R represents a multidimensional array in a real number field, T is the length of a time sequence, 2 represents two different views, P i (t) represents the number of the skeleton point under the t-th frame, C i (t) represents the coordinates of the corresponding skeleton point, ΔX represents the difference between the coordinates of adjacent frames, and M is a weight matrix representing the spatial mapping relationship between the two views;
[0054] The skeleton point graph topology of each view is represented as:
[0055]
[0056] wherein A i,j represents the connection weight between the skeleton point i and the skeleton point j;
[0057] The graph convolution network GCN is used to perform convolution operation on the skeleton point graph topology of each view, and the differential value ΔX is integrated into the graph convolution operation to enhance the sensitivity of the model to action changes and extract local structure features. The graph convolution operation is represented as:
[0058]
[0059] wherein, is the node feature matrix of the lth layer, D (l) represents the dimension of the node feature vector of the lth layer of the graph convolution network, the initial dimension is designed as 64, and is adjusted according to the performance in the model training process, is the degree matrix, represents the self-connection of the node, W (l) is the weight matrix of the lth layer, and σ is the nonlinear activation function ReLU. After L-layer graph convolution, the output features of each view are:
[0060]
[0061] The feature sequence after graph convolution is input into the recurrent neural network RNN. The recurrent neural network RNN captures the time sequence features and performs average pooling on the features of each time step to obtain the feature representation of each time step. The outputs of the recurrent neural networks RNN of the two views are fused through a fusion layer to obtain the final feature vector.
[0062] wherein, for each view, the input of the recurrent neural network RNN is the feature sequence after graph convolution:
[0063]
[0064] The features of each time step are averaged and pooled to obtain the feature representation of each time step:
[0065]
[0066] The output of the RNN is represented as:
[0067]
[0068] The RNN outputs of the two perspectives are fused through a fusion layer to obtain the final feature vector, which is then fused using a weighted summation method:
[0069]
[0070] Where M is the view weight matrix, represents matrix multiplication, and finally, the dimension of the fused feature vector is:
[0071] H fusion ∈R T×H
[0072] The feature vector is input into two fully connected layers and the output layer, and the score of each parameter is obtained through the Sigmoid activation function. The scores of each parameter are weighted and summed to output the health assessment score.
[0073] On the other hand, the present invention also provides a multi-perspective children's motor coordination ability assessment system based on deep learning, the multi-perspective children's motor coordination ability assessment system based on deep learning includes:
[0074] The model building module is used to construct a dedicated children's motion dataset, capture a motion evaluation dataset based on the children's motion dataset, integrate the children's motion dataset and the motion evaluation dataset into a modeling dataset, use the modeling dataset to train a skeleton point detection model and a motion ability evaluation model, and output a converged skeleton point detection model and a motion ability evaluation model;
[0075] The data acquisition module collects real-time video streams of people entering the venue based on the acquisition camera, preprocesses the real-time video streams, and outputs preprocessed data sets;
[0076] The visual recognition module is used to load the preprocessed data set. The skeleton point detection model is based on the combination of deep convolutional neural network and regression method to identify and analyze the preprocessed data set, identify the skeleton key points in the preprocessed data set, and obtain the skeleton key point detection results;
[0077] The motion state detection module obtains the skeleton key point detection results, performs multi-view data association on the skeleton key point detection results, associates the person ID with the corresponding skeleton key points, obtains the person matching results containing the person ID and skeleton key points, and performs multi-person detection on the person matching results based on the motion area division method to obtain multi-person detection results;
[0078] The personalized assessment module takes the test results of multiple people as input and executes the sports ability assessment model. The sports ability assessment model quantitatively analyzes the test results of multiple people and generates a personalized health assessment report.
[0079] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0080] The application realizes automatic evaluation of children's dynamic actions by combining visual recognition and motion state detection technology, and using a front view angle and a side view angle dual-camera layout in cooperation with a deep learning algorithm. Compared with manual measurement or single view angle analysis, the application can completely capture the action process, ensure the consistency of evaluation standards, eliminate human errors, improve detection accuracy and efficiency, and support simultaneous detection of multiple people. The application uses the field-of-view partitioning technology of the acquisition camera to assign a unique identity number to each child entering the site and perform independent motion trajectory tracking, ensuring that the action data of different individuals are not confused, and realizing efficient multi-target action analysis.
[0081] In the embodiment of the application, the collected video stream is analyzed frame by frame and numbered and stored in chronological order, so that each frame of image has a clear timestamp and sequence identification. This not only facilitates maintaining the time sequence consistency of data during subsequent processing, but also enables quick positioning and backtracking of the motion state at a specific moment when needed, providing convenience for motion trajectory analysis and anomaly detection. The preprocessed data set has uniform format, resolution and quality standards and can be directly input into the skeleton point detection model and other analysis modules. This avoids the need for additional conversion or preprocessing work due to inconsistent data formats or uneven quality in the subsequent processing stage, significantly improving the running efficiency and real-time performance of the entire evaluation system.
[0082] In the embodiment of the application, the skeleton point detection model extracts key features of the preprocessed data set through convolution layers, effectively identifies important information related to skeleton points in the image, provides accurate basis for subsequent skeleton point positioning, reduces interference from irrelevant features, makes skeleton point recognition more accurate, uses clustering algorithms combined with regression analysis to classify and accurately classify the initially positioned skeleton points according to human body structure, fully considers the anatomical features and motion rules of human body skeleton, can effectively correct possible deviations in initial positioning, further improves the accuracy of skeleton key point recognition, and makes the recognition results of each skeleton point more consistent with the actual human body posture. A skeleton point detection model trained specifically for children's data is used instead of a general action recognition model, which has a significant advantage in the accuracy of children's action recognition. The model can more accurately capture the unique motion characteristics of children, improve the reliability of evaluation, and ensure strong evaluation pertinence and wide adaptability.
[0083] In the embodiment of the present application, by correlating the key point detection results of the skeleton with multi-view data, aligning the image data collected by the dual camera through the unified host timestamp, the consistency of the two video streams in the time dimension is ensured, the matching error caused by the time asynchronization is avoided, and the foundation for accurately correlating the skeleton point data is laid, which is crucial for analyzing the action patterns of children changing over time. Based on the strict mapping relationship between the pre-calibrated physical space of the site and the camera screen, the spatial position features are matched, the skeleton point data of the same person under different angles can be accurately matched, the mismatch caused by the angle difference or scene complexity is effectively reduced, the accuracy of data correlation is improved, and the correct evaluation of the motion state of children is ensured. Integrating the skeleton point data of the front and side views can comprehensively and stereoscopically reflect the motion of children in the three-dimensional space, provide more abundant and complete information for evaluating key indicators such as coordination, continuity and symmetry of actions, and help to find potential motion problems such as action asymmetry and abnormal posture, and provide more sufficient basis for subsequent development of personalized motion training programs. BRIEF DESCRIPTION OF DRAWINGS
[0084] Figure 1 is the implementation flowchart of the multi-view child motion coordination ability evaluation method based on deep learning provided by the present application.
[0085] Figure 2 shows the detection site division diagram when the key point detection results of the skeleton are correlated with multi-view data.
[0086] Figure 3 is the structural diagram of the multi-view child motion coordination ability evaluation system based on deep learning provided by the present application.
[0087] Figure 4 shows the method flowchart of the coordination evaluation of the multi-person detection results based on the motion ability evaluation model. DETAILED DESCRIPTION
[0088] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used in the specification of the application is only for the purpose of describing specific embodiments and is not intended to limit the application; the specification and claims of the application and the above description of the drawings, the terms "include" and "have" and any variations thereof, are intended to cover non-exclusive inclusion. The terms "first", "second" and the like in the specification and claims of the application or the above description of the drawings are used to distinguish different objects, not to describe a specific order.
[0089] Reference to“an embodiment” herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase“in an embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily all referring to a particular embodiment logically divided into parts. It is explicitly contemplated that embodiments described herein can be combined to include claims directed to combinations of the embodiments.
[0090] The prior art mainly relies on manual observation and simple physical tests when evaluating children's movements. Due to the limited experience and attention of the evaluators, it is difficult to synchronously track multiple angles and key details of children's actions, making it impossible for existing methods to automatically, accurately and efficiently evaluate children's dynamic actions, especially when complex movements are involved. Traditional methods cannot provide comprehensive, efficient and accurate standard unified evaluation. To address the above problems, we propose a multi-view child movement coordination ability evaluation system and method based on deep learning. The present application identifies and analyzes the preprocessed data set based on the combination of deep convolutional neural network and regression method, and performs multi-person detection on the personnel matching results based on the motion region division method. It combines visual recognition and motion state detection technology, and uses a front view and a side view dual-camera layout, combined with a deep learning algorithm, to realize automatic evaluation of children's dynamic actions. Compared with manual measurement or single-view analysis, it can completely capture the action process, ensure the consistency of the evaluation standard, eliminate human errors, improve the detection accuracy and efficiency, and support multi-person simultaneous detection. The field-of-view partitioning technology of the acquisition camera is used to assign a unique identity number to each child entering the field and track their independent motion trajectories, ensuring that the action data of different individuals are not confused, and achieving efficient multi-target action analysis.
[0091] The embodiment of the present application provides a multi-view child movement coordination ability evaluation method based on deep learning, Figure 1 An implementation process schematic diagram of the multi-view child movement coordination ability evaluation method based on deep learning is shown, and the multi-view child movement coordination ability evaluation method based on deep learning specifically comprises:
[0092] Step S10, constructing a special child movement data set, grabbing a movement evaluation data set based on the child movement data set, integrating the child movement data set and the movement evaluation data set into a modeling data set, training the modeling data set to obtain a skeleton point detection model and a movement ability evaluation model, and outputting the converged skeleton point detection model and the movement ability evaluation model;
[0093] Step S20, acquiring real-time video streams of personnel entering the field based on an acquisition camera, preprocessing the real-time video streams, and outputting a preprocessed data set;
[0094] Step S30, load the preprocessed data set, the skeleton point detection model is based on the deep convolutional neural network and the regression method is combined to identify and analyze the preprocessed data set, identify the skeleton key points in the preprocessed data set, and obtain the skeleton key point detection result;
[0095] Step S40, obtain the skeleton key point detection result, perform multi-view data correlation on the skeleton key point detection result, associate the person ID with the corresponding skeleton key point, obtain the personnel matching result containing the person ID and the skeleton key point, perform multi-person detection on the personnel matching result based on the motion region division method, and obtain the multi-person detection result;
[0096] Step S50, taking the multi-person detection result as input, executing the motion ability evaluation model, the motion ability evaluation model quantitatively analyzes the multi-person detection result, and generates a personalized health evaluation report.
[0097] The application combines visual recognition and motion state detection technology, and adopts a front view and a side view dual-camera layout, cooperates with a deep learning algorithm, and realizes automatic dynamic action evaluation of children. Compared with manual measurement or single-view analysis, the action process can be completely captured, the consistency of the evaluation standard is ensured, human errors are eliminated, the detection accuracy and efficiency are improved, and multiple people can be detected at the same time. The field-of-view partition technology of the acquisition camera is used to assign a unique identity number to each child entering the site and perform independent motion trajectory tracking, so that the action data of different individuals cannot be confused, and efficient multi-target action analysis is realized.
[0098] The embodiment of the application provides a method for constructing a child motion data set, and the method specifically comprises the following steps:
[0099] Step S101, a multi-angle and full-body entry shooting mode is used to collect special video data of children in different scenes, wherein the video data covers common motion actions such as running, jumping and side sliding, and also covers natural actions of children in play or daily activities, these images provide diversified motion data sources, which are helpful to improve the recognition ability of the model under different motion conditions;
[0100] Step S102, load the special video data, analyze the special video data frame by frame to obtain at least one set of special images, mark the child skeleton points in the special images to obtain a child motion data set containing skeleton point marking results, wherein when marking the child skeleton points in the special images, 17 skeleton points of the child are set with fixed numbers, which are respectively: 1. nose, 2. left eye, 3. right eye, 4. left ear, 5. right ear, 6. left shoulder, 7. right shoulder, 8. left elbow, 9. right elbow, 10. left wrist, 11. right wrist, 12. left hip joint, 13. right hip joint, 14. left knee, 15. right knee, 16. left ankle, 17. right ankle;
[0101] It should be noted that when the special video data is analyzed frame by frame to obtain at least one set of special images, the special images contain image data of various motion modes of the child, and all the skeleton points are manually labeled by experts to ensure that the model can accurately identify the key skeleton point positions in the child motion.
[0102] Step S103, based on the child motion data set, the motion evaluation data set is grabbed, wherein the motion evaluation data set is based on the abnormal action data collection of the child with motion coordination deviation, and the child motion data set and the motion evaluation data set are integrated into the modeling data set.
[0103] In this embodiment, based on the child motion data set, the action data of the normally developed children and the children with motion development difference is further collected, and a motion evaluation data set specially used for evaluating the coordination ability of the child is constructed. For example, in the side sliding step test, the children with poor coordination may have insufficient development of vestibular sense or weak control ability of core muscle group, resulting in asymmetric rotation of shoulder and pelvis, and it is difficult to keep the upper body stable. Some children may have trunk lateral flexion compensation action, that is, the body is inclined to one side unconsciously during the sliding process to maintain balance. In addition, the children with abnormal gait coordination may have inconsistent direction of shoulder and foot movement, or obvious pause during the sliding process, and cannot complete the action smoothly. However, the common model has low recognition accuracy for such data, and it is difficult to support accurate quantitative evaluation.
[0104] To improve the recognition ability of the skeletal point detection model and the movement ability evaluation model for the above abnormal situations, abnormal action data of children with movement coordination deviation is specially collected. For example, some children have different left and right body coordination abilities due to sensory integration disorder, which is manifested as one shoulder being significantly higher than the other during sliding, and the angle between the shoulder skeletal points and the ground is significantly deviated from the horizontal reference (such as more than 5°). For another example, children with insufficient core stability have difficulty maintaining a stable center of gravity during sliding, resulting in a significant deviation of the sliding trajectory, and the step size difference shown by the skeletal points is too large. The collection and analysis of these abnormal data help the skeletal point detection model and the movement ability evaluation model accurately identify the deviation of the children's movement pattern, and provide more reliable basis for subsequent evaluation.
[0105] The embodiment of the present application provides a method for training skeletal point detection model and movement ability evaluation model by using modeling data set, and the method specifically comprises the following steps:
[0106] Step S201, load the modeling data set, and divide the modeling data set into a training set and a validation set, wherein the training set is used to train the skeletal point detection model and the movement ability evaluation model, and the validation set is used to verify the generalization ability of the skeletal point detection model and the movement ability evaluation model, and the proportion of the training set and the validation set is 4:1;
[0107] Step S202, iteratively train the skeletal point detection model by using the training set and the validation set, and output the converged skeletal point detection model, wherein the training set includes skeletal point labeling data and movement data of normal and abnormal children, and the training set is used to train the skeletal point detection model and the movement ability evaluation model, so that the skeletal point detection model can identify the joint positions of children in different movement actions and distinguish normal and abnormal movement patterns;
[0108] Step S203, load the skeletal point detection model, and the skeletal point detection model combines the data of the front view angle and the side view angle to stably identify the skeletal points of the children under different view angles, and outputs the skeletal point recognition result of the skeletal point detection model;
[0109] Step S204, obtain the skeletal point recognition result, and optimize the movement ability evaluation model based on the skeletal point recognition result, and output the movement ability evaluation model based on the multi-view sequence skeletal point image data.
[0110] In the embodiment, by combining the data of the front view angle and the side view angle, the skeleton point detection model can stably identify the skeleton points of the child at different view angles. Meanwhile, by the optimization algorithm for the child posture feature, the recognition deviation caused by the height and body size difference is further reduced. After the skeleton point recognition model is trained, the skeleton point recognition result output by the skeleton point detection model is input into the movement ability evaluation model of the second stage, the movement state of the child is evaluated through the analysis of the movement data, and after such training and optimization, the recognition accuracy of the movement ability evaluation model for the child posture feature is higher than that of the ordinary model.
[0111] The embodiment of the present application provides a method for pre-processing a real-time video stream, and the method specifically comprises the following steps:
[0112] In step S301, a camera is used to collect a video stream in real time at a fixed frame rate of 60 FPS to obtain a real-time video stream.
[0113] It should be noted that the camera can be two groups, which are arranged on the front and side respectively, and the camera collects the video stream in real time at a fixed frame rate of 60 FPS to ensure complete recording of the whole process of the child's movement.
[0114] In step S302, the real-time video stream is loaded, and the real-time video stream is processed frame by frame, and the real-time images processed frame by frame are numbered and stored in time sequence.
[0115] In step S303, the real-time images processed in time sequence are subjected to format conversion, grayscale normalization processing is performed to reduce the influence of light changes, and a smoothing filter method is used to remove noise in the images to reduce slight errors caused by sensor noise or environmental interference, and a pre-processed data set is obtained. In addition, all image frames are standardized to a uniform resolution.
[0116] In the embodiment of the present application, the collected video stream is processed frame by frame and numbered and stored in time sequence, so that each image has a clear timestamp and sequence identification. This not only facilitates maintaining the time sequence consistency of the data in subsequent processing, but also enables quick positioning and backtracking of the movement state at a specific time when needed, providing convenience for movement trajectory analysis and anomaly detection. The pre-processed data set has a uniform format, resolution and quality standard, and can be directly input into the skeleton point detection model and other analysis modules. This avoids the need for additional conversion or preprocessing work due to inconsistent data formats or uneven quality in the subsequent processing stage, significantly improving the running efficiency and real-time performance of the entire evaluation system.
[0117] The embodiment of the present application provides a method for identifying skeleton key points in a pre-processed data set, and the method specifically comprises the following steps:
[0118] Step S401, load the preprocessed data set, the skeleton point detection model extracts the key features of the preprocessed data set through the convolution layer, identifies a plurality of skeleton point candidate positions in the preprocessed data set, and preliminarily locates the skeleton points;
[0119] Step S402, obtain the preliminary positioning result of the skeleton points, classify and accurately classify the preliminary positioning result of the skeleton points according to the human body structure through a clustering algorithm combined with regression analysis, and obtain the skeleton key point detection result;
[0120] Wherein, when the preliminary positioning result of the skeleton points is classified and accurately classified according to the human body structure, each class is regarded as a target, and the skeleton point number set of each target is output first:
[0121] P i (t)={p1,p2,...,p n}
[0122] Where p k represents the skeleton point number: 1. nose, 2. left eye, 3. right eye, 4. left ear, 5. right ear, 6. left shoulder, 7. right shoulder, 8. left elbow, 9. right elbow, 10. left wrist, 11. right wrist, 12. left hip joint, 13. right hip joint, 14. left knee, 15. right knee, 16. left ankle, 17. right ankle;
[0123] Each number corresponds to a two-dimensional coordinate, indicating the position of the skeleton point in the image:
[0124] C i (t)={(x1,y1),(x2,y2),...,(x n ,y n )}
[0125] Where (x k ,y k ) is the coordinate of the skeleton point p k on the image.
[0126] In the embodiment, the skeleton point detection model extracts key features of the preprocessed data set through the convolution layer, can effectively identify important information related to the skeleton point in the image, provides accurate basis for subsequent skeleton point positioning, reduces the interference of irrelevant features, makes the skeleton point recognition more accurate, uses the clustering algorithm combined with regression analysis, classifies and accurately classifies the preliminarily positioned skeleton points according to the human body structure, fully considers the anatomical characteristics and motion law of human skeleton, can effectively correct the deviation that may occur in preliminary positioning, further improves the accuracy of skeleton key point recognition, makes the recognition result of each skeleton point more consistent with the actual human body posture, adopts the skeleton point detection model specially trained for children data, rather than the general action recognition model, has a significant advantage in the accuracy of children action recognition. The model can more accurately capture the motion characteristics specific to children, improve the reliability of evaluation, and ensure that the evaluation is highly targeted and widely adaptable.
[0127] The embodiment of the present application provides a method for multi-view data association of skeleton key point detection results, and the method for multi-view data association of skeleton key point detection results specifically comprises:
[0128] In step S501, load the skeleton key point detection result, identify the image-related acquisition camera in the skeleton key point detection result, and synchronously acquire video streams by two acquisition cameras, each frame of image has a time stamp t, and output image data. In order to ensure that the images captured by the two acquisition cameras are consistent in time sequence, the system ensures that the time difference of the pull stream between the cameras is less than 100 ms, takes the host time stamp as a unified standard, aligns the time stamps of the acquisition videos of the acquisition cameras, and outputs image data with synchronized time stamps: I front (t),I side (t);
[0129] Step S502, since the system adopts two acquisition cameras (front view angle, side view angle), each individual entering the experimental site will be captured at two view angles at the same time, in multi-view data association, the system adopts a double-view target matching method based on physical space position, pre-calibrates the strict mapping relationship between the actual physical space of the site and the acquisition camera screen, and performs multi-view data association on the skeleton key point detection results according to the pre-calibrated mapping relationship to obtain personnel matching results containing person ID and skeleton key points. The mapping relationship is established through accurate acquisition camera installation position, measurement and view angle calibration experiment. When the target is located in the upper left corner of the physical area of the site, according to the calibrated mapping relationship, the system can determine that the target must appear in the right side area in the front view angle screen, and must appear in the left side area in the left side camera screen. When the target is located in the upper left corner of the physical area of the site, according to the pre-calibrated mapping relationship, matching is realized through the following specific steps: first, real-time detection of all target skeleton point data appearing in the right side area of the front view angle screen, and detection of target skeleton point data in the left side area of the side view angle screen, then spatial consistency verification, checking whether the relative position relationship of the two sets of skeleton points conforms to the pre-calibrated spatial mapping rule, then performing motion state verification, verifying whether the motion trends reflected by the two sets of skeleton points are consistent, and when all the verifications pass, assigning the same ID number to the two sets of skeleton point data, that is:
[0130]
[0131] Step S503, combining the ID number with the skeleton point data of each person to represent the corresponding relationship between the number and the skeleton point data:
[0132]
[0133] Step S504, multi-person detection is performed on the personnel matching results based on the motion region division method to obtain multi-person detection results.
[0134] In this embodiment, in order to realize multi-person simultaneous detection, the motion region division method is adopted to reasonably divide the detection site, so as to ensure that the motion of each individual can be completely captured and avoid occlusion interference in the front view angle and side view angle images. The detection site is divided into a plurality of independent rectangular regions as shown in Figure 2 The regions are arranged along the diagonal lines, similar to the diagonal line distribution of a chessboard, and each rectangular region does not overlap, so as to ensure that only one individual moves in each region at the same time. Such design can ensure that the camera screens of the front view angle and the side view angle do not overlap between individuals, thereby improving the detection accuracy.
[0135] In the multi-person detection process, the system attributes the individual to the preset motion area according to the initial position of the individual in the camera picture, and ensures that each area corresponds to only one individual. The key to this is to avoid cross occlusion between individuals, so that the skeletal point data of each person can be captured independently, avoiding data confusion caused by character overlap or mis-matching in the detection process. Since the individual in each area always remains independent, the system can analyze the motion data in different areas in parallel when processing, without the need for an additional individual separation step, thereby reducing the amount of calculation and improving the processing speed. This method effectively reduces the interference when multiple people are detected at the same time, ensures the stability and accuracy of skeletal point matching, optimizes the processing efficiency of data flow, and improves the detection accuracy and real-time performance of the overall evaluation scheme. After the above multiple processes, real-time and accurate skeletal point detection of multiple people from multiple angles can be realized, and the motion coordination evaluation model can be input for motion coordination evaluation.
[0136] In the embodiment of the application, by correlating the skeletal key point detection results with multi-view data, aligning the image data collected by the dual camera through a unified host timestamp, the consistency of the two video streams in the time dimension is ensured, and matching errors caused by different time synchronization are avoided, which lays a foundation for subsequent accurate correlation of skeletal point data. This is crucial for analyzing the action patterns of children over time. Based on the strict mapping relationship between the pre-calibrated physical space of the site and the camera picture, the spatial position features are matched, the skeletal point data of the same person under different viewing angles can be accurately matched, the mis-matching caused by viewing angle difference or scene complexity is effectively reduced, the accuracy of data correlation is improved, and the correct evaluation of the motion state of children is ensured. Integrating the skeletal point data of the front view and the side view can comprehensively and stereoscopically reflect the motion of children in three-dimensional space, provide more abundant and complete information for evaluating key indicators such as coordination, continuity and symmetry of actions, and help to find potential motion problems such as action asymmetry and abnormal posture, and provide a more sufficient basis for subsequent development of individualized motion training programs.
[0137] The embodiment of the application also provides a method for quantitative analysis of multi-person detection results by a motion ability evaluation model, and the method specifically comprises:
[0138] Step S601: loading the multi-person detection results, evaluating the coordination of the multi-person detection results based on the motion ability evaluation model, and obtaining a child health evaluation score corresponding to a health standard;
[0139] It should be noted that the motion ability evaluation model in the embodiment includes a feature extraction module and a regression scoring module, wherein the feature extraction module includes a graph construction layer, a graph convolution network GCN, a recurrent neural network RNN, an average pooling layer and a feature fusion layer connected in sequence, and the regression scoring module is used for finally outputting a health evaluation score, and the regression scoring module includes a full connection layer 1, a full connection layer 2 and an output layer connected in sequence, wherein the full connection layer 1 is used for acquiring a feature vector output by the feature fusion layer, the full connection layer 1 and the full connection layer 2 pass data through a ReLu activation function, the full connection layer 2 and the output layer also pass data through a ReLu activation function, and the output layer outputs a score vector through a Sigmoid activation function.
[0140] In the embodiment of the application, a method for coordinative evaluation of multi-person detection results based on a motion ability evaluation model is also provided. Figure 4 A flowchart of a method for coordinative evaluation of multi-person detection results based on a motion ability evaluation model is shown, and the method for coordinative evaluation of multi-person detection results based on a motion ability evaluation model includes:
[0141] In step S6011, multi-person detection results are acquired, and pixel coordinates of skeletal points in a front view and a side view, difference values of pixel coordinates of skeletal points between multiple frames and spatial mapping in the front view and the side view in the multi-person detection results are identified, a graph is constructed for skeletal key point detection results in the front view and the side view, and the skeletal key point detection results in the front view and the side view are regarded as a skeletal point graph topology. The node is represented as V E represents an edge between corresponding skeletal points The weight of the edge is defined according to the Euclidean distance between the skeletal points.
[0142] In step S6012, a graph convolution network GCN is used to perform convolution operation on the skeletal point graph topology of each view, and the difference value ΔX is integrated into the graph convolution operation to enhance the sensitivity of the model to action changes, and local structural features are extracted.
[0143] In step S6013, the feature sequence after graph convolution is input into a recurrent neural network RNN, the recurrent neural network RNN captures time sequence features, and performs average pooling on the features of each time step to obtain feature representation of each time step, and the outputs of the recurrent neural network RNN of the two views are fused through a fusion layer to obtain a final feature vector.
[0144] Specifically, when the skeletal key point detection results in the front view and the side view are graph constructed, the pixel coordinates of the skeletal points in the front view and the side view, the difference values of the pixel coordinates of the skeletal points between multiple frames and the spatial mapping in the two views are taken as input ends, and the mathematical representation is as follows:
[0145]
[0146] where R represents a multidimensional array over a real field, T is the length of time series, 2 represents two different views, P i (t) represents the bone point number in the t-th frame, C i (t) represents the coordinates of the corresponding bone point, ΔX represents the difference between the coordinates of adjacent frames, M is a weight matrix representing the spatial mapping relationship between the two views;
[0147] For each view (front view, side view), the bone point sequence is regarded as a bone point graph topology where the node E represents the edge between the corresponding bone points The weight of the edge is defined according to the Euclidean distance between the bone points;
[0148] The bone point graph topology of each view can be represented as:
[0149]
[0150] where A i,j represents the connection weight between bone point i and bone point j. The graph convolution network (GCN) is used to perform convolution operation on the bone point graph of each view, and the difference value ΔX is integrated into the graph convolution operation to enhance the sensitivity of the model to action changes and extract local structural features. The graph convolution operation can be represented as:
[0151]
[0152] where, is the node feature matrix of the l-th layer, D (l) represents the dimension of the node feature vector of the l-th layer of the graph convolution network, the initial dimension is designed as 64, which can be adjusted according to the performance in the model training process, is the degree matrix, represents the self-connection of the node, W (l) is the weight matrix of the l-th layer, σ is the nonlinear activation function ReLU, and after L-layer graph convolution, the output feature of each view is:
[0153]
[0154] The feature sequence after graph convolution is input into the recurrent neural network (RNN) to capture the time series features. For each view, the input of the recurrent neural network RNN is the feature sequence after graph convolution:
[0155]
[0156] For simplicity, we can average pool the features at each time step to get a feature representation for each time step:
[0157]
[0158] The output of the RNN is represented as:
[0159]
[0160] The outputs of the RNNs from the two views are fused through a fusion layer to get the final feature vector. The fusion is done using weighted summation:
[0161]
[0162] where M is the view weight matrix and · denotes matrix multiplication. Finally, the dimension of the fused feature vector is:
[0163] H fusion ∈R T×H
[0164] In step S6014, the feature vector is input to two dense layers and an output layer to get the score of each parameter through the Sigmoid activation function, and the scores of each parameter are weighted and summed to output the health assessment score.
[0165] In the regression score stage, the feature vector is input to two dense layers and an output layer to calculate the score of each parameter of the health standard. Assuming that the health standard has K parameters, the score of each parameter ranges from 0 to 100. The mathematical representation of the dense layers is as follows:
[0166] Z1 = ReLU(W1H fusion +b1),
[0167] Z2 = ReLU(W2Z1 + b2),
[0168] where W1, W2 and are weight matrices, and are bias vectors, and are output values. Then the score of each parameter is obtained through the Sigmoid activation function, which is mathematically represented as follows:
[0169] S = σ(W3Z2 + b3),
[0170] where W3 is a weight matrix, b1 K is a bias vector, and S T×Kis the output rating vector, and σ is the activation function Sigmoid. In order to obtain a total rating value, we perform a weighted summation of the ratings of each parameter:
[0171]
[0172] Among them, w k is the weight of the kth parameter, S k is the rating of the kth parameter. Subsequently, the mean squared error (MSE) is used as the loss function for the regression task to calculate the difference between the predicted rating and the true rating:
[0173]
[0174] Among them, S k is the predicted score, is the actual score. Using the Adam optimizer, the learning rate is adjusted according to the actual situation:
[0175]
[0176] Among them, θ t are the model parameters for the tth iteration, and α is the learning rate.
[0177] It should be noted that the motor ability assessment model determines whether the child's motor ability meets health standards by analyzing key indicators such as the child's movement trajectory, stability, and range of motion, and promptly detects possible motor control problems or neurodevelopmental delays, especially subtle developmental abnormalities that are difficult to detect through routine observation.
[0178] In an embodiment of the present invention, the training process of the athletic ability assessment model involves feature extraction and scoring stages. First, the image data is converted into the pixel coordinates of the skeleton points, the differential values between multiple frames, and the spatial mapping relationship between the perspectives. A skeleton point map is constructed for each perspective, and GCN is applied to extract local structural features. Subsequently, the time series features are captured by RNN, and then a feature vector is formed by feature fusion. In the regression scoring stage, the feature vector is input to the fully connected layer, using MSE as the loss function, and the parameters are updated by the Adam optimizer to calculate the parameter scores of the health standard.
[0179] For example, in the side step test, the motor ability assessment model first analyzes the child's movement trajectory and quantitatively assesses the key movement postures. The test requires the child to step in the direction of the slide with the body side, and to be evaluated under front and side views. Among them, the standard of "shoulder parallel to the ground" is calculated based on the bone point data of the front view. The motor ability assessment model extracts the coordinate values of the left shoulder (bone point number 6) and the right shoulder (bone point number 7), and calculates the included angle of them relative to the horizontal axis. If the included angle is less than 5°, it is determined that the shoulder is basically kept horizontal, otherwise it is considered as posture deviation. In addition, the motor ability assessment model side view can further analyze the consistency of the foot movement direction and the shoulder direction to determine whether the child maintains a normal posture during lateral movement during the sliding process. This evaluation method based on the spatial relationship of bone points converts subjective action standards into measurable geometric parameters, ensuring the objectivity and repeatability of the evaluation results.
[0180] The motor ability assessment model also detects whether there is a moment of taking off the ground. Based on bone point data, the system tracks the vertical coordinate changes of both ankles (bone point numbers 16 and 17), and if both coordinates are simultaneously higher than a fixed threshold in the previous frame during the sliding process, and there is no bone point below the reference horizontal line of the ground contact point, it is determined that the action meets the requirements; if at least one foot (16 or 17) is always in contact with the ground, it is considered as not completed. In addition, when the motor ability assessment model detects that the child completes the double-foot landing-double-foot take-off action, it is considered that the child has completed a cycle of action. The motor ability assessment model counts whether the child has the ability to complete the corresponding task according to the evaluation requirements (for example: complete a complete 4 cycles, and complete the sliding in the opposite direction).
[0181] Through these detailed evaluation standards, the motor ability assessment model not only accurately judges whether the child has completed the action, but also analyzes the detailed performance of the child's motor ability, and provides targeted improvement suggestions. This process ensures the accuracy and comprehensiveness of the evaluation, and provides scientific motor coordination ability evaluation basis for parents and professionals.
[0182] Step S602, the motor ability assessment model analyzes the health assessment score based on the health standard, judges whether the health assessment score is qualified, and obtains the health assessment result;
[0183] It should be noted that the health assessment score is calculated based on the child's motor ability data and the set health standard, and the calculation formula considers the weight parameters of each index feature in each movement, comprehensively evaluates the key indexes such as motor stability, coordination, and movement amplitude, and judges whether the movement is completed. After the movement is judged to be completed, the score is added.
[0184] In the evaluation process, each indicator in each test item is independently determined as "qualified" or "unqualified". If a certain indicator is not completed, it cannot be scored and will affect the final evaluation result.
[0185] Qualified (S≥70): The movement ability meets the health standard, and each evaluation indicator meets the requirements. The movement coordination and stability are normal or only slightly deviated.
[0186] Unqualified (S<70): The movement ability does not meet the standard, and there are obvious movement abnormalities or coordination problems. It is recommended to conduct targeted training and intervention.
[0187] In step S603, the movement ability evaluation model generates a personalized health evaluation report combined with the health evaluation result.
[0188] The personalized health evaluation report includes: evaluation indicator analysis, health score and suggestion, and training suggestion.
[0189] It should be noted that the evaluation indicator analysis: gives the completion of each movement, analyzes each movement indicator in detail and gives the corresponding conclusion, helping parents and educators identify potential problems in movement.
[0190] Health score and suggestion: according to the evaluation score, provide the evaluation of the current movement ability level, combined with specific data to provide personalized training suggestions and improvement direction, help children optimize movement performance and promote healthy development, for example, in the side sliding step test, the system detects that the child's shoulder cannot be kept horizontal during sliding, the calculation result shows that the angle between the left shoulder (skeletal point number 6) and the right shoulder (skeletal point number 7) is 8°, which exceeds the set threshold of 5°. At the same time, the angle between the step direction and the shoulder direction is 12°, which is higher than the coordination standard of 10°, indicating that the child has compensatory movement of body twisting during sliding. When the child's movement is severely uncoordinated, it indicates that the child may have neurological development disorders.
[0191] According to the evaluation result of the child, the movement ability evaluation model will give individualized improvement suggestions in the evaluation report:
[0192] Training suggestions:
[0193] If the child feels integrated training, as well as cognitive ability training in daily life, neural motor task training, etc. Secondly, psychological treatment can help improve, such as parents giving correct guidance to improve the child's self-awareness.
[0194] In the embodiment of the application, the movement ability evaluation model can monitor the dynamic movement of children in real time (such as jumping, single-leg jumping, side sliding step, etc.), and provide accurate evaluation according to complete space-time feature data to generate personalized movement ability report and improvement suggestions.
[0195] In another aspect, the present application also provides a multi-view child motion coordination ability evaluation system based on deep learning, Figure 3 A multi-view child motion coordination ability evaluation system based on deep learning is shown, which comprises:
[0196] A model construction module 100 is configured to construct a special child motion dataset, to capture a motion evaluation dataset based on the child motion dataset, to integrate the child motion dataset and the motion evaluation dataset into a modeling dataset, to train a skeletal point detection model and a motion ability evaluation model by using the modeling dataset, and to output the converged skeletal point detection model and the motion ability evaluation model.
[0197] A data acquisition module 200 is configured to acquire real-time video streams of personnel entering a site based on an acquisition camera, to pre-process the real-time video streams, and to output a pre-processing dataset.
[0198] A visual recognition module 300 is configured to load the pre-processing dataset, to recognize and analyze the pre-processing dataset based on a combination of a deep convolutional neural network and a regression method, to identify skeletal key points in the pre-processing dataset, and to obtain a skeletal key point detection result.
[0199] A motion state detection module 400 is configured to obtain the skeletal key point detection result, to perform multi-view data correlation on the skeletal key point detection result, to associate a person ID with corresponding skeletal key points, to obtain a personnel matching result containing the person ID and the skeletal key points, to perform multi-person detection on the personnel matching result based on a motion region division method, and to obtain a multi-person detection result.
[0200] A personalized evaluation module 500 is configured to take the multi-person detection result as an input, to execute the motion ability evaluation model, to quantitatively analyze the multi-person detection result by using the motion ability evaluation model, and to generate a personalized health evaluation report.
[0201] It should be noted that the model construction module 100, the data acquisition module 200, the visual recognition module 300, the motion state detection module 400, and the personalized evaluation module 500 in the multi-view child motion coordination ability evaluation system based on deep learning correspond to the multi-view child motion coordination ability evaluation method based on deep learning, and the explanations, examples, beneficial effects, and other parts of the related content can refer to the corresponding content in the multi-view child motion coordination ability evaluation method based on deep learning, which will not be repeated here.
[0202] In summary, the application provides a multi-view child motion coordination ability evaluation system and method based on deep learning. The application identifies and analyzes the preprocessed data set based on the combination of deep convolutional neural network and regression method through a skeleton point detection model, and performs multi-person detection on the personnel matching result based on a motion region division method. The application combines visual recognition and motion state detection technology, and adopts a front view and side view dual-camera layout, cooperates with a deep learning algorithm, and realizes automatic child dynamic action evaluation. Compared with manual measurement or single view analysis, the application can completely capture the action process, ensure the consistency of evaluation standards, eliminate human errors, improve detection accuracy and efficiency, support multi-person simultaneous detection, use the field of view partition technology of the acquisition camera, assign a unique identity number to each child entering the site, and perform independent motion trajectory tracking, ensure that the action data of different individuals are not confused, and realize efficient multi-target action analysis.
[0203] It should be noted that for the foregoing embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the application is not limited by the described action sequence, because according to the application, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily necessary for the application.
[0204] The above embodiments are only used to illustrate the technical solutions of the present application, but not to limit the protection scope of the application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on these embodiments, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of the present application. Although the present application has been described in detail with reference to the above embodiments, those of ordinary skill in the art can still combine, add or delete the features of the embodiments of the present application according to the circumstances without creative labor, so as to obtain different other technical solutions which do not deviate from the concept of the present application in essence. These technical solutions also belong to the scope of the present application.
Claims
1. A multi-perspective children's motor coordination ability assessment method based on deep learning, characterized by: The method comprises: Construct a dedicated children's motion dataset, capture a motion evaluation dataset based on the children's motion dataset, integrate the children's motion dataset and the motion evaluation dataset into a modeling dataset, use the modeling dataset to train a skeleton point detection model and a motion ability evaluation model, and output a converged skeleton point detection model and a motion ability evaluation model; Based on the real-time video stream of people entering the venue collected by the acquisition camera, the real-time video stream is preprocessed and the preprocessed data set is output; The preprocessed dataset is loaded. The skeleton point detection model recognizes and analyzes the preprocessed dataset based on a combination of deep convolutional neural networks and regression methods, identifies the skeleton key points in the preprocessed dataset, and obtains the skeleton key point detection results. Obtain the skeleton key point detection results, perform multi-view data association on the skeleton key point detection results, associate the person ID with the corresponding skeleton key points, obtain the person matching results containing the person ID and skeleton key points, perform multi-person detection on the person matching results based on the motion area division method, and obtain the multi-person detection results; Taking the test results of multiple people as input, the sports ability assessment model is executed. The sports ability assessment model quantitatively analyzes the test results of multiple people and generates a personalized health assessment report.
2. The multi-perspective children's motor coordination ability assessment method based on deep learning as claimed in claim 1, characterized in that: The method for constructing a children's movement dataset comprises: Multi-angle, full-body shooting is used to collect specialized video data of children in different scenarios. The video data covers common movements such as running, jumping, and side-sliding, as well as children's natural movements during play or daily activities. Loading dedicated video data, parsing the dedicated video data frame by frame to obtain at least one set of dedicated images, marking the child's skeletal points in the dedicated images, and obtaining a child motion dataset containing the skeletal point marking results, wherein when marking the child's skeletal points in the dedicated images, fixed numbers are set for the child's 17 skeletal points, namely:
1. nose, 2. left eye, 3. right eye, 4. left ear, 5. right ear, 6. left shoulder, 7. right shoulder, 8. left elbow, 9. right elbow, 10. left wrist, 11. right wrist, 12. left hip joint, 13. right hip joint, 14. left knee, 15. right knee, 16. left ankle, 17. right ankle; The motion evaluation dataset is captured based on the children's motion dataset. The motion evaluation dataset is based on the abnormal movement data collection of children with motion coordination deviations, and the children's motion dataset and the motion evaluation dataset are integrated into the modeling dataset.
3. The multi-perspective children's motor coordination ability assessment method based on deep learning as claimed in claim 2, characterized in that: The method of using the modeling data set to train a skeleton point detection model and a motion ability assessment model includes: Load the modeling dataset and divide it into a training set and a validation set. The training set is used to train the skeleton point detection model and the motion ability assessment model, while the validation set is used to verify the generalization ability of the skeleton point detection model and the motion ability assessment model. The ratio of the training set to the validation set is 4:
1. The skeleton point detection model is iteratively trained using a training set and a validation set, outputting a converged skeleton point detection model. The training set includes skeleton point annotation data and movement data of normal and abnormal children. The training set is used to train the skeleton point detection model and the movement ability assessment model, enabling the skeleton point detection model to identify the joint positions of children in different movements and distinguish between normal and abnormal movement patterns. Load the skeleton point detection model. The skeleton point detection model combines data from the frontal and side views to stably identify the child's skeleton points under different viewing angles and outputs the skeleton point detection model's skeleton point recognition results. Obtaining the skeleton point recognition results, using the skeleton point recognition results to train and optimize the motion ability assessment model, and outputting the motion ability assessment model based on multi-view sequence skeleton point image data; Among them, the athletic ability assessment model includes a feature extraction module and a regression scoring module. The feature extraction module includes a graph construction layer, a graph convolutional network GCN, a recurrent neural network RNN, an average pooling layer, and a feature fusion layer connected in sequence. The regression scoring module is used to finally output the health assessment score. The regression scoring module includes a fully connected layer 1, a fully connected layer 2, and an output layer connected in sequence. Among them, the fully connected layer 1 is used to obtain the feature vector output by the feature fusion layer. Data is transmitted between the fully connected layer 1 and the fully connected layer 2 through the ReLu activation function. Data is also transmitted between the fully connected layer 2 and the output layer through the ReLu activation function. The output layer outputs the scoring vector through the Sigmoid activation function.
4. The multi-perspective children's motor coordination ability assessment method based on deep learning as claimed in claim 1, characterized in that: The method for preprocessing a real-time video stream comprises: The acquisition camera collects video streams in real time at a fixed frame rate of 60FPS to obtain real-time video streams; Loading real-time video stream, parsing and processing the real-time video stream frame by frame, and numbering and storing the real-time images processed frame by frame in chronological order; The real-time images processed according to the time sequence are format converted, grayscale normalization is performed, and the noise in the image is removed by smoothing filtering method to obtain the preprocessed data set.
5. The multi-perspective children's motor coordination ability assessment method based on deep learning as claimed in claim 4, characterized in that: The method for identifying skeleton key points in a preprocessed data set comprises: The preprocessed dataset is loaded. The skeleton point detection model extracts the key features of the preprocessed dataset through the convolution layer, identifies multiple skeleton point candidate positions in the preprocessed dataset, and preliminarily locates the skeleton points. Obtain the preliminary positioning results of the skeleton points, classify and accurately categorize the preliminary positioning results of the skeleton points according to the human body structure through clustering algorithm combined with regression analysis, and obtain the skeleton key point detection results; Among them, when the preliminary positioning results of the skeleton points are classified and accurately classified according to the human body structure, each class is regarded as a target, and the skeleton point number set of each target is first output: P i (t)={p1,p2,...,p n } where p k Skeletal point numbers:
1. Nose, 2. Left eye, 3. Right eye, 4. Left ear, 5. Right ear, 6. Left shoulder, 7. Right shoulder, 8. Left elbow, 9. Right elbow, 10. Left wrist, 11. Right wrist, 12. Left hip, 13. Right hip, 14. Left knee, 15. Right knee, 16. Left ankle, 17. Right ankle; Each number corresponds to a two-dimensional coordinate, indicating the position of the bone point in the image: C i (t)={(x1,y1),(x2,y2),...,(x n ,y n )} Where (x k ,y k ) is the bone point p k Coordinates on the image.
6. The multi-perspective children's motor coordination ability assessment method based on deep learning according to claim 5, characterized in that: The method for performing multi-view data association on skeleton key point detection results includes: Load the skeleton key point detection results, identify the acquisition camera associated with the image in the skeleton key point detection results, use the host timestamp as a unified standard, align the timestamps of the acquisition cameras' acquisition videos, and output the image data with synchronized timestamps: I front (t),I side (t); Pre-calibrate the strict mapping relationship between the actual physical space of the site and the captured camera image, and perform multi-view data association on the skeleton key point detection results based on the pre-calibrated mapping relationship to obtain the person matching results including the person ID and skeleton key points. The mapping relationship is established through precise acquisition camera installation position, measurement and perspective calibration experiments. When the target is located in the physical area in the upper left corner of the site, according to the pre-calibrated mapping relationship, the matching is achieved through the following specific steps: First, all target skeleton point data appearing in the right area of the front view image are detected in real time, and the target skeleton point data in the left area of the side view image are detected at the same time. Then, a spatial consistency check is performed to check whether the relative position relationship of the two groups of skeleton points conforms to the pre-calibrated spatial mapping rules. Then, a motion state check is performed to verify whether the motion trends reflected by the two groups of skeleton points are consistent. When all the checks are passed, the same ID number is assigned to the two groups of skeleton point data, that is: Combine the ID number with each person's skeleton point data, and use the number to indicate the corresponding relationship with the skeleton point data: (P IDi (t),C IDi (t))=(P i (t),C i (t)) Based on the motion area division method, multi-person detection is performed on the person matching results to obtain multi-person detection results.
7. The multi-perspective children's motor coordination ability assessment method based on deep learning according to claim 6, characterized in that: The method for quantitatively analyzing the test results of multiple people using the sports ability assessment model includes: Load multiple test results, conduct coordination evaluation on these results based on the motor ability assessment model, and obtain the child health assessment score corresponding to the health standard; The sports ability assessment model analyzes the health assessment score based on the health standard, determines whether the health assessment score is qualified, and obtains the health assessment result; The sports ability assessment model is combined with the health assessment results to generate a personalized health assessment report; Among them, the personalized health assessment report includes: assessment indicator analysis, health scores and suggestions, and training suggestions.
8. The multi-perspective children's motor coordination ability assessment method based on deep learning according to claim 7, characterized in that: The method for evaluating the coordination of multiple test results based on the sports ability assessment model includes: Obtain multi-person detection results, and identify the pixel coordinates of the skeleton points in the front and side views, the difference in the pixel coordinates of the skeleton points between multiple frames, and the spatial mapping under the front and side views. Graph the skeleton key point detection results of the front and side views respectively, and regard the skeleton key point detection results of the front and side views as a skeleton point graph topology. The nodes E represents the edge between corresponding bone points The edge weight is defined based on the Euclidean distance between skeleton points; Among them, when constructing the graph of the skeleton key point detection results of the front view and the side view, the pixel coordinates of the skeleton points of the front view and the side view, the difference value of the pixel coordinates of the skeleton points between multiple frames, and the spatial mapping under the two views are used as input. Its mathematical expression is as follows: Among them, R represents a multidimensional array in the real field, T is the length of the time series, 2 represents two different perspectives, and P i (t) represents the bone point number in the t-th frame, C i (t) represents the coordinates of the corresponding bone point, ΔX represents the coordinate difference between adjacent frames, and M is a weight matrix that represents the spatial mapping relationship between the two perspectives; The skeleton point graph topology of each view is represented as: Among them, A i,j Represents the connection weight between bone point i and bone point j; The graph convolutional network GCN is used to perform convolution operations on the skeleton point graph topology of each view, and the differential value ΔX is integrated into the graph convolution operation to enhance the model's sensitivity to action changes and extract local structural features. The graph convolution operation is expressed as: in, is the node feature matrix of the lth layer, D (l) The dimension of the feature vector of the node in the first layer of the graph convolutional network is 64, which is adjusted according to the performance of the model during training. is the degree matrix, represents the self-connection of the node, W (l) is the weight matrix of the lth layer, σ is the nonlinear activation function ReLU, and after L layers of graph convolution, the output features of each view are: The feature sequence after graph convolution is input into the recurrent neural network (RNN). The RNN captures the time series features and performs average pooling on the features of each time step to obtain the feature representation of each time step. The RNN outputs of the two perspectives are fused through a fusion layer to obtain the final feature vector. Among them, for each perspective, the input of the recurrent neural network RNN is the feature sequence after graph convolution: Perform average pooling on the features of each time step to obtain the feature representation of each time step: The output of the RNN is represented as: The RNN outputs of the two perspectives are fused through a fusion layer to obtain the final feature vector, which is then fused using a weighted summation method: Where M is the view weight matrix, represents matrix multiplication, and finally, the dimension of the fused feature vector is: H fusion ∈R T×H The feature vector is input into two fully connected layers and the output layer, and the score of each parameter is obtained through the Sigmoid activation function. The scores of each parameter are weighted and summed to output the health assessment score.
9. A multi-perspective children's motor coordination ability assessment system based on deep learning, for implementing the multi-perspective children's motor coordination ability assessment method based on deep learning according to any one of claims 1 to 8, characterized in that: The multi-perspective children's motor coordination ability assessment system based on deep learning includes: The model building module is used to construct a dedicated children's motion dataset, capture a motion evaluation dataset based on the children's motion dataset, integrate the children's motion dataset and the motion evaluation dataset into a modeling dataset, use the modeling dataset to train a skeleton point detection model and a motion ability evaluation model, and output a converged skeleton point detection model and a motion ability evaluation model; The data acquisition module collects real-time video streams of people entering the venue based on the acquisition camera, preprocesses the real-time video streams, and outputs preprocessed data sets; The visual recognition module is used to load the preprocessed data set. The skeleton point detection model is based on the combination of deep convolutional neural network and regression method to identify and analyze the preprocessed data set, identify the skeleton key points in the preprocessed data set, and obtain the skeleton key point detection results; The motion state detection module obtains the skeleton key point detection results, performs multi-view data association on the skeleton key point detection results, associates the person ID with the corresponding skeleton key points, obtains the person matching results containing the person ID and skeleton key points, and performs multi-person detection on the person matching results based on the motion area division method to obtain multi-person detection results; The personalized assessment module takes the test results of multiple people as input and executes the sports ability assessment model. The sports ability assessment model quantitatively analyzes the test results of multiple people and generates a personalized health assessment report.
Citation Information
Patent Citations
Posture health monitoring method and system based on human body key point detection
CN117457193A
Motion estimation method and device based on multi-modal perception and electronic equipment
CN119533462A
End-to-end human behavior recognition method and model based on skeleton nodes
CN114613013A
Three-dimensional human body posture estimation method based on multi-view geometry
CN116206328A
Posture assessment method based on human skeleton point and action recognition
CN117373109A
Cited By
Real-time safety monitoring system based on motion state analysis
CN121122741A
Newborn feeding behavior image recognition and oral movement evaluation system
CN121482844A
Method and system for evaluating structural performance of tension band of sportswear
CN121503169A
Neural circuit feedback-based dyskinesia treatment analysis method and system
CN121709268A
Method and system for analysis of movement disorder treatment based on neural circuit feedback
CN121709268B