Multi-view child motion coordination ability evaluation system and method based on deep learning
Patent Information
- Application Number
- CN202510916539.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2045-07-03
AI Technical Summary
[0007]本发明的目的在于针对现有技术的不足之处,提供基于深度学习的多视角儿童运动协调能力评估系统及方法,解决了现有技术对儿童运动评估主要依赖人工观察和简单物理测试,由于评估者的经验和注意力有限,难以同步追踪儿童动作的多个角度和关键细节,使得现有方法无法自动、准确、高效评估儿童动态动作,特别是在涉及复杂运动时,传统方法难以提供全面、高效、精准标准统一的评估的问题
[0080]本发明通过骨骼点检测模型基于深度卷积神经网络与回归方法相结合对预处理数据集识别分析,并基于运动区域划分法对人员匹配结果进行多人检测,结合了视觉识别与运动状态检测技术,且采用正视角与侧视角双摄像头布局,配合深度学习算法,实现了自动化的儿童动态动作评估。相较于人工测量或单一视角分析,能够完整捕捉动作过程,确保评估标准的一致性,消除人为误差,提高了检测精准度与效率,且支持多人同时检测,利用采集摄像头的视野分区技术,为进入场地的每位儿童分配唯一身份编号,并进行独立的运动轨迹跟踪,确保不同个体的动作数据不会混淆,实现高效的多目标动作分析。
Smart Images

Figure CN120809152B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of sports assessment technology, specifically relating to a multi-perspective assessment system and method for children's motor coordination ability based on deep learning. Background Technology
[0002] Motor coordination in children is a core indicator of neurodevelopment, and abnormalities in this area are closely related to developmental coordination disorder (DCD). According to the World Health Organization, approximately 5%-6% of school-aged children worldwide have varying degrees of motor coordination deficits, manifesting as abnormal gait symmetry, delayed motor initiation, and bilateral coordination impairment. In severe cases, this can lead to decreased learning ability and social avoidance behaviors. The internationally recognized "Movement Assessment Battery for Children - Second Edition" (MABC-2) tool indicates that a fine motor control error exceeding 1.5 standard deviations is considered abnormal. In specific movement testing scenarios such as the lateral gliding test, accurately assessing indicators such as shoulder and foot coordination, movement continuity, and bilateral symmetry can reveal a child's neuromuscular coordination ability. For example, the parallelism between the shoulder and the direction of movement, and the parallelism of the shoulder area to the ground can reflect trunk stability, while the angle between the shoulder direction and the foot movement direction characterizes overall coordination and gait control ability. If a child exhibits significant shoulder swaying or a tilt angle exceeding 5° during the test, it may indicate insufficient core muscle control or delayed vestibular function development. Uneven gliding amplitude of both legs and short or insufficient airtime may be related to limited cerebellar function regulation. However, traditional assessment methods are difficult to capture millisecond-level coordination deviations in dynamic movements.
[0003] Traditional pediatric motor assessments primarily rely on manual observation and simple physical tests. Due to the limited experience and attention of assessors, it's difficult to simultaneously track multiple angles and key details of a child's movements, such as horizontal shoulder deviation, gait imbalances, or insufficient airborne phases. This leads to inconsistent assessment standards, and the same child may receive different results from different assessors, affecting the reliability of the final assessment. Furthermore, traditional methods lack high-precision temporal data and trajectory analysis, making it difficult to provide quantifiable and repeatable objective indicators. This makes it difficult to accurately identify early abnormalities such as developmental coordination disorder (DCD) or sensory integration dysfunction, potentially leading to missed intervention opportunities.
[0004] Chinese patent CN119533462A discloses a motion estimation method, device, and electronic device based on multimodal perception. The method includes acquiring motion assessment-related data for a target area, including IMU sensor data and flexible deformation sensor data. Based on the IMU sensor data, the pose information of the area where the IMU sensor is configured can be determined. Then, based on the positional relationship between the configured IMU sensor and the flexible deformation sensor, and the acquired flexible deformation sensor data, the pose information of the area where the flexible deformation sensor is configured is calculated. This allows for obtaining pose information for locations without IMU sensors, even when the number of configured IMU sensors is limited. However, existing methods require wearable devices, can only measure partial body movements, cannot cover all parts of the body, and have limited accuracy. In scenarios involving the assessment of whole-body coordination, the inability to comprehensively acquire motion data from all body parts may affect the comprehensiveness and accuracy of the assessment results.
[0005] Chinese patent CN117457193A discloses a method and system for monitoring postural health based on human key point detection. The method includes the following steps: optical human image acquisition, human key point detection, human detection box recognition, multi-human tracking, posture detection of key human body parts, and early warning of poor posture. This method can detect the postural health of multiple individuals appearing within the field of view and provide early warnings for each, generating a postural health report. However, existing methods mainly focus on single-posture detection to assess children's skeletal development, making it difficult to adapt to children's flexible and varied movement characteristics and behavioral patterns, and unable to assess dimensions such as motor coordination, frequency, and duration.
[0006] Therefore, existing technologies have significant shortcomings in addressing the specific needs of children's coordination and motor skills assessment: children have a large range of motion and their movements change rapidly, making efficient daily group assessments difficult. This necessitates that the system possess core capabilities such as multi-angle synchronous acquisition, real-time tracking of multiple users, and high-precision motion analysis. Currently, there is an urgent need to develop a new assessment scheme that maintains the naturalness of non-contact measurement while accurately capturing rapidly changing motion characteristics through multi-view fusion technology, enabling the identification of motor coordination, completion time, and number of repetitions. Simultaneously, it should support simultaneous assessment by multiple users, truly meeting the large-scale, routine monitoring needs of children's motor abilities in practical scenarios such as kindergartens and schools. To address these issues, we propose a deep learning-based multi-view children's motor coordination ability assessment system and method. Summary of the Invention
[0007] The purpose of this invention is to address the shortcomings of existing technologies by providing a multi-perspective assessment system and method for children's motor coordination ability based on deep learning. This solves the problem that existing technologies mainly rely on manual observation and simple physical tests to assess children's motor skills. Due to the limited experience and attention of the assessors, it is difficult to simultaneously track multiple angles and key details of children's movements. As a result, existing methods cannot automatically, accurately, and efficiently assess children's dynamic movements, especially when it comes to complex movements. Traditional methods are unable to provide a comprehensive, efficient, accurate, and standardized assessment.
[0008] This invention is implemented as follows: a deep learning-based multi-view assessment method for children's motor coordination ability, comprising:
[0009] A dedicated children's sports dataset is constructed. A sports assessment dataset is extracted based on the children's sports dataset. The children's sports dataset and the sports assessment dataset are integrated into a modeling dataset. The modeling dataset is used to train a skeletal point detection model and a sports ability assessment model. The converged skeletal point detection model and sports ability assessment model are output.
[0010] Based on the real-time video stream of people entering the venue captured by the acquisition camera, the real-time video stream is preprocessed and a preprocessed dataset is output.
[0011] Load the preprocessed dataset, and the skeleton point detection model uses a combination of deep convolutional neural network and regression method to identify and analyze the preprocessed dataset, identify the key skeleton points in the preprocessed dataset, and obtain the skeleton key point detection results;
[0012] Obtain the skeletal keypoint detection results, perform multi-view data association on the skeletal keypoint detection results, associate the person ID with the corresponding skeletal keypoint, obtain the person matching results containing the person ID and skeletal keypoint, and perform multi-person detection on the person matching results based on the motion region segmentation method to obtain multi-person detection results.
[0013] Using the test results of multiple individuals as input, an exercise capacity assessment model is executed. The exercise capacity assessment model performs quantitative analysis on the test results of multiple individuals and generates a personalized health assessment report.
[0014] Preferably, the method for constructing a children's motion dataset includes:
[0015] The system uses multi-angle, full-body shooting to collect specialized video data of children in different scenarios. The video data covers common movements such as running, jumping, and side-sliding, as well as children's natural movements during play or daily activities.
[0016] Load dedicated video data, parse and process the dedicated video data frame by frame to obtain at least one set of dedicated images, process the children's skeletal point markings in the dedicated images to obtain a children's motion dataset containing the skeletal point marking results. In the process of processing the children's skeletal point markings in the dedicated images, 17 skeletal points of the children are assigned fixed numbers, namely: 1. nose, 2. left eye, 3. right eye, 4. left ear, 5. right ear, 6. left shoulder, 7. right shoulder, 8. left elbow, 9. right elbow, 10. left wrist, 11. right wrist, 12. left hip joint, 13. right hip joint, 14. left knee, 15. right knee, 16. left ankle, 17. right ankle;
[0017] Based on the children's movement dataset, a movement assessment dataset is extracted. The movement assessment dataset is based on abnormal movement data of children with motor coordination deviations. The children's movement dataset and the movement assessment dataset are integrated into the modeling dataset.
[0018] Preferably, the method for training the skeletal point detection model and the motion capability assessment model using a modeling dataset includes:
[0019] Load the modeling dataset and divide it into a training set and a validation set. The training set is used to train the skeleton point detection model and the motion ability assessment model, while the validation set is used to verify the generalization ability of the skeleton point detection model and the motion ability assessment model. The ratio of the training set to the validation set is 4:1.
[0020] The skeletal point detection model is iteratively trained using a training set and a validation set to output a converged skeletal point detection model. The training set includes skeletal point annotation data and movement data of normal and abnormal children. The training set is used to train the skeletal point detection model and the movement ability assessment model, so that the skeletal point detection model can identify the joint positions of children in different movements and distinguish between normal and abnormal movement patterns.
[0021] Load the skeleton point detection model. The skeleton point detection model combines data from the frontal and side views to stably identify the skeleton points of children under different views and output the skeleton point recognition results of the skeleton point detection model.
[0022] Obtain the skeleton point recognition results, use the skeleton point recognition results to train and optimize the motion ability assessment model, and output the motion ability assessment model based on multi-view sequence skeleton point image data.
[0023] The athletic ability assessment model includes a feature extraction module and a regression scoring module. The feature extraction module consists of a graph construction layer, a graph convolutional network (GCN), a recurrent neural network (RNN), an average pooling layer, and a feature fusion layer, all connected in sequence. The regression scoring module is used to output the final health assessment score. It consists of a fully connected layer 1, a fully connected layer 2, and an output layer, all connected in sequence. The fully connected layer 1 is used to obtain the feature vector output by the feature fusion layer. Data is passed between fully connected layer 1 and fully connected layer 2 through the ReLU activation function. Data is also passed between fully connected layer 2 and the output layer through the ReLU activation function. The output layer outputs the scoring vector through the Sigmoid activation function.
[0024] Preferably, the method for preprocessing real-time video streams includes:
[0025] The camera captures video streams in real time at a fixed frame rate of 60 FPS to obtain a real-time video stream;
[0026] Load the real-time video stream, parse and process the real-time video stream frame by frame, and number and store the real-time images processed frame by frame in chronological order.
[0027] The real-time images, after being processed according to time sequence, are converted to a new format, grayscale normalization is performed, and a smoothing filter is used to remove noise from the images, resulting in a preprocessed dataset.
[0028] Preferably, the method for identifying skeletal key points in the preprocessed dataset includes:
[0029] Load the preprocessed dataset, and the skeleton point detection model extracts key features from the preprocessed dataset through convolutional layers, identifies multiple candidate locations of skeleton points in the preprocessed dataset, and performs preliminary localization of the skeleton points.
[0030] The preliminary location results of skeletal points are obtained. Then, the preliminary location results of skeletal points are classified and accurately categorized according to human body structure by using clustering algorithms combined with regression analysis to obtain the detection results of skeletal key points.
[0031] When classifying and precisely categorizing the preliminary skeletal point localization results according to human anatomy, each class is treated as a target, and the set of skeletal point numbers for each target is output first:
[0032] P i (t)={p1,p2,...,p n}
[0033] Where p kThe skeletal point numbers are as follows: 1. Nose, 2. Left eye, 3. Right eye, 4. Left ear, 5. Right ear, 6. Left shoulder, 7. Right shoulder, 8. Left elbow, 9. Right elbow, 10. Left wrist, 11. Right wrist, 12. Left hip joint, 13. Right hip joint, 14. Left knee, 15. Right knee, 16. Left ankle, 17. Right ankle;
[0034] Each number corresponds to a two-dimensional coordinate, representing the position of that skeletal point in the image:
[0035] C i (t)={(x1,y1),(x2,y2),...,(x n ,y n )}
[0036] Where (x) k ,y k ) represents the skeletal point p k Coordinates on the image.
[0037] Preferably, the method for multi-view data association of skeletal key point detection results includes:
[0038] Load the skeletal keypoint detection results, identify the acquisition cameras associated with the images in the skeletal keypoint detection results, align the timestamps of the acquired videos from the acquisition cameras using the host timestamp as a unified standard, and output time-stamp synchronized image data: I front (t),I side (t);
[0039] A strict mapping relationship between the actual physical space of the site and the captured camera images is pre-calibrated. Based on this pre-calibrated mapping relationship, multi-view data association is performed on the skeletal keypoint detection results to obtain personnel matching results containing person IDs and skeletal keypoints. The mapping relationship is established through precise camera installation positions, measurements, and viewpoint calibration experiments. When the target is located in the upper left corner of the site, matching is achieved through the following specific steps based on the pre-calibrated mapping relationship: First, all target skeletal point data appearing in the right area of the frontal view are detected in real time, and target skeletal point data in the left area of the side view are also detected. Then, spatial consistency verification is performed to check whether the relative positional relationship of the two sets of skeletal points conforms to the pre-calibrated spatial mapping rules. Then, motion state verification is performed to verify whether the motion trends reflected by the two sets of skeletal points are consistent. When all verifications pass, the same ID number is assigned to the two sets of skeletal point data, i.e.:
[0040]
[0041] The ID number is combined with each person's skeletal point data, and the ID number represents the correspondence between the ID number and the skeletal point data:
[0042]
[0043] The motion region segmentation method is used to perform multi-person detection on the matching results of the personnel, and the multi-person detection results are obtained.
[0044] Preferably, the method for quantitative analysis of multiple test results using the athletic ability assessment model includes:
[0045] Load the test results of multiple people, evaluate the coordination of the test results based on the motor ability assessment model, and obtain the children's health assessment score corresponding to the health standard;
[0046] The athletic ability assessment model analyzes health assessment scores based on health standards, determines whether the health assessment scores are qualified, and obtains the health assessment results.
[0047] The athletic ability assessment model, combined with health assessment results, generates a personalized health assessment report.
[0048] The personalized health assessment report includes: assessment indicator analysis, health score and recommendations, and training suggestions.
[0049] The method for evaluating coordination of multiple-person test results based on a motion ability assessment model includes:
[0050] Obtain multi-person detection results and identify the skeletal pixel coordinates of the frontal and side views, the differences in skeletal pixel coordinates between multiple frames, and the spatial mapping under the frontal and side views. Construct graphs for the skeletal keypoint detection results of the frontal and side views respectively, treating the skeletal keypoint detection results of the frontal and side views as a skeletal point graph topology. Among the nodes E represents the edge between corresponding bone points. The weight of the edge is defined based on the Euclidean distance between the skeletal points;
[0051] Specifically, when constructing graphs from the skeletal keypoint detection results of the frontal and side views, the skeletal pixel coordinates of the frontal and side views, the differences in skeletal pixel coordinates between multiple frames, and the spatial mappings under the two views are used as inputs, and their mathematical representation is as follows:
[0052]
[0053] Where R represents a multidimensional array over the real number field, T is the length of the time series, 2 represents two different perspectives, and P i (t) represents the skeleton point number in frame t, C i (t) represents the coordinates of the corresponding skeleton point, ΔX represents the coordinate difference obtained by calculating the coordinate difference between adjacent frames, and M is a weight matrix that represents the spatial mapping relationship between two viewpoints;
[0054] The topological representation of the skeletal point map for each viewpoint is as follows:
[0055]
[0056] Among them, A i,j This represents the connection weight between skeletal point i and skeletal point j;
[0057] The Graph Convolutional Network (GCN) is used to perform convolution operations on the skeletal point graph topology for each viewpoint, and the difference value ΔX is integrated into the graph convolution operation to enhance the model's sensitivity to action changes and extract local structural features. The graph convolution operation is represented as follows:
[0058]
[0059] in, D is the node feature matrix of the l-th layer. (l) This represents the dimension of the feature vector of the l-th layer node in the graph convolutional network. The initial dimension was designed to be 64, and it was adjusted based on the model's performance during training. It is a degree matrix. W represents a self-join of nodes. (l) Here is the weight matrix of the l-th layer, and σ is the non-linear activation function ReLU. After L layers of graph convolution, the output features for each viewpoint are:
[0060]
[0061] The feature sequence after graph convolution is input into a recurrent neural network (RNN). The RNN captures the time series features and performs average pooling on the features at each time step to obtain the feature representation at each time step. The outputs of the two RNNs from different perspectives are fused through a fusion layer to obtain the final feature vector.
[0062] For each viewpoint, the input to the recurrent neural network (RNN) is the feature sequence after graph convolution:
[0063]
[0064] Average pooling is performed on the features at each time step to obtain the feature representation at each time step:
[0065]
[0066] The output of the RNN is represented as:
[0067]
[0068] The outputs of the two RNN perspectives are fused through a fusion layer to obtain the final feature vector, and then a weighted summation is used for the fusion:
[0069]
[0070] Where M is the viewpoint weight matrix, · denotes matrix multiplication, and finally, the dimension of the fused feature vector is:
[0071] H fusion ∈R T×H
[0072] The feature vector is input into two fully connected layers and the output layer. The score of each parameter is obtained by passing the sigmoid activation function. The scores of each parameter are weighted and summed to output the health assessment score.
[0073] On the other hand, the present invention also provides a deep learning-based multi-view children's motor coordination ability assessment system, the deep learning-based multi-view children's motor coordination ability assessment system comprising:
[0074] The model building module is used to build a dedicated children's sports dataset. Based on the children's sports dataset, a sports assessment dataset is extracted. The children's sports dataset and the sports assessment dataset are integrated into the modeling dataset. The modeling dataset is used to train the skeleton point detection model and the sports ability assessment model, and the converged skeleton point detection model and sports ability assessment model are output.
[0075] The data acquisition module collects real-time video streams of people entering the venue from the acquisition cameras, preprocesses the real-time video streams, and outputs a preprocessed dataset.
[0076] The visual recognition module is used to load the preprocessed dataset. The skeleton point detection model is based on a combination of deep convolutional neural networks and regression methods to identify and analyze the preprocessed dataset, identify the key skeleton points in the preprocessed dataset, and obtain the key skeleton point detection results.
[0077] The motion state detection module acquires the skeletal key point detection results, performs multi-view data association on the skeletal key point detection results, associates the person ID with the corresponding skeletal key points, and obtains the person matching results containing the person ID and skeletal key points. Based on the motion region division method, the person matching results are used to perform multi-person detection and obtain multi-person detection results.
[0078] The personalized assessment module takes the test results of multiple people as input, executes the exercise ability assessment model, and performs quantitative analysis on the test results of multiple people to generate a personalized health assessment report.
[0079] Compared with the prior art, the embodiments of this application have the following main advantages:
[0080] This invention utilizes a skeletal point detection model based on a combination of deep convolutional neural networks and regression methods to identify and analyze preprocessed datasets. It then employs a motion region segmentation method to perform multi-person detection based on the matching results. Combining visual recognition and motion state detection technologies, and employing a dual-camera layout with both frontal and side views, along with deep learning algorithms, it achieves automated dynamic motion assessment of children. Compared to manual measurement or single-view analysis, it can completely capture the motion process, ensuring consistency in assessment standards, eliminating human error, and improving detection accuracy and efficiency. Furthermore, it supports simultaneous detection of multiple individuals. By utilizing the field-of-view segmentation technology of the acquisition cameras, each child entering the venue is assigned a unique identification number and their movement trajectory is tracked independently, ensuring that the motion data of different individuals are not mixed up, thus achieving efficient multi-target motion analysis.
[0081] In this embodiment of the invention, the acquired video stream is parsed frame by frame and numbered and stored in chronological order, ensuring that each frame has a clear timestamp and sequence identifier. This not only facilitates maintaining the temporal consistency of the data during subsequent processing but also enables rapid location and retracing of motion states at specific moments when needed, providing convenience for motion trajectory analysis and anomaly detection. The preprocessed dataset has a unified format, resolution, and quality standard, allowing it to be directly input into the skeleton point detection model and other analysis modules. This avoids the need for additional conversion or preprocessing work in subsequent processing stages due to inconsistent data formats or varying quality, significantly improving the overall efficiency and real-time performance of the evaluation system.
[0082] In this embodiment of the invention, the skeletal point detection model extracts key features from the preprocessed dataset through convolutional layers, effectively identifying important information related to skeletal points in the image. This provides a precise basis for subsequent skeletal point localization, reduces interference from irrelevant features, and makes skeletal point recognition more accurate. Using clustering algorithms combined with regression analysis, the initially located skeletal points are classified and precisely categorized based on human anatomy, fully considering the anatomical characteristics and movement patterns of the human skeleton. This effectively corrects potential deviations in the initial localization, further improving the accuracy of skeletal key point recognition. The recognition results for each skeletal point are more consistent with actual human postures. The model employs a skeletal point detection model specifically trained on children's data, rather than a general action recognition model, exhibiting significant advantages in the accuracy of children's action recognition. This model can more accurately capture children's unique movement characteristics, improving the reliability of the assessment and ensuring strong targeting and broad adaptability.
[0083] In this embodiment of the invention, by performing multi-view data association on the detection results of skeletal key points and aligning the image data acquired by the dual cameras using a unified host timestamp, the consistency of the two video streams in the temporal dimension is ensured, avoiding matching errors caused by time asynchrony. This lays the foundation for accurate association of skeletal point data in the future, which is crucial for analyzing children's movement patterns over time. Based on the strict mapping relationship between the pre-calibrated physical space of the venue and the camera images, matching is performed using spatial position features, which can accurately match the skeletal point data of the same person from different perspectives. This effectively reduces mismatches caused by differences in perspective or scene complexity, improves the accuracy of data association, and thus ensures the correct assessment of children's movement status. Integrating skeletal point data from the frontal and side perspectives can comprehensively and three-dimensionally reflect children's movement in three-dimensional space, providing richer and more complete information for assessing key indicators such as coordination, continuity, and symmetry of movements. This helps to discover potential movement problems, such as asymmetry of movements and abnormal postures, and provides a more sufficient basis for developing personalized sports training programs. Attached Figure Description
[0084] Figure 1 This is a schematic diagram illustrating the implementation process of the deep learning-based multi-perspective children's motor coordination ability assessment method provided by the present invention.
[0085] Figure 2 This diagram illustrates the division of the detection site when performing multi-view data association on the detection results of skeletal key points.
[0086] Figure 3 This is a schematic diagram of the structure of the deep learning-based multi-view children's motor coordination ability assessment system provided by the present invention.
[0087] Figure 4 A flowchart illustrating a method for evaluating coordination in multi-person testing based on a motor ability assessment model is presented. Detailed Implementation
[0088] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0089] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0090] Existing technologies for assessing children's motor skills primarily rely on manual observation and simple physical tests. Due to the limited experience and attention of assessors, it is difficult to simultaneously track multiple angles and key details of children's movements. This makes existing methods unable to automatically, accurately, and efficiently assess children's dynamic movements, especially when dealing with complex movements. Traditional methods struggle to provide comprehensive, efficient, accurate, and standardized assessments. To address these issues, we propose a multi-view children's motor coordination ability assessment system and method based on deep learning. This invention uses a skeletal point detection model combined with deep convolutional neural networks and regression methods to identify and analyze preprocessed datasets. It also uses a motion region segmentation method to perform multi-person detection based on matching results. Combining visual recognition and motion state detection technologies, and employing a dual-camera layout with frontal and side views, along with deep learning algorithms, it achieves automated assessment of children's dynamic movements. Compared to manual measurement or single-view analysis, it can completely capture the movement process, ensuring consistency in assessment standards, eliminating human error, and improving detection accuracy and efficiency. Furthermore, it supports simultaneous detection of multiple individuals. Utilizing the field-of-view partitioning technology of the acquisition cameras, each child entering the venue is assigned a unique identification number and their movement trajectory is tracked independently, ensuring that the movement data of different individuals are not mixed up, achieving efficient multi-target movement analysis.
[0091] This invention provides a multi-perspective method for assessing children's motor coordination ability based on deep learning. Figure 1 A schematic diagram illustrating the implementation process of a deep learning-based multi-view assessment method for children's motor coordination ability is shown. This deep learning-based multi-view assessment method for children's motor coordination ability specifically includes:
[0092] Step S10: Construct a dedicated children's sports dataset. Based on the children's sports dataset, capture a sports assessment dataset. Integrate the children's sports dataset and the sports assessment dataset as a modeling dataset. Use the modeling dataset to train a skeletal point detection model and a sports ability assessment model. Output a converged skeletal point detection model and a sports ability assessment model.
[0093] Step S20: Based on the real-time video stream of people entering the venue captured by the acquisition camera, preprocess the real-time video stream and output the preprocessed dataset;
[0094] Step S30: Load the preprocessed dataset. The skeletal point detection model uses a combination of deep convolutional neural networks and regression methods to identify and analyze the preprocessed dataset, identify key skeletal points in the preprocessed dataset, and obtain the key skeletal point detection results.
[0095] Step S40: Obtain the skeletal key point detection results, perform multi-view data association on the skeletal key point detection results, associate the person ID with the corresponding skeletal key points, obtain the person matching results containing the person ID and skeletal key points, and perform multi-person detection on the person matching results based on the motion region division method to obtain multi-person detection results.
[0096] Step S50: Using the test results of multiple people as input, execute the exercise ability assessment model. The exercise ability assessment model performs quantitative analysis on the test results of multiple people and generates a personalized health assessment report.
[0097] This invention utilizes a skeletal point detection model based on a combination of deep convolutional neural networks and regression methods to identify and analyze preprocessed datasets. It then employs a motion region segmentation method to perform multi-person detection based on the matching results. Combining visual recognition and motion state detection technologies, and employing a dual-camera layout with both frontal and side views, along with deep learning algorithms, it achieves automated dynamic motion assessment of children. Compared to manual measurement or single-view analysis, it can completely capture the motion process, ensuring consistency in assessment standards, eliminating human error, and improving detection accuracy and efficiency. Furthermore, it supports simultaneous detection of multiple individuals. By utilizing the field-of-view segmentation technology of the acquisition cameras, each child entering the venue is assigned a unique identification number and their movement trajectory is tracked independently, ensuring that the motion data of different individuals are not mixed up, thus achieving efficient multi-target motion analysis.
[0098] This invention provides a method for constructing a children's sports dataset, the method specifically including:
[0099] Step S101: Use multi-angle, full-body shooting to collect dedicated video data of children in different scenarios. The video data covers common sports movements such as running, jumping, and side-sliding steps, as well as natural movements of children during play or daily activities. These images provide diverse sources of sports data, which helps improve the model's ability to recognize different sports situations.
[0100] Step S102: Load dedicated video data, analyze and process the dedicated video data frame by frame to obtain at least one set of dedicated images, process the children's skeletal point markings in the dedicated images to obtain a children's motion dataset containing the skeletal point marking results. When processing the children's skeletal point markings in the dedicated images, set fixed numbers for the 17 skeletal points of the children, which are: 1. nose, 2. left eye, 3. right eye, 4. left ear, 5. right ear, 6. left shoulder, 7. right shoulder, 8. left elbow, 9. right elbow, 10. left wrist, 11. right wrist, 12. left hip joint, 13. right hip joint, 14. left knee, 15. right knee, 16. left ankle, 17. right ankle;
[0101] It should be noted that when the dedicated video data is analyzed frame by frame to obtain at least one set of dedicated images, the dedicated images contain image data of various movement patterns of children. All skeletal points are manually annotated by experts to ensure that the model can accurately identify the location of key skeletal points in children's movements.
[0102] Step S103: Based on the children's movement dataset, capture the movement assessment dataset. The movement assessment dataset is based on the abnormal movement data of children with movement coordination deviations. Integrate the children's movement dataset and the movement assessment dataset into a modeling dataset.
[0103] In this embodiment, based on the children's movement dataset, further movement data were collected from children with normal development and those with differences in motor development, constructing a movement assessment dataset specifically for evaluating children's coordination abilities. For example, in the lateral gliding test, children with poor coordination may have insufficient vestibular development or weak core muscle control, leading to asymmetrical shoulder and pelvic rotation and difficulty maintaining upper body stability. Some children, due to insufficient proprioception, may exhibit compensatory trunk lateral flexion, that is, their bodies unconsciously tilt to one side during the gliding process to maintain balance. In addition, children with abnormal gait coordination may show a misalignment between shoulder direction and foot movement direction, or significant pauses during the gliding process, making it impossible to complete the movement smoothly. Conventional models have low accuracy in recognizing such data, making it difficult to support precise quantitative assessment.
[0104] To improve the ability of skeletal point detection models and motor ability assessment models to identify the above-mentioned abnormalities, abnormal movement data of children with motor coordination deviations were specifically collected. For example, some children have sensory integration dysfunction, resulting in asynchronous coordination between their left and right sides. This manifests as one shoulder being significantly higher than the other during gliding, and the angle between the shoulder skeletal point and the ground significantly deviating from the horizontal baseline (e.g., exceeding 5°). Another example is children with insufficient core stability who have difficulty maintaining a stable center of gravity during gliding, leading to significant deviations in their gliding trajectory and excessively large differences in stride length indicated by skeletal points. The collection and analysis of this abnormal data helps skeletal point detection models and motor ability assessment models accurately identify deviations in children's movement patterns and provides a more reliable basis for subsequent assessments.
[0105] This invention provides a method for training a skeletal point detection model and a motion capability assessment model using a modeling dataset. The method specifically includes:
[0106] Step S201: Load the modeling dataset and divide it into a training set and a validation set. The training set is used to train the skeleton point detection model and the motion ability assessment model, while the validation set is used to verify the generalization ability of the skeleton point detection model and the motion ability assessment model. The ratio of the training set to the validation set is 4:1.
[0107] Step S202: The skeletal point detection model is iteratively trained using a training set and a validation set to output a converged skeletal point detection model. The training set includes skeletal point annotation data and movement data of normal and abnormal children. The training set is used to train the skeletal point detection model and the movement ability assessment model so that the skeletal point detection model can identify the joint positions of children in different movement movements and distinguish between normal and abnormal movement patterns.
[0108] Step S203: Load the skeleton point detection model. The skeleton point detection model combines data from the frontal and side views to stably identify the skeleton points of the child under different views and output the skeleton point recognition results of the skeleton point detection model.
[0109] Step S204: Obtain the skeletal point recognition results, use the skeletal point recognition results to train and optimize the motion ability assessment model, and output the motion ability assessment model based on multi-view sequence skeletal point image data.
[0110] In this embodiment, by combining data from both frontal and side views, the skeletal point detection model can stably identify children's skeletal points under different perspectives. Simultaneously, through optimization algorithms tailored to children's physical characteristics, recognition biases caused by differences in height and body shape are further reduced. After the skeletal point recognition model is trained, the skeletal point recognition results output by the skeletal point detection model are input into the second-stage motor ability assessment model. Through analysis of motor data, the child's motor state is evaluated. Through this training and optimization, compared to ordinary models, the motor ability assessment model has a higher accuracy rate in recognizing children's physical characteristics.
[0111] This invention provides a method for preprocessing real-time video streams, the method specifically including:
[0112] Step S301: The camera captures a video stream in real time at a fixed frame rate of 60 FPS to obtain a real-time video stream;
[0113] It should be noted that the acquisition camera can be in two sets, set on the front and side respectively. The acquisition camera acquires video stream in real time at a fixed frame rate of 60FPS to ensure complete recording of the child's entire movement process.
[0114] Step S302: Load the real-time video stream, parse and process the real-time video stream frame by frame, and number and store the real-time images processed frame by frame in chronological order.
[0115] Step S303: The real-time images processed according to the time sequence are converted into a new format and grayscale normalization is performed to reduce the impact of illumination changes. A smoothing filter is used to remove noise from the images and reduce minor errors caused by sensor noise or environmental interference, resulting in a preprocessed dataset. In addition, all image frames are normalized to a uniform resolution.
[0116] In this embodiment of the invention, the acquired video stream is parsed frame by frame and numbered and stored in chronological order, ensuring that each frame has a clear timestamp and sequence identifier. This not only facilitates maintaining the temporal consistency of the data during subsequent processing but also enables rapid location and retracing of motion states at specific moments when needed, providing convenience for motion trajectory analysis and anomaly detection. The preprocessed dataset has a unified format, resolution, and quality standard, allowing it to be directly input into the skeleton point detection model and other analysis modules. This avoids the need for additional conversion or preprocessing work in subsequent processing stages due to inconsistent data formats or varying quality, significantly improving the overall efficiency and real-time performance of the evaluation system.
[0117] This invention provides a method for identifying skeletal key points in a preprocessed dataset. The method specifically includes:
[0118] Step S401: Load the preprocessed dataset. The skeleton point detection model extracts key features from the preprocessed dataset through convolutional layers, identifies multiple candidate skeleton point locations in the preprocessed dataset, and performs preliminary localization of the skeleton points.
[0119] Step S402: Obtain preliminary skeletal point localization results. Use clustering algorithm combined with regression analysis to classify and accurately categorize the preliminary skeletal point localization results according to human body structure, and obtain skeletal key point detection results.
[0120] When classifying and precisely categorizing the preliminary skeletal point localization results according to human anatomy, each class is treated as a target, and the set of skeletal point numbers for each target is output first:
[0121] P i (t)={p1,p2,...,p n}
[0122] Where p k The skeletal point numbers are as follows: 1. Nose, 2. Left eye, 3. Right eye, 4. Left ear, 5. Right ear, 6. Left shoulder, 7. Right shoulder, 8. Left elbow, 9. Right elbow, 10. Left wrist, 11. Right wrist, 12. Left hip joint, 13. Right hip joint, 14. Left knee, 15. Right knee, 16. Left ankle, 17. Right ankle;
[0123] Each number corresponds to a two-dimensional coordinate, representing the position of that skeletal point in the image:
[0124] C i (t)={(x1,y1),(x2,y2),...,(x n ,y n )}
[0125] Where (x) k ,y k ) represents the skeletal point p k Coordinates on the image.
[0126] In this embodiment, the skeletal point detection model extracts key features from the preprocessed dataset through convolutional layers, effectively identifying important information related to skeletal points in the image. This provides a precise basis for subsequent skeletal point localization, reduces interference from irrelevant features, and makes skeletal point recognition more accurate. Using clustering algorithms combined with regression analysis, the initially located skeletal points are classified and precisely categorized based on human anatomy, fully considering the anatomical characteristics and movement patterns of the human skeleton. This effectively corrects potential deviations in the initial localization, further improving the accuracy of skeletal key point recognition. The recognition results for each skeletal point are more consistent with actual human postures. A skeletal point detection model specifically trained on children's data, rather than a general action recognition model, is used, exhibiting significant advantages in the accuracy of children's action recognition. This model can more accurately capture children's unique movement characteristics, improving the reliability of the assessment and ensuring strong targeting and broad adaptability.
[0127] This invention provides a method for multi-view data association of skeletal keypoint detection results. The method specifically includes:
[0128] Step S501: Load the skeletal keypoint detection results, identify the acquisition cameras associated with the images in the skeletal keypoint detection results, and synchronously acquire video streams from the two acquisition cameras. Each frame of the image has a timestamp t, and the image data is output. To ensure that the images captured by the two acquisition cameras are consistent in timing, the system guarantees that the time difference between the streaming of the cameras is less than 100ms. Using the host timestamp as a unified standard, the timestamps of the acquired videos from the acquisition cameras are aligned, and the image data with synchronized timestamps is output: I front (t),I side (t);
[0129] In step S502, since the system uses two acquisition cameras (frontal view and side view), each individual entering the experimental site will be captured simultaneously from both perspectives. In multi-view data association, the system adopts a dual-view target matching method based on physical spatial location. A strict mapping relationship between the actual physical space of the site and the acquisition camera images is pre-calibrated. Based on the pre-calibrated mapping relationship, multi-view data association is performed on the skeletal keypoint detection results to obtain personnel matching results including person ID and skeletal keypoints. The mapping relationship is established through precise acquisition camera installation positions, measurements, and perspective calibration experiments. When the target is located in the upper left corner of the site's physical area, the system can determine the target's location based on the pre-calibrated mapping relationship. The target is determined to appear in the right-hand area of the frontal view and in the left-hand camera view, respectively. When the target is located in the upper left corner of the field, matching is achieved through the following steps based on the pre-calibrated mapping relationship: First, all target skeletal point data appearing in the right-hand area of the frontal view are detected in real time, and target skeletal point data in the left-hand area of the side-view view are also detected. Then, a spatial consistency check is performed to check whether the relative positional relationship of the two sets of skeletal points conforms to the pre-calibrated spatial mapping rules. Next, a motion state check is performed to verify whether the motion trends reflected by the two sets of skeletal points are consistent. When all checks pass, the same ID number is assigned to the two sets of skeletal point data, i.e.:
[0130]
[0131] Step S503: Combine the ID number with the skeletal point data of each person, using the ID number to represent the correspondence between the ID number and the skeletal point data.
[0132]
[0133] Step S504: Perform multi-person detection based on the motion region segmentation method on the personnel matching results to obtain multi-person detection results.
[0134] In this embodiment, to enable simultaneous detection of multiple individuals, a motion region division method is employed to rationally partition the detection area, ensuring that the movements of each individual are fully captured and avoiding occlusion interference in frontal and side-view images. The detection area is divided as follows: Figure 2 The diagram shows multiple independent rectangular regions arranged diagonally, similar to a checkerboard pattern. Each region is non-overlapping, ensuring that only one individual is active in each region at any given time. This design prevents overlap between individuals in the frontal and side-view camera feeds, thus improving detection accuracy.
[0135] During multi-person detection, the system assigns individuals to pre-defined motion areas based on their initial positions in the camera feed, ensuring that each area corresponds to only one individual. This is crucial to prevent cross-occlusion between individuals, allowing for independent capture of each person's skeletal data and avoiding data confusion caused by overlapping individuals or mismatches. Because each individual remains independent, the system can analyze motion data from different areas in parallel without additional individual separation steps, thus reducing computation and increasing processing speed. This method effectively reduces interference during simultaneous multi-person detection, ensuring the stability and accuracy of skeletal point matching, while optimizing data flow processing efficiency and improving the overall accuracy and real-time performance of the evaluation scheme. Through these processes, real-time accurate detection of skeletal points from multiple angles for multiple individuals can be achieved and input into a motion ability assessment model for motion coordination evaluation.
[0136] In this embodiment of the invention, by performing multi-view data association on the detection results of skeletal key points and aligning the image data acquired by the dual cameras using a unified host timestamp, the consistency of the two video streams in the temporal dimension is ensured, avoiding matching errors caused by time asynchrony. This lays the foundation for accurate association of skeletal point data in the future, which is crucial for analyzing children's movement patterns over time. Based on the strict mapping relationship between the pre-calibrated physical space of the venue and the camera images, matching is performed using spatial position features, which can accurately match the skeletal point data of the same person from different perspectives. This effectively reduces mismatches caused by differences in perspective or scene complexity, improves the accuracy of data association, and thus ensures the correct assessment of children's movement status. Integrating skeletal point data from the frontal and side perspectives can comprehensively and three-dimensionally reflect children's movement in three-dimensional space, providing richer and more complete information for assessing key indicators such as coordination, continuity, and symmetry of movements. This helps to discover potential movement problems, such as asymmetry of movements and abnormal postures, and provides a more sufficient basis for developing personalized sports training programs.
[0137] This invention also provides a method for quantitatively analyzing the test results of multiple individuals using a sports ability assessment model. The method specifically includes:
[0138] Step S601: Load the results of multiple tests, evaluate the coordination of the results of multiple tests based on the motor ability assessment model, and obtain the children's health assessment score corresponding to the health standard.
[0139] It should be noted that the sports ability assessment model in this embodiment includes a feature extraction module and a regression scoring module. The feature extraction module includes a graph construction layer, a graph convolutional network (GCN), a recurrent neural network (RNN), an average pooling layer, and a feature fusion layer connected in sequence. The regression scoring module is used to output the final health assessment score. The regression scoring module includes a fully connected layer 1, a fully connected layer 2, and an output layer connected in sequence. The fully connected layer 1 is used to obtain the feature vector output by the feature fusion layer. Data is passed between the fully connected layer 1 and the fully connected layer 2 through the ReLU activation function. Data is also passed between the fully connected layer 2 and the output layer through the ReLU activation function. The output layer outputs the scoring vector through the Sigmoid activation function.
[0140] In this embodiment of the invention, a method for evaluating the coordination of multiple-person test results based on a motion ability assessment model is also provided. Figure 4 A flowchart illustrating a method for evaluating coordination in multi-person testing based on a motor ability assessment model is provided. The method includes:
[0141] Step S6011: Obtain multi-person detection results, and identify the skeletal pixel coordinates of the frontal and side views, the difference in skeletal pixel coordinates between multiple frames, and the spatial mapping under the frontal and side views in the multi-person detection results. Construct graphs for the skeletal keypoint detection results of the frontal and side views respectively, and regard the skeletal keypoint detection results of the frontal and side views as a skeletal point graph topology. Among the nodes E represents the edge between corresponding bone points. The weight of the edge is defined based on the Euclidean distance between the skeletal points;
[0142] Step S6012: Use Graph Convolutional Network (GCN) to perform convolution operations on the topology of the skeleton point graph for each viewpoint, and integrate the difference value ΔX into the graph convolution operation to enhance the model's sensitivity to action changes and extract local structural features.
[0143] Step S6013: Input the feature sequence after graph convolution into the recurrent neural network (RNN). The RNN captures the time series features and performs average pooling on the features at each time step to obtain the feature representation at each time step. The outputs of the RNN from the two perspectives are fused through a fusion layer to obtain the final feature vector.
[0144] Specifically, when constructing graphs from the skeletal keypoint detection results of the frontal and side views, the skeletal pixel coordinates of the frontal and side views, the differences in skeletal pixel coordinates between multiple frames, and the spatial mappings under the two views are used as inputs, and their mathematical representation is as follows:
[0145]
[0146] Where R represents a multidimensional array over the real number field, T is the length of the time series, 2 represents two different perspectives, and P i (t) represents the skeleton point number in frame t, C i (t) represents the coordinates of the corresponding skeleton point, ΔX represents the coordinate difference obtained by calculating the coordinate difference between adjacent frames, and M is a weight matrix that represents the spatial mapping relationship between two viewpoints;
[0147] For each viewpoint (frontal viewpoint, side viewpoint), the sequence of skeletal points is considered as a skeletal point graph topology. Among the nodes E represents the edge between corresponding bone points. The weight of the edge is defined based on the Euclidean distance between the skeletal points;
[0148] The topology of the skeletal point map for each viewpoint can be represented as:
[0149]
[0150] Among them, A i,j This represents the connection weight between skeletal point i and skeletal point j. A Graph Convolutional Network (GCN) is used to convolve the skeletal point map for each viewpoint, and the difference value ΔX is integrated into the graph convolution operation to enhance the model's sensitivity to action changes and extract local structural features. The graph convolution operation can be represented as:
[0151]
[0152] in, D is the node feature matrix of the l-th layer. (l) This represents the dimension of the feature vector of the l-th layer node in the graph convolutional network. The initial dimension was designed to be 64, and it was adjusted based on the model's performance during training. It is a degree matrix. W represents a self-join of nodes. (l) Here is the weight matrix of the l-th layer, and σ is the non-linear activation function ReLU. After L layers of graph convolution, the output features for each viewpoint are:
[0153]
[0154] The feature sequence after graph convolution is input into a recurrent neural network (RNN) to capture time-series features. For each viewpoint, the input to the RNN is the feature sequence after graph convolution:
[0155]
[0156] To simplify, we can perform average pooling on the features at each time step to obtain the feature representation for each time step:
[0157]
[0158] The output of the RNN is represented as:
[0159]
[0160] The outputs of the two RNNs from different perspectives are fused through a fusion layer to obtain the final feature vector. A weighted summation method is used for the fusion.
[0161]
[0162] Where M is the viewpoint weight matrix, and · denotes matrix multiplication. Finally, the dimension of the fused feature vector is:
[0163] H fusion ∈R T×H
[0164] In step S6014, the feature vector is input into two fully connected layers and the output layer. The score of each parameter is obtained through the Sigmoid activation function, and the scores of each parameter are weighted and summed to output the health assessment score.
[0165] In this embodiment of the invention, during the regression scoring stage, to calculate the scores of each parameter of the health standard, the feature vector is input into two fully connected (dense) layers and an output layer. Assume the health standard has K parameters, and the score range for each parameter is [0, 100]. The mathematical representation of the fully connected layer is as follows:
[0166] Z1 = ReLU(W1H) fusion +b1),
[0167] Z2 = ReLU(W2Z1 + b2),
[0168] in, and It is a weight matrix. and It is a bias vector. and This is the output value. Then, the Sigmoid activation function is used to obtain a score for each parameter, mathematically represented as follows:
[0169] S=σ(W3Z2+b3),
[0170] in, It is a weight matrix, b1∈R K It is a bias vector, S∈R T×KThis is the output score vector, and σ is the sigmoid activation function. To obtain a total score, we perform a weighted summation of the scores for each parameter:
[0171]
[0172] Among them, w k It is the weight of the k-th parameter, S k This represents the score of the k-th parameter. Subsequently, the mean squared error (MSE) is used as the loss function for the regression task to calculate the difference between the predicted and actual scores:
[0173]
[0174] Among them, S k It is a predicted score. These are actual scores. Using the Adam optimizer, the learning rate setting is adjusted based on the actual situation:
[0175]
[0176] Where, θ t α represents the model parameters for the t-th iteration, and α is the learning rate.
[0177] It should be noted that the motor ability assessment model analyzes key indicators such as a child's movement trajectory, stability, and range of motion to determine whether their motor ability meets health standards and to promptly detect potential motor control problems or neurodevelopmental delays, especially subtle developmental abnormalities that are difficult to detect through routine observation.
[0178] In this embodiment of the invention, the training process of the motion ability assessment model involves feature extraction and scoring stages. First, image data is converted into skeletal point pixel coordinates, differences between multiple frames, and spatial mapping relationships between viewpoints. A skeletal point map is constructed for each viewpoint, and a Generic Network (GCN) is applied to extract local structural features. Subsequently, a Recurrent Neural Network (RNN) captures time-series features, and then feature fusion is used to form a feature vector. In the regression scoring stage, the feature vector is input into a fully connected layer, using MSE as the loss function, and the Adam optimizer is used to update parameters and calculate the scores for various parameters of the health standard.
[0179] For example, in the lateral gliding test, the motor ability assessment model first analyzes the child's movement trajectory and quantifies key motor postures. The test requires the child to face sideways to the gliding direction and be assessed from both frontal and side views. The standard of "shoulder parallel to the ground" is calculated based on skeletal point data from the frontal view. The motor ability assessment model extracts the coordinates of the left shoulder (skeletal point number 6) and right shoulder (skeletal point number 7) skeletal points and calculates their angles relative to the horizontal axis. If the angle is less than 5°, the shoulder is considered basically horizontal; otherwise, it is considered a posture deviation. Furthermore, the side view of the motor ability assessment model can further analyze the consistency between the foot movement direction and the shoulder direction to determine whether the child maintains a normal posture during lateral movement. This assessment method based on the spatial relationship of skeletal points transforms subjective movement standards into measurable geometric parameters, ensuring the objectivity and repeatability of the assessment results.
[0180] The motor ability assessment model also detects the presence of a moment of takeoff. Based on skeletal point data, the system tracks the vertical coordinate changes of both ankles (skeleton point numbers 16 and 17). If, during the glide, both ankle coordinates are simultaneously higher than a fixed threshold from the previous frame, and there are no skeletal points below the reference horizontal line of the ground contact point, the movement is considered to meet the requirements. If at least one foot (16 or 17) is always in contact with the ground, it is considered incomplete. Furthermore, when the motor ability assessment model detects that the child has completed a double-foot landing-double-foot takeoff movement, it considers the child to have completed a set of movement cycles. The motor ability assessment model counts and statistically analyzes whether the child is capable of completing the corresponding task according to the assessment requirements (e.g., completing four complete cycles and performing a reverse glide).
[0181] Through these refined assessment criteria, the motor ability assessment model can not only accurately determine whether a child has completed a movement, but also analyze the details of their motor ability and provide targeted improvement suggestions. This process ensures the accuracy and comprehensiveness of the assessment, providing parents and professionals with a scientific basis for evaluating motor coordination ability.
[0182] Step S602: The athletic ability assessment model analyzes the health assessment score based on health standards, determines whether the health assessment score is qualified, and obtains the health assessment result.
[0183] It should be noted that the health assessment score is calculated based on the child's motor ability data and the established health standards. The calculation formula takes into account the weight parameters of each indicator feature in each sport, and comprehensively evaluates key indicators such as movement stability, coordination, and range of motion to determine whether the movement is completed. If the movement is judged to be completed, points are awarded.
[0184] During the evaluation process, each indicator in each test item is independently judged as "qualified" or "unqualified". If an indicator is not completed, no points will be awarded for that item, which will affect the final evaluation result.
[0185] Qualified (S≥70): The athletic ability meets the health standards, all assessment indicators meet the requirements, and the performance in terms of coordination and stability is normal or has only slight deviations.
[0186] Unsatisfactory (S < 70): The motor ability does not meet the standard, there are obvious abnormal movements or coordination problems, and targeted training and intervention are recommended.
[0187] Step S603: The athletic ability assessment model combines the health assessment results to generate a personalized health assessment report;
[0188] The personalized health assessment report includes: assessment indicator analysis, health score and recommendations, and training suggestions.
[0189] It should be noted that the assessment indicator analysis provides the completion status of each movement, analyzes the indicators of each movement in detail, and gives corresponding conclusions to help parents and educators identify potential problems in movement.
[0190] Health Score and Recommendations: Based on the assessment score, the system provides an evaluation of the child's current athletic ability level. Combined with specific data, it offers personalized training suggestions and improvement directions to help children optimize their athletic performance and promote healthy development. For example, in the lateral gliding test, the system detected that a child's shoulders failed to remain level during the gliding motion. The calculation showed that the angle between the left shoulder (skeletal point number 6) and the right shoulder (skeletal point number 7) was 8°, exceeding the set threshold of 5°. Simultaneously, the angle between the step direction and the shoulder direction was 12°, higher than the coordination standard of 10°, indicating that the child exhibited compensatory body twisting movements during the gliding motion. When a child's movements are severely uncoordinated, it may indicate a neurodevelopmental disorder.
[0191] Based on the child's assessment results, the motor ability assessment model will provide personalized improvement suggestions in the assessment report:
[0192] Training suggestions:
[0193] Sensory integration training, as well as cognitive ability training and neuromotor task training in daily life, can be helpful. Secondly, psychological therapy can also help improve the child's self-awareness, such as proper guidance from parents.
[0194] In this embodiment of the invention, the motor ability assessment model can monitor children's dynamic movements (such as jumping, hopping on one foot, gliding sideways, etc.) in real time, and provide accurate assessments based on complete spatiotemporal feature data, generating personalized motor ability reports and improvement suggestions.
[0195] On the other hand, this invention also provides a multi-perspective assessment system for children's motor coordination ability based on deep learning. Figure 3 The diagram shows the structure of a deep learning-based multi-view children's motor coordination ability assessment system, which includes:
[0196] The model building module 100 is used to build a dedicated children's sports dataset. Based on the children's sports dataset, a sports evaluation dataset is extracted. The children's sports dataset and the sports evaluation dataset are integrated into the modeling dataset. The modeling dataset is used to train the skeletal point detection model and the sports ability evaluation model, and the converged skeletal point detection model and the sports ability evaluation model are output.
[0197] The data acquisition module 200 acquires real-time video streams of people entering the site from the acquisition camera, preprocesses the real-time video streams, and outputs a preprocessed dataset.
[0198] The visual recognition module 300 is used to load the preprocessed dataset. The skeleton point detection model is based on a combination of deep convolutional neural network and regression method to identify and analyze the preprocessed dataset, identify the skeleton key points in the preprocessed dataset, and obtain the skeleton key point detection results.
[0199] The motion state detection module 400 acquires the skeletal key point detection results, performs multi-view data association on the skeletal key point detection results, associates the person ID with the corresponding skeletal key points, and obtains the person matching results containing the person ID and skeletal key points. Based on the motion region division method, it performs multi-person detection on the person matching results and obtains multi-person detection results.
[0200] The personalized assessment module 500 takes the test results of multiple people as input, executes the exercise ability assessment model, and performs quantitative analysis on the test results of multiple people to generate a personalized health assessment report.
[0201] It should be noted that the model building module 100, data acquisition module 200, visual recognition module 300, motion state detection module 400, and personalized assessment module 500 in the deep learning-based multi-perspective children's motor coordination ability assessment system correspond to the deep learning-based multi-perspective children's motor coordination ability assessment method described above. The explanations, examples, and beneficial effects of the relevant content can be found in the corresponding content of the deep learning-based multi-perspective children's motor coordination ability assessment method, and will not be repeated here.
[0202] In summary, this invention provides a multi-view children's motor coordination ability assessment system and method based on deep learning. This invention uses a skeletal point detection model combined with deep convolutional neural networks and regression methods to identify and analyze preprocessed datasets, and performs multi-person detection based on motion region segmentation methods to match personnel. It combines visual recognition and motion state detection technologies, employs a dual-camera layout with frontal and side views, and utilizes deep learning algorithms to achieve automated assessment of children's dynamic movements. Compared to manual measurement or single-view analysis, it can completely capture the movement process, ensure consistency of assessment standards, eliminate human error, and improve detection accuracy and efficiency. It also supports simultaneous detection of multiple individuals. By utilizing the field-of-view partitioning technology of the acquisition cameras, each child entering the venue is assigned a unique identification number and their movement trajectory is tracked independently, ensuring that the movement data of different individuals are not mixed up, thus achieving efficient multi-target movement analysis.
[0203] It should be noted that, for the sake of simplicity, the foregoing embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to the present invention. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0204] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit the scope of protection of the invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on these embodiments, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art can still combine, add, delete, or otherwise adjust the features of the various embodiments of the present invention according to the circumstances without conflict or creative effort, thereby obtaining different technical solutions that do not fundamentally depart from the concept of the present invention. These technical solutions also fall within the scope of protection of the present invention.
Claims
1. A multi-perspective assessment method for children's motor coordination ability based on deep learning, characterized in that, The method includes: A children's movement dataset is constructed, which includes skeletal point labeling results. A movement assessment dataset is extracted based on the children's movement dataset, wherein the movement assessment dataset is based on abnormal movement data of children with motor coordination deviations. The children's movement dataset and the movement assessment dataset are integrated into a modeling dataset. A skeletal point detection model and a movement ability assessment model are trained using the modeling dataset, and a converged skeletal point detection model and a movement ability assessment model are output. Based on the real-time video stream of people entering the venue captured by the acquisition camera, the real-time video stream is preprocessed and a preprocessed dataset is output. Load the preprocessed dataset, and the skeleton point detection model uses a combination of deep convolutional neural network and regression method to identify and analyze the preprocessed dataset, identify the skeleton key points in the preprocessed dataset, and obtain the skeleton key point detection results; Obtain the skeletal keypoint detection results, perform multi-view data association on the skeletal keypoint detection results, associate the person ID with the corresponding skeletal keypoint, obtain the person matching results containing the person ID and skeletal keypoint, and perform multi-person detection on the person matching results based on the motion region segmentation method to obtain multi-person detection results; Using the test results of multiple people as input, the exercise ability assessment model is executed. The exercise ability assessment model performs quantitative analysis on the test results of multiple people and generates a personalized health assessment report. The process involves using the results of multiple tests as input, executing a fitness assessment model, and then quantitatively analyzing the results to generate a personalized health assessment report. Specifically, this includes: Load the test results of multiple people, evaluate the coordination of the test results based on the motor ability assessment model, and obtain the children's health assessment score corresponding to the health standard; The athletic ability assessment model analyzes health assessment scores based on health standards, determines whether the health assessment scores are qualified, and obtains the health assessment results. The athletic ability assessment model, combined with health assessment results, generates a personalized health assessment report. The personalized health assessment report includes: assessment indicator analysis, health score and recommendations, and training recommendations; The coordination evaluation of multiple-person test results based on the motion ability assessment model specifically includes: Obtain multi-person detection results and identify the skeletal pixel coordinates of the frontal and side views, the differences in skeletal pixel coordinates between multiple frames, and the spatial mapping under the frontal and side views. Construct graphs for the skeletal keypoint detection results of the frontal and side views respectively, treating the skeletal keypoint detection results of the frontal and side views as a skeletal point graph topology. , where nodes E represents the edge between corresponding skeletal points ( The weight of the edge is defined based on the Euclidean distance between the skeletal points; Specifically, when constructing graphs from the skeletal keypoint detection results of the frontal and side views, the skeletal pixel coordinates of the frontal and side views, the differences in skeletal pixel coordinates between multiple frames, and the spatial mappings under the two views are used as inputs, and their mathematical representation is as follows: , , , in, Represents a multidimensional array over the real number field. This represents the length of the time series, and 2 indicates two different perspectives. This represents the skeleton point number in frame t. This represents the coordinates of the corresponding skeletal point. This is indicated by calculating the coordinate difference between adjacent frames. It is a weight matrix that represents the spatial mapping relationship between two perspectives; The topological representation of the skeletal point map for each viewpoint is as follows: , in, This represents the connection weight between bone point i and bone point j; A Graph Convolutional Network (GCN) is used to convolve the skeletal point graph topology for each viewpoint, and the difference values are then processed. Integrating this into graph convolution operations enhances the model's sensitivity to action changes and extracts local structural features. Graph convolution operations are represented as follows: , in, It is the feature vector of the node in the l-th layer of the graph convolutional network. This represents the dimension of the feature vector of the l-th layer node in the graph convolutional network. The initial dimension was designed to be 64, and it was adjusted based on the model's performance during training. It is a degree matrix. Indicates a self-connection of nodes. Here is the weight matrix of the l-th layer, and σ is the non-linear activation function ReLU. After L layers of graph convolution, the output features of each viewpoint are: , The feature sequence after graph convolution is input into a recurrent neural network (RNN). The RNN captures the time series features and performs average pooling on the features at each time step to obtain the feature representation at each time step. The outputs of the two RNNs from different perspectives are fused through a fusion layer to obtain the final feature vector. For each viewpoint, the input to the recurrent neural network (RNN) is the feature sequence after graph convolution: , Average pooling is performed on the features at each time step to obtain the feature representation at each time step: , The output of the RNN is represented as: , The outputs of the two RNNs from different perspectives are fused through a fusion layer to obtain the final feature vector, where a weighted summation method is used for fusion. , in, It is the viewpoint weight matrix. This represents matrix multiplication. Ultimately, the dimension of the fused feature vector is: , The fused feature vector is input into two fully connected layers and an output layer connected in sequence. The score of each parameter is obtained by passing the Sigmoid activation function, and the scores of each parameter are weighted and summed to output the health assessment score.
2. The multi-view assessment method for children's motor coordination ability based on deep learning as described in claim 1, characterized in that: The construction of the children's sports dataset specifically includes: The system uses multi-angle, full-body shooting to collect specialized video data of children in different scenarios. The specialized video data covers common movements such as running, jumping, and side-sliding, as well as children's natural movements during play or daily activities. Load dedicated video data, parse and process the dedicated video data frame by frame to obtain at least one set of dedicated images, process the children's skeletal point markings in the dedicated images to obtain a children's motion dataset containing the skeletal point marking results. In the process of processing the children's skeletal point markings in the dedicated images, 17 skeletal points of the children are assigned fixed numbers, namely:
1. nose, 2. left eye, 3. right eye, 4. left ear, 5. right ear, 6. left shoulder, 7. right shoulder, 8. left elbow, 9. right elbow, 10. left wrist, 11. right wrist, 12. left hip joint, 13. right hip joint, 14. left knee, 15. right knee, 16. left ankle, 17. right ankle.
3. The multi-view assessment method for children's motor coordination ability based on deep learning as described in claim 2, characterized in that: The method for training a skeletal point detection model and a motion capability assessment model using a modeling dataset includes: Load the modeling dataset and divide it into a training set and a validation set. The training set is used to train the skeleton point detection model and the motion ability assessment model, while the validation set is used to verify the generalization ability of the skeleton point detection model and the motion ability assessment model. The ratio of the training set to the validation set is 4:
1. The skeletal point detection model is iteratively trained using a training set and a validation set to output a converged skeletal point detection model. The training set includes skeletal point annotation data and movement data of normal and abnormal children. The training set is used to train the skeletal point detection model and the movement ability assessment model, so that the skeletal point detection model can identify the joint positions of children in different movements and distinguish between normal and abnormal movement patterns. Load the skeleton point detection model. The skeleton point detection model combines data from the frontal and side views to stably identify the skeleton points of children under different views and output the skeleton point recognition results of the skeleton point detection model. Obtain the skeleton point recognition results, use the skeleton point recognition results to train and optimize the motion ability assessment model, and output the motion ability assessment model based on multi-view sequence skeleton point image data. The athletic ability assessment model includes a feature extraction module and a regression scoring module. The feature extraction module consists of a graph construction layer, a graph convolutional network (GCN), a recurrent neural network (RNN), an average pooling layer, and a feature fusion layer, all connected in sequence. The regression scoring module is used to output the final health assessment score. It consists of a fully connected layer 1, a fully connected layer 2, and an output layer, all connected in sequence. The fully connected layer 1 is used to obtain the feature vector output by the feature fusion layer. Data is passed between fully connected layer 1 and fully connected layer 2 through the ReLU activation function. Data is also passed between fully connected layer 2 and the output layer through the ReLU activation function. The output layer outputs the scoring vector through the Sigmoid activation function.
4. The multi-view assessment method for children's motor coordination ability based on deep learning as described in claim 1, characterized in that: The process involves preprocessing the real-time video streams of people entering the venue captured by cameras and outputting a preprocessed dataset, specifically including: The camera captures video streams in real time at a fixed frame rate of 60 FPS to obtain a real-time video stream; Load the real-time video stream, parse and process the real-time video stream frame by frame, and number and store the real-time images processed frame by frame in chronological order. The real-time images, after being processed according to time sequence, are converted to a new format, grayscale normalization is performed, and a smoothing filter is used to remove noise from the images, resulting in a preprocessed dataset.
5. The multi-view assessment method for children's motor coordination ability based on deep learning as described in claim 4, characterized in that: The preprocessed dataset is loaded, and the skeletal point detection model, based on a combination of deep convolutional neural networks and regression methods, identifies and analyzes the preprocessed dataset to identify key skeletal points and obtain skeletal key point detection results, specifically including: Load the preprocessed dataset, and the skeleton point detection model extracts key features from the preprocessed dataset through convolutional layers, identifies multiple candidate locations of skeleton points in the preprocessed dataset, and performs preliminary localization of the skeleton points. The preliminary location results of skeletal points are obtained. Then, the preliminary location results of skeletal points are classified and accurately categorized according to human body structure by using clustering algorithms combined with regression analysis to obtain the detection results of skeletal key points. When classifying and precisely categorizing the preliminary skeletal point localization results according to human anatomy, each class is treated as a target, and the set of skeletal point numbers for each target is output first: , in The skeletal point numbers are as follows:
1. Nose, 2. Left eye, 3. Right eye, 4. Left ear, 5. Right ear, 6. Left shoulder, 7. Right shoulder, 8. Left elbow, 9. Right elbow, 10. Left wrist, 11. Right wrist, 12. Left hip joint, 13. Right hip joint, 14. Left knee, 15. Right knee, 16. Left ankle, 17. Right ankle; Each number corresponds to a two-dimensional coordinate, representing the position of that skeletal point in the image: , in Skeletal points Coordinates on the image.
6. The multi-view assessment method for children's motor coordination ability based on deep learning as described in claim 5, characterized in that: The process of performing multi-view data association on the skeletal keypoint detection results, associating the person ID with the corresponding skeletal keypoints, and obtaining person matching results containing the person ID and skeletal keypoints specifically includes: Load the skeletal keypoint detection results, identify the acquisition cameras associated with the images in the skeletal keypoint detection results, align the timestamps of the acquired videos from the acquisition cameras using the host timestamp as a unified standard, and output image data with synchronized timestamps: ; A strict mapping relationship between the actual physical space of the site and the captured camera images is pre-calibrated. Based on this pre-calibrated mapping relationship, multi-view data association is performed on the skeletal keypoint detection results to obtain personnel matching results containing person IDs and skeletal keypoints. The mapping relationship is established through precise camera installation positions, measurements, and viewpoint calibration experiments. When the target is located in the upper left corner of the site, matching is achieved through the following specific steps based on the pre-calibrated mapping relationship: First, all target skeletal point data appearing in the right area of the frontal view are detected in real time, and target skeletal point data in the left area of the side view are also detected. Then, spatial consistency verification is performed to check whether the relative positional relationship of the two sets of skeletal points conforms to the pre-calibrated spatial mapping rules. Then, motion state verification is performed to verify whether the motion trends reflected by the two sets of skeletal points are consistent. When all verifications pass, the same ID number is assigned to the two sets of skeletal point data, i.e.: , The ID number is combined with each person's skeletal point data, and the ID number represents the correspondence between the ID number and the skeletal point data: 。 7. A deep learning-based multi-view children's motor coordination ability assessment system, the system being used to implement the deep learning-based multi-view children's motor coordination ability assessment method as described in any one of claims 1-6, characterized in that: The evaluation system includes: The model building module is used to construct a children's motion dataset, which includes skeletal point labeling results. Based on the children's motion dataset, a motion assessment dataset is extracted. The motion assessment dataset is based on abnormal movement data of children with motor coordination deviations. The children's motion dataset and the motion assessment dataset are integrated into a modeling dataset. The modeling dataset is used to train a skeletal point detection model and a motion ability assessment model, and outputs a converged skeletal point detection model and a motion ability assessment model. The data acquisition module collects real-time video streams of people entering the venue from the acquisition cameras, preprocesses the real-time video streams, and outputs a preprocessed dataset. The visual recognition module is used to load the preprocessed dataset. The skeleton point detection model is based on a combination of deep convolutional neural networks and regression methods to identify and analyze the preprocessed dataset, identify the key skeleton points in the preprocessed dataset, and obtain the key skeleton point detection results. The motion state detection module acquires the skeletal key point detection results, performs multi-view data association on the skeletal key point detection results, associates the person ID with the corresponding skeletal key points, and obtains the person matching results containing the person ID and skeletal key points. Based on the motion region division method, the person matching results are used to perform multi-person detection and obtain multi-person detection results. The personalized assessment module takes the test results of multiple people as input, executes the exercise ability assessment model, and performs quantitative analysis on the test results of multiple people to generate a personalized health assessment report.
Citation Information
Patent Citations
Posture health monitoring method and system based on human body key point detection
CN117457193A
Motion estimation method and device based on multi-modal perception and electronic equipment
CN119533462A
Infant physical fitness action standard identification and evaluation method based on convolutional neural network
CN118570871A
Rehabilitation training action evaluation method and device based on multi-view vision
CN120183042A