A classification system, classification method, device, and storage medium applicable to the cerebral palsy scenario
The classification system addresses the high technical demand and time-consuming nature of brain palsy assessments by converting two-dimensional keypoint information to three-dimensional data for improved classification using pre-trained models, enhancing efficiency and accuracy.
Patent Information
- Application Number
- CN202510000914.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-01-02
AI Technical Summary
The existing technology has high technical requirements and takes a long time to evaluate cerebral palsy, making it difficult to promote and apply it at the grassroots level.
By obtaining the video data of the target object, extracting two-dimensional key point information and converting it into three-dimensional key point information, the pre-trained classification model is used for classification processing, and the evaluation efficiency is improved.
The time of the cerebral palsy detection process is shortened and the accuracy and efficiency of classification is improved.
Smart Images

Figure CN119399797B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical technology, and in particular, to a classification system, a classification method, a device, and a storage medium applicable to the cerebral palsy scenario. Background Art
[0002] Research in the field of cerebral palsy is an important branch of the field of medical research. High-quality evidence from the "Chinese Rehabilitation Guidelines for Cerebral Palsy (2022)" shows that abnormal general movements (GMs) assessment or the scoring trajectory of the Hammersmith Infant Neurological Examination (HINE), combined with abnormal MRI examination, is more accurate than a single clinical assessment. It can combine assessments with strong predictive validity and clinical inferences for early diagnosis before the corrected age of 6 months.
[0003] Although the GMs or HINE assessment techniques have extremely high application value in the early identification of cerebral palsy and the prediction of severe developmental delays, due to the high technical requirements for medical staff and the long time consumption, there are obvious limitations in large-scale promotion, especially in grass-roots promotion and application. Summary of the Invention
[0004] The present invention provides a classification system, a classification method, a device, and a storage medium applicable to the cerebral palsy scenario to solve the problems of high technical requirements and long time consumption in the cerebral palsy assessment process.
[0005] According to one aspect of the present invention, a classification system applicable to the cerebral palsy scenario is provided. The classification system includes a data processing module and a classification module, wherein:
[0006] The data processing module is configured to obtain video data of a target object, where the video data includes a plurality of video frames; extract two-dimensional key point information of the target object in each video frame, and convert the two-dimensional key point information into three-dimensional key point information;
[0007] The classification module is configured to perform classification processing on the three-dimensional key point information corresponding to each video frame based on a pre-trained classification model to obtain a classification result of the target object in the cerebral palsy scenario. The accuracy of classifying the target object in the cerebral palsy scenario is improved.
[0008] According to another aspect of the present invention, a classification method applicable to the cerebral palsy scenario is provided. The classification method includes:
[0009] Obtain video data of a target object, where the video data includes a plurality of video frames; extract two-dimensional key point information of the target object in each video frame, and convert the two-dimensional key point information into three-dimensional key point information;
[0010] Based on a pre-trained classification model, classify the three-dimensional key point information corresponding to each video frame to obtain the classification result of the target object in the cerebral palsy scenario.
[0011] According to another aspect of the present invention, there is provided an electronic device, which includes:
[0012] At least one processor; and
[0013] A memory communicatively connected to the at least one processor; wherein,
[0014] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the classification method applicable to the cerebral palsy scenario according to any embodiment of the present invention.
[0015] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the classification method applicable to the cerebral palsy scenario according to any embodiment of the present invention when executed.
[0016] The technical solution of the embodiment of the present invention solves the problems of high technical requirements and long time consumption in the cerebral palsy evaluation process by acquiring the video data of the target object, extracting the two-dimensional key point information of the target object in each video frame, converting the two-dimensional key point information into three-dimensional key point information, and classifying the three-dimensional key point information corresponding to each video frame based on a pre-trained classification model to obtain the classification result of the target object in the cerebral palsy scenario, shortens the time required for the target object in the cerebral palsy detection process, and improves the classification efficiency of the target object in the cerebral palsy scenario.
[0017] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0019] Figure 1 is a schematic structural diagram of a classification system applicable to the cerebral palsy scenario provided in Embodiment 1 of the present invention;
[0020] Figure 2It is a schematic structural diagram of a classification model provided by an embodiment of the present invention;
[0021] Figure 3 It is a flowchart of a classification method applicable to the cerebral palsy scenario provided by the second embodiment of the present invention;
[0022] Figure 4 It is a schematic structural diagram of an electronic device provided by the third embodiment of the present invention. Detailed implementation manners
[0023] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0025] Embodiment 1
[0026] Figure 1 It is a schematic structural diagram of a classification system applicable to the cerebral palsy scenario provided by the first embodiment of the present invention. This embodiment is applicable to collecting video data of a target object and classifying the target object in the cerebral palsy scenario through the video data. The classification system applicable to the cerebral palsy scenario can be integrated in an electronic device, and the electronic device can include electronic devices such as a mobile terminal, a server, or a computer. The mobile terminal can include a mobile phone and a tablet computer, etc.
[0027] Such as Figure 1As shown in the figure, the classification system includes a data processing module 110 and a classification module 120. Among them, the data processing module 110 is used to obtain the video data of the target object, and the video data includes multiple video frames; extract the two-dimensional key point information of the target object in each video frame and convert the two-dimensional key point information into three-dimensional key point information; the classification module 120 is used to classify the three-dimensional key point information corresponding to each video frame based on a pre-trained classification model to obtain the classification result of the target object in the cerebral palsy scenario.
[0028] In this embodiment, the target object can be an object that needs to be classified in the cerebral palsy scenario, such as an infant, a toddler, or a high-risk infant, etc. Among them, a high-risk infant refers to a special population that has various risk factors unfavorable to the growth and development of the fetus and infant during the mother's pregnancy and childbirth period, neonatal period, and infant and toddler period. High-risk infants often suffer from brain damage due to various reasons, resulting in the impairment of the development of one or more functional areas such as movement, and even sequelae such as mental retardation and cerebral palsy in the long term.
[0029] In the process of classifying the target object in the cerebral palsy scenario, classification can be performed by analyzing the video data of the target object. The video data of the target object can be composed of multiple video frames. Among them, a video frame can be a static image that constitutes the video of the target object. The video data can be information of the target object recorded in the form of a video during a preset time period, and can be collected by a video acquisition device. The preset time period can be a time period set in advance for video data acquisition. For example, the preset time period can be 10 minutes. The video acquisition device can be various terminal intelligent devices, etc. Among them, the terminal intelligent devices include mobile phones, tablet computers, and PC machines, etc. It can be understood that the video acquisition device and the electronic device in this embodiment can be the same device. Taking the electronic device as a mobile phone as an example, the video data of the target object can be collected through the camera function of the mobile phone, and the collected video data is processed by the data processing module 110 configured in the mobile phone. The video acquisition device and the electronic device in this embodiment can be different devices. For example, the video acquisition device is a camera or a mobile terminal, and the electronic device is a computer; for example, the video acquisition device can be a client device, and the electronic device is a doctor's device. Among them, the client device collects video data and transmits it to the doctor's device, and the doctor's device receives the video data transmitted by the video acquisition device. Among them, the client device can be a camera or a mobile terminal, etc., and the doctor's device can be a mobile terminal, a server, or a computer, etc.
[0030] The video data of the target object can also be retrieved from the database. The database stores the video data of multiple objects, and the video data of the target object is matched in the database through the identifier of the target object. The target object identifier can be the identity information of the target object or other unique identifiers, etc.
[0031] The key points can be the main joints or main parts of the target object required in the classification process of the cerebral palsy scenario. Optionally, the key points include one or more of the nose, left and right eyes, left and right ears, left and right shoulder joints, left and right joints, left and right wrist joints, left and right hip joints, and left and right knee joints and left and right ankle joints. The two-dimensional key point information can be the two-dimensional coordinates of the key points, and the three-dimensional key point information can be the three-dimensional coordinates of the key points. Optionally, the two-dimensional key point information can be the two-dimensional coordinates of the key points in the same coordinate system. The classification result can be used to determine whether the target object has cerebral palsy. For example, the classification result can be that the target object does not have cerebral palsy.
[0032] The two-dimensional key point information can be obtained through a machine learning model. For example, it can be an object detection model. Optionally, the machine learning model for obtaining the two-dimensional key point information can be configured in the data processing module.
[0033] Optionally, the data processing module 110 is further configured to perform key point recognition on multiple video frames respectively through the object detection model to obtain the two-dimensional key point information of the target object in each video frame.
[0034] Specifically, the object detection model can be used to detect the position of the target object. For example, it can be a convolutional neural network model or a Transformer model, which can be selected according to the user's needs and is not limited here. Optionally, the object detection model can be a Detectron2 model. The Detectron2 model is a deep learning framework that can be used for object detection. In the classification of the cerebral palsy scenario, it can perform key point recognition on multiple video frames to determine the two-dimensional key point information of the target object in each video frame, providing detailed data support for subsequent analysis.
[0035] In the process of obtaining the two-dimensional key point information, there may be abnormal situations of the two-dimensional key point information. For example, in the process of obtaining the two-dimensional key point information, it is impossible to obtain the two-dimensional key point information of some key points, resulting in the situation that there are missing values in the corresponding three-dimensional key point information converted from the two-dimensional key point information of some key points, or the two-dimensional key point information provided is abnormal, resulting in the extracted two-dimensional key point information exceeding the reasonable data range. In order to reduce the impact of abnormal situations, the video frames corresponding to the abnormal two-dimensional key point information can be deleted.
[0036] Optionally, the data processing module 110 is further configured to perform abnormal determination on the two-dimensional key point information corresponding to each video frame, where the abnormal determination includes one or more of missing value determination and data range abnormal determination; in the case where the two-dimensional key point information corresponding to any video frame is abnormal, the video frame is deleted.
[0037] Specifically, the missing value determination can be used to determine whether there is missing two-dimensional key point information, and the data range anomaly determination can be used to determine whether the two-dimensional key point information exceeds a preset range. For example, the preset range can be [-2000, 2000]. It can be understood that the preset range can be determined according to the resolution of the video data.
[0038] For any video frame, if the two-dimensional key point information of any key point in the video frame is missing or out of the data range, determine that the video frame is an abnormal video frame, and delete the abnormal video frame to avoid interference of the abnormal video frame on the classification process.
[0039] During the video acquisition process, the target object is in a free movement state. For example, in the video frame, due to occlusion, the two-dimensional key point information of a certain key point cannot be read, resulting in the missing of the two-dimensional key point information of this key point. Exemplarily, during the rotation of the target object's head, the side face of the target object is captured in the video frame, and only the two-dimensional key point information of the left eye can be captured, while the two-dimensional key point information of the right eye cannot be captured, resulting in the missing of the two-dimensional key point information.
[0040] For example, the preset range of the two-dimensional key point information can be [-2000, 2000]. The two-dimensional key point information of a certain key point is (x, y). When there is at least one data in x or y that is less than -2000 or greater than 2000, it can be determined that the data range of the two-dimensional key point information of this key point is abnormal.
[0041] The resolutions of video data collected by different devices may vary, which may cause coordinate errors to a certain extent. Therefore, before converting the two-dimensional key point information into three-dimensional key point information, the two-dimensional key point information can be normalized.
[0042] Optionally, the data processing module 110 is further configured to normalize the two-dimensional key point information based on the resolution information of the video frame before converting the two-dimensional key point information into three-dimensional key point information, so as to obtain the two-dimensional key point information located in the video frame with the standard resolution.
[0043] Specifically, the resolution information can be used to represent the clarity of the video data, and can include the width and height. Different video data may have different resolution information, such as 1980*1080 and 1280*720, etc. The normalization process can be used to limit the two-dimensional key point information within a certain range, which is helpful for subsequent processing. Taking 17 key points such as the nose, left and right eyes, left and right ears, left and right shoulder joints, left and right joints, left and right wrist joints, left and right hip joints, and left and right knee joints and left and right ankle joints as an example, through the width and height Standardize the two-dimensional key point information, convert the two-dimensional key point information to the range of [-1, 1], and represent the 17 key points in each video frame as , and represent the two-dimensional key point information of each key point as . The standardization formula is:
[0044] .
[0045] Standardize the two-dimensional key point information through the resolution information. When and are both divided by the same width w, it can ensure that the obtained two-dimensional key point information has a consistent ratio, rather than different scalings in different directions, thus ensuring the consistency of the input data. By subtracting [1, h / w], the y coordinate can be adjusted to maintain the correct aspect ratio. This standardization method not only eliminates the influence of the resolution information on the two-dimensional key point information, but also ensures the generalization ability of the data processing module 110 for the video data of different target objects.
[0046] Optionally, the two-dimensional key point information of each key point can also be standardized by setting the standard resolution information and calculating the change ratios of the width and height of the current resolution information to the width and height of the standard resolution information respectively. Among them, the standard resolution information can be the resolution information set in advance. By calculating the first change ratio of the width of the current resolution information to the width of the standard resolution information, and calculating the second change ratio of the height of the current resolution information to the height of the standard resolution information, the width of the current resolution information is converted through the first change ratio, and the height of the current resolution information is converted based on the second change ratio, so as to standardize the two-dimensional key point information. For example, the width and height of the current resolution information are 1980 and 1080 respectively, the width and height of the standard resolution information are 1280 and 720 respectively, the change ratio of the width of the current resolution information is -35%, and the change ratio of the height of the current resolution information is -33%. The "-" indicates a decrease, and the "+" indicates an increase. The and of the two-dimensional key point information of each key point are decreased by 35% and 33% respectively.
[0047] Optionally, the data processing module 110 is further configured to perform dimension conversion processing on the two-dimensional key point information based on a data conversion model to obtain three-dimensional key point information corresponding to the two-dimensional key point information, where the data conversion model includes a plurality of one-dimensional convolutional blocks.
[0048] Specifically, a data conversion model can be a model used for data conversion processing, which can perform dimensional conversion on data. For example, a machine learning model can be used to convert two-dimensional data into three-dimensional data. Optionally, the data conversion model can be a Videopose3D model. The Videopose3D model includes multiple one-dimensional convolutional blocks. Through multi-layer one-dimensional convolution, three-dimensional coordinates are generated: [N,T,C_2] -> [N,T,C_3], where N represents the batch size, T represents the length of the time series, C_2 represents two-dimensional key point information, and C_3 represents three-dimensional key point information. Among them, the one-dimensional convolutional block can be a one-dimensional convolutional block with lengths of 3 and 1 and a stride of 1.
[0049] Since the target object can be a human body, by converting the two-dimensional key point information of the target object into three-dimensional key point information, it is convenient to obtain the spatial features of the key point data during the classification process, providing a more comprehensive feature basis for the classification process and further improving the accuracy of classification.
[0050] Optionally, the classification model includes two feature extraction units and a classification unit. Among them, any one of the feature extraction units includes a long short-term memory processing network block and a dropout layer, and the classification unit includes a fully connected layer and a pooling layer.
[0051] Specifically, the feature extraction unit can be used to extract the features of the input data. Taking 17 key points such as the nose, left and right eyes, left and right ears, left and right shoulder joints, left and right joints, left and right wrist joints, left and right hip joints, left and right knee joints, and left and right ankle joints as an example, the input data can be the three-dimensional key point information of 17 key points extracted from each video frame. The three-dimensional key point information extracted from each video frame is stacked into an array, that is, [N,17,3], where N represents the number of video frames included in the video, 17 is the number of key points, and 3 is the dimension of each key point.
[0052] The feature extraction unit includes a long short-term memory processing network block (Long Short-Term Memory, LSTM) and a dropout layer. Optionally, the long short-term memory processing network block includes 2 layers of long short-term memory networks. The 2 layers of long short-term memory networks can extract features of different depths of the input data, improving the feature extraction ability of the long short-term memory processing network block. The dropout layer is a commonly used regularization technique, which can be used to randomly discard a certain proportion of neurons after the first feature extraction unit extracts the features, so that the classification model does not depend on specific information and retains more features, preventing overfitting of the long short-term memory processing network block. The feature extraction unit adopts the structure of the long short-term memory processing network block and the dropout layer, which can retain more features and prevent overfitting of the long short-term memory network at the same time.
[0053] The classification unit can be used to classify the features extracted by the feature extraction unit. The classification unit includes a fully connected layer and a pooling layer. Among them, the fully connected layer can map the features extracted by the feature extraction unit into a binary classification vector, and the binary classification vector can be used to determine whether the target object is cerebral palsy. The pooling layer can output the result of the binary classification vector. Optionally, the pooling layer can be a max pooling layer.
[0054] Exemplarily, Figure 2 is a schematic structural diagram of a classification model provided by an embodiment of the present invention. The classification model includes two feature extraction units and a classification unit. Each feature extraction unit includes two LSTM blocks and a dropout layer connected in series. The classification unit includes a fully connected layer and a pooling layer.
[0055] In this embodiment, the multi-instance learning method in the three-dimensional coordinate system is applied to the field of gesture recognition. The classification model uses a method of cascading multiple LSTM layers to transmit the output of the previous layer to a re-initialized LSTM module, ensuring that each layer can independently learn different temporal features. Compared with the simple stacking of multiple layers in the traditional LSTM in terms of temporal feature processing, the cascaded LSTM structures in the classification model are in a separated state, which helps to capture dynamic changes at different levels and improve the processing ability for long-time series data. While increasing the network depth, it gradually extracts and refines the temporal information, capturing the patterns in the data from coarse-grained to fine-grained. Making the network more suitable for complex temporal tasks and capturing more temporal features.
[0056] Optionally, the classification system further includes a model training module, which is used to obtain clipped video data. The clipped video data includes marked video frames in the sample video data, and the marked video frames are local video frames in the sample video data; pre-train the classification model to be trained based on the clipped video data and the classification labels corresponding to the clipped video data to obtain a pre-trained classification model; train the pre-trained classification model based on the sample video data and the classification labels corresponding to the sample video data to obtain a trained classification model.
[0057] Specifically, the model training module can be used to train the classification model to be trained. The sample video data can be collected through a video device. For example, the video data of the sample object in a preset time period can be collected through a camera; it can also be obtained by retrieving the video data stored in the database, and the database stores the video data of multiple sample objects.
[0058] The video data includes valid video frames that affect the decision-making for cerebral palsy, as well as invalid video frames that have no impact on the decision-making for cerebral palsy. The invalid video frames cause certain interference to the processing of the classification model. To improve the training efficiency of the classification model, the valid video frames in the video data are marked, and the clipped video data is formed by marking the video frames. For example, doctors or experts can mark the valid video frames in the video data. The marked video frames are local video frames in the sample video data and are also the marked video frames. The clipped video data can be obtained by clipping the video data. For example, the marked video frames in the video data are identified, other video frames outside the marked video frames are deleted, the marked video frames are retained, and the retained marked video frames are clipped according to the timestamps to obtain the clipped video data. This clipped video data can be used as new sample data for pre-training the classification model to be trained.
[0059] Optionally, the number of video frames in the clipped video data can be greater than N. By restricting the minimum number of video frames in the clipped video data, the video frames in the clipped video data can provide features for decision-making, avoiding the situation where insufficient features lead to the inability to infer the classification result.
[0060] Optionally, the number of video frames in the clipped video data can be less than M, where N < M. By restricting the maximum number of video frames in the clipped video data, while ensuring that features for decision-making are provided for the classification model, the data processing volume of the classification model is reduced, and the training efficiency of the classification model is improved.
[0061] Optionally, a sample data set is obtained. The sample data set can include multiple sample video data. Local sample video data in the multiple sample video data can be marked to obtain the clipped video data corresponding to the local sample video data respectively. The classification model to be trained is pre-trained with a small amount of clipped video data to achieve fast learning of the classification model in the clipped video data and obtain a pre-trained classification model. Further, the pre-trained classification model is further trained with the sample video data in the sample data set to achieve learning of the classification model for the sample video data and obtain a trained classification model.
[0062] The classification label can be used to represent the actual classification result of the sample video data in the cerebral palsy scenario. Pre-training the classification model to be trained with the clipped video data and the classification label corresponding to the clipped video data can enable faster learning of the features of the video data that affect the decision-making for cerebral palsy. Optionally, the classification model to be trained can be pre-trained by the K-fold training method to obtain multiple pre-trained classification models, and the above multiple pre-trained classification models are evaluated according to one or more of the indicators such as AUC, accuracy, specificity, sensitivity, and precision to determine the final pre-trained classification model.
[0063] The pre-trained classification model is trained with sample video data and the classification labels corresponding to the sample video data to obtain a trained classification model, which improves the generalization ability of the classification model, can more accurately predict the characteristics of the key points of the target object, and improves the accuracy of the classification model prediction.
[0064] Optionally, the model training module is further configured to obtain sample video data, train the classification model to be trained based on the sample video data and the classification labels corresponding to the sample video data, and adjust the model parameters of the classification model to be trained according to one or more of the indexes such as AUC, accuracy, specificity, sensitivity, and precision to obtain a trained classification model.
[0065] In the technical solution of this embodiment, the video data of the target object is obtained through the data processing module, the key points of the target object are identified for multiple video frames respectively through the target detection model to obtain the two-dimensional key point information of the target object in each video frame, the abnormality determination is performed on the two-dimensional key point information corresponding to each video frame, the video frame is deleted when the two-dimensional key point information corresponding to any video frame is abnormal, the two-dimensional key point information is standardized based on the resolution information of the video frame to obtain the two-dimensional key point information located in the video frame with the standard resolution, the dimension conversion process is performed on the two-dimensional key point information located in the video frame with the standard resolution through the data conversion model to obtain the three-dimensional key point information corresponding to the two-dimensional key point information, the pre-trained classification model is obtained by pre-training the classification model to be trained with the clipped video data and the classification labels corresponding to the clipped video data, and the trained classification model is obtained by training the pre-trained classification model with the sample video data and the classification labels corresponding to the sample video data. The three-dimensional key point information corresponding to each video frame is classified through the trained classification model to obtain the classification result of the target object in the cerebral palsy scenario, which improves the classification efficiency of the target object in the cerebral palsy scenario.
[0066] Embodiment 2
[0067] Figure 3 It is a flowchart of a classification method applicable to the cerebral palsy scenario provided by the second embodiment of the present invention. This embodiment is applicable to the situation of classifying the target object in the cerebral palsy scenario. As Figure 3 shown, the classification method includes:
[0068] S210. Obtain the video data of the target object, where the video data includes multiple video frames; extract the two-dimensional key point information of the target object in each video frame and convert the two-dimensional key point information into three-dimensional key point information.
[0069] Optionally, the key points include one or more of the nose, left and right eyes, left and right ears, left and right shoulder joints, left and right joints, left and right wrist joints, left and right hip joints, and left and right knee joints and left and right ankle joints.
[0070] The two-dimensional key point information can be obtained through relevant models, for example, it can be an object detection model.
[0071] Optionally, extracting the two-dimensional key point information of the target object in each video frame includes: respectively performing key point recognition on multiple video frames through an object detection model to obtain the two-dimensional key point information of the target object in each video frame.
[0072] Optionally, perform anomaly determination on the two-dimensional key point information corresponding to each video frame, where the anomaly determination includes one or more of missing value determination and data range anomaly determination; in the case where the two-dimensional key point information corresponding to any video frame is abnormal, delete the video frame.
[0073] Since the resolutions of video data collected by different devices may vary, it may cause coordinate errors to a certain extent. Therefore, before converting the two-dimensional key point information into three-dimensional key point information, the two-dimensional key point information can be normalized.
[0074] Optionally, before converting the two-dimensional key point information into three-dimensional key point information, normalize the two-dimensional key point information based on the resolution information of the video frame to obtain the two-dimensional key point information located in the video frame with the standard resolution.
[0075] Optionally, perform dimension conversion processing on the two-dimensional key point information based on a data conversion model to obtain the three-dimensional key point information corresponding to the two-dimensional key point information, where the data conversion model includes multiple one-dimensional convolutional blocks.
[0076] During the process of obtaining the two-dimensional key point information and converting the two-dimensional key point information into three-dimensional key point information, there may be abnormal situations of the three-dimensional key point information. For example, during the process of obtaining the two-dimensional key point information, the two-dimensional key point information of some key points cannot be obtained, resulting in the situation that the corresponding three-dimensional key point information converted from the two-dimensional key point information of some key points is missing, or the three-dimensional key point information exceeds a certain range. To reduce the impact of abnormal situations, the video frames corresponding to the abnormal three-dimensional key point information can be deleted.
[0077] S220. Based on the pre-trained classification model, perform classification processing on the three-dimensional key point information corresponding to each video frame to obtain the classification result of the target object in the cerebral palsy scenario.
[0078] Optionally, the classification model includes two feature extraction units and a classification unit, where any one of the feature extraction units includes a long short-term memory processing network block and a dropout layer, and the classification unit includes a fully connected layer and a pooling layer.
[0079] Optionally, the classification method further includes a model training method for obtaining clipped video data, where the clipped video data includes labeled video frames in the sample video data, and the labeled video frames are partial video frames in the sample video data; pre-training a classification model to be trained based on the clipped video data and the classification labels corresponding to the clipped video data to obtain a pre-trained classification model; and training the pre-trained classification model based on the sample video data and the classification labels corresponding to the sample video data to obtain a trained classification model.
[0080] In the technical solution of this embodiment, by obtaining the video data of the target object, respectively performing key point recognition on multiple video frames through the target detection model to obtain the two-dimensional key point information of the target object in each video frame, performing normalization processing on the two-dimensional key point information based on the resolution information of the video frame, performing dimension conversion processing on the normalized two-dimensional key point information through the data conversion model to obtain the corresponding three-dimensional key point information, performing anomaly determination on the three-dimensional key point information corresponding to each video frame, deleting the video frame when the three-dimensional key point information corresponding to any video frame is abnormal, and performing classification processing on the three-dimensional key point information corresponding to each video frame based on the pre-trained classification model to obtain the classification result of the target object in the cerebral palsy scenario, which shortens the time required for the target object in the cerebral palsy detection process and improves the classification efficiency of the target object in the cerebral palsy scenario.
[0081] Embodiment III
[0082] Figure 4 It is a schematic structural diagram of an electronic device provided by Embodiment III of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0083] As Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the random access memory (RAM) 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the read-only memory (ROM) 12, and the random access memory (RAM) 13 are connected to each other via a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0084] Multiple components in the electronic device 10 are connected to the input / output (I / O) interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0085] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the classification method applicable to the cerebral palsy scenario.
[0086] In some embodiments, the classification method applicable to the cerebral palsy scenario can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the read-only memory (ROM) 12 and / or the communication unit 19. When the computer program is loaded into the random access memory (RAM) 13 and executed by the processor 11, one or more steps of the classification method applicable to the cerebral palsy scenario described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the classification method applicable to the cerebral palsy scenario in any other appropriate way (for example, by means of firmware).
[0087] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0088] The computer program for implementing the classification method applicable to the cerebral palsy scenario of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0089] Embodiment IV
[0090] Embodiment IV of the present invention further provides a computer-readable storage medium storing computer instructions for causing a processor to execute a classification method applicable to the cerebral palsy scenario, the method including:
[0091] Obtaining video data of a target object, where the video data includes a plurality of video frames; extracting two-dimensional key point information of the target object in each video frame and converting the two-dimensional key point information into three-dimensional key point information; and performing classification processing on the three-dimensional key point information corresponding to each video frame based on a pre-trained classification model to obtain a classification result of the target object in the cerebral palsy scenario.
[0092] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0093] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0094] The systems and techniques described herein can be implemented in a computing system that includes backend components (such as, for example, a data server), or a computing system that includes middleware components (such as, for example, an application server), or a computing system that includes frontend components (such as, for example, a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (such as, for example, a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0095] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0096] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0097] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A classification system applicable to the cerebral palsy scenario, characterized in that, Comprising: A data processing module and a classification module, wherein, The data processing module is used to obtain video data of a target object, and the video data includes a plurality of video frames; extract two-dimensional key point information of the target object in each of the video frames, and convert the two-dimensional key point information into three-dimensional key point information; wherein, the target object is an object that needs to be classified in a cerebral palsy scenario; The classification module is used to extract features from the three-dimensional key point information corresponding to each of the video frames in the video data based on a trained classification model, and perform classification processing on the video data based on the extracted features to obtain a classification result of the target object in the cerebral palsy scenario; the classification result is used to determine whether the target object has cerebral palsy; The system further includes a model training module, which is used to obtain clipped video data, and the clipped video data includes marked video frames in the sample video data, and the marked video frames are partial video frames in the sample video data; Pre-train a classification model to be trained based on the clipped video data and the classification label corresponding to the clipped video data to obtain a pre-trained classification model; Train the pre-trained classification model based on the sample video data and the classification label corresponding to the sample video data to obtain a trained classification model; The classification model includes two feature extraction units and a classification unit, wherein, any one of the feature extraction units includes a long short-term memory processing network block and a dropout layer, and the classification unit includes a fully connected layer and a pooling layer; Wherein, the long short-term memory processing network block includes 2 layers of long short-term memory networks.
2. The system according to claim 1, wherein The key points include one or more of the nose, left and right eyes, left and right ears, left and right shoulder joints, left and right joints, left and right wrist joints, left and right hip joints, and left and right knee joints and left and right ankle joints; The data processing module is further used to respectively identify key points of the plurality of video frames through a target detection model to obtain two-dimensional key point information of the target object in each of the video frames.
3. The system according to claim 1, wherein The data processing module is further used for: Perform dimensionality conversion processing on the two-dimensional key point information based on a data conversion model to obtain three-dimensional key point information corresponding to the two-dimensional key point information, wherein the data conversion model includes a plurality of one-dimensional convolutional blocks.
4. The system according to claim 1 or 3, characterized in that The data processing module is further used to, before converting the two-dimensional key point information into three-dimensional key point information, perform normalization processing on the two-dimensional key point information based on the resolution information of the video frame to obtain two-dimensional key point information located in a standard resolution video frame.
5. The system according to claim 1, wherein The data processing module is further used for: Perform anomaly determination on the two-dimensional key point information corresponding to each of the video frames, wherein the anomaly determination includes one or more of missing value determination and data range anomaly determination; Delete the video frame in the case where the two-dimensional key point information corresponding to any video frame is abnormal.
6. A classification method applicable to the cerebral palsy scenario, characterized in that, Comprising: Obtain video data of a target object, where the video data includes multiple video frames; extract two-dimensional key point information of the target object in each of the video frames, and convert the two-dimensional key point information into three-dimensional key point information; wherein, the target object is an object that needs to be classified in a cerebral palsy scenario. Based on the trained classification model, extract features from the three-dimensional key point information corresponding to each of the video frames in the video data, and perform classification processing on the video data based on the extracted features to obtain a classification result of the target object in the cerebral palsy scenario; the classification result is used to determine whether the target object has cerebral palsy. The method further includes: Obtain clipped video data, where the clipped video data includes marked video frames in the sample video data, and the marked video frames are partial video frames in the sample video data. Based on the clipped video data and the classification labels corresponding to the clipped video data, pre-train a classification model to be trained to obtain a pre-trained classification model. Based on the sample video data and the classification labels corresponding to the sample video data, train the pre-trained classification model to obtain a trained classification model. The classification model includes two feature extraction units and a classification unit. Among them, any one of the feature extraction units includes a long short-term memory processing network block and a dropout layer, and the classification unit includes a fully connected layer and a pooling layer. Among them, the long short-term memory processing network block includes 2 layers of long short-term memory networks.
7. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute a classification method applicable to a cerebral palsy scenario as claimed in claim 6.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a processor to implement a classification method applicable to a cerebral palsy scenario as claimed in claim 6 when executed.
Citation Information
Patent Citations
Motion test scoring method, apparatus and device, and storage medium
CN108921907A
Three-dimensional skeleton generation method and computer equipment
CN110874865A