An Interactive Video Action Comprehensive Recognition and Evaluation System and Method

Through the interactive video action comprehensive recognition and evaluation system, the problem of the existing technology being unable to recognize and evaluate complex actions is solved, and detailed recognition and evaluation of single-person, multiple-person, and people-object interaction actions are realized, and a general video action recognition and evaluation method is provided.

CN114677765BActive Publication Date: 2025-05-30ORIENTAL MIND (WUHAN) COMPUTING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210448232.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-24
Publication Date
2025-05-30
Estimated Expiration
2042-04-24

AI Technical Summary

Technical Problem

Existing video action recognition and evaluation methods cannot effectively identify and evaluate actions that include local forms and changes outside the human body skeleton and common changes between the human body and external things. The evaluation methods are relatively vague and cannot provide details and corrective directions.

Method used

An interactive video action comprehensive recognition and evaluation system and method are proposed, including data acquisition, data annotation, action recognition and action evaluation components. The system can collect video data in real time, process it through the user's new action recognition and evaluation algorithm model, and generate comprehensive video action feature recognition results and evaluation results.

Benefits of technology

It expands the scope of traditional action description, can identify single, multi-person, person-object interaction and other actions, and provides detailed comprehensive evaluation indicators to form a general video action recognition and evaluation method, which has good scalability and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114677765B_ABST
    Figure CN114677765B_ABST
Patent Text Reader

Abstract

The present invention provides an interactive video action comprehensive recognition and evaluation system and method; the present invention expands the action description scope that is traditionally only based on human skeleton modeling algorithms, can model actions generated by single-person, human-object interaction, and multi-person interaction, and provides a rich comprehensive evaluation index system, forming a general video action recognition and evaluation method solution, which can describe the differences between the actions to be evaluated and the standards from multiple qualitative and quantitative levels, and can integrate the latest data, data processing technologies, and the most advanced algorithm models with the development of technology to achieve the best action recognition and evaluation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of computer technology, and in particular, to an interactive video action comprehensive recognition and evaluation system and method. Background Art

[0002] The recognition and evaluation of video actions are widely used in industries such as security and sports. It generally uses static or moving cameras to collect videos of human movements in specific scenarios, and based on artificial intelligence algorithms, it recognizes and analyzes the angles of human limbs, action categories, etc., and compares them with set standards to determine whether there are abnormalities or evaluate the standardization of actions.

[0003] In the prior art regarding the algorithms for video action recognition and evaluation, some scholars have proposed several expert rules and artificial intelligence algorithms for certain specific application scenarios.

[0004] In the prior art, a way to recognize video actions is to use a Kinect sensor to collect the coordinates of human joint points, and then calculate the angles between the joint points through a series of artificially designed expert rules, and compare them with preset standard values to determine whether the sit-up action is qualified.

[0005] Another way to recognize video actions is an action recognition method based on artificial intelligence. It mainly uses a classic human skeleton sequence feature classification model, which includes 3 sub-models, respectively for extracting human skeleton features of single-frame images, encoding temporal features, and action classification. Most existing human action classification methods are based on this mode, only with improvements in the performance of specific algorithms.

[0006] There is also a way to recognize video actions that simultaneously considers the features of the target object itself and the features of the background area. By a series of methods, the influence value of the background area is determined, and after fusing with the target object features, action classification is performed. Considering that human actions are very likely to be related to the background environment, this method broadens the video action recognition to more information outside the human body range, and uses richer features to ultimately improve the recognition accuracy of the algorithm.

[0007] However, the definition of actions in the above three existing technologies is still one-sided. Most of them only consider the actions reflected by the human skeleton itself and its temporal changes (such as walking, running, sit-ups), or use some vague external features (such as the background area influence value mentioned above) to improve the recognition accuracy. A broader definition of actions not only includes the actions reflected by the human skeleton morphology and changes, but also includes local morphology and changes unrelated to the human skeleton (such as facial expressions, visual directions), actions reflected by the co-variation of the human body and external things (such as drinking water and raising a flag, the meaning of the action will be uncertain if external objects such as water cups and flags are not considered), and actions reflected by multi-person behaviors (such as fighting). These scenarios are very common in fields such as sports and security, but the existing methods cannot solve them. In addition, the existing action evaluation methods generally rely on expert rules or overall similarity calculations to give a similarity score or similarity level, which is relatively vague and cannot give details of differences and correction directions (such as action delay).

[0008] The above defects make the existing methods can only be carefully designed for specific tasks, with weak generalization ability, and cannot be used as a general solution for video action recognition and evaluation methods. In addition, with the continuous accumulation of business data and the continuous breakthrough of algorithm research, there is a possibility of improving the model accuracy, and there is a possibility of expanding or changing the action recognition requirements. The existing methods cannot meet these well.

[0009] The above content is only used to assist in understanding the technical solution of the present invention, and does not represent an admission that the above content is prior art. Summary of the Invention

[0010] To solve the above technical problems mentioned in the background art, the present invention proposes an interactive video action comprehensive recognition and evaluation system and method. The system includes:

[0011] The system includes a data acquisition component 100, a data annotation component 200, an action recognition component 300, and an action evaluation component 400:

[0012] The data acquisition component 100 is used for collecting original video data;

[0013] The action recognition component 300 is used for receiving a user's request for adding a new action recognition algorithm model component and adding it to the action recognition algorithm model library;

[0014] The action evaluation component 400 is used for receiving a user's request for adding a new action evaluation algorithm model component and adding it to the action evaluation algorithm model library;

[0015] The action recognition component 300 is used to receive the action recognition scope and algorithm model combination configuration set by the user, form a comprehensive video action recognition method, and entrust the data annotation component 200 to start the data annotation service for the corresponding task for the algorithm that needs to be trained by the action recognition component, so that the user can perform data annotation based on the interactive interface to generate the first annotation data;

[0016] The action evaluation component 400 receives the action evaluation index and algorithm model combination configuration set by the user, forms a comprehensive video action evaluation method, and entrusts the data annotation component 200 to start the data annotation service for the corresponding task for the algorithm that needs to be trained by the action evaluation component 400, so that the user can perform data annotation based on the interactive interface to generate the second annotation data; wherein, the action evaluation component 400 uses the second annotation data to train the algorithm that needs to be trained in the comprehensive video action evaluation method, and obtains and saves the corresponding evaluation model;

[0017] The action recognition component 300 is used to perform inference on the video data collected in real time by the data collection component based on the comprehensive video action recognition method, and output the comprehensive video action feature recognition result;

[0018] The action evaluation component 400 performs inference based on the comprehensive video action evaluation method on the video data collected in real time by the data collection component and the recognition result of the comprehensive video action feature output by the above action recognition component, and outputs the comprehensive video action evaluation result.

[0019] Correspondingly, in the comprehensive video action feature recognition result output by the action recognition component 300, it at least includes the human body's own features and the external object features that have related changes with the human body's own features.

[0020] In addition, to achieve the above object, the present invention also proposes a comprehensive interactive video action recognition and evaluation method, and the method includes the following steps:

[0021] Call the data collection component to collect the original video data;

[0022] The action recognition component receives the user's request for adding a new action recognition algorithm model component, and adds it to the action recognition algorithm model library;

[0023] The action evaluation component receives the user's request for adding a new action evaluation algorithm model component, and adds it to the action evaluation algorithm model library;

[0024] The action recognition component receives the action recognition scope and algorithm model combination configuration set by the user to form a comprehensive video action recognition method. For the algorithms that need to be trained by the action recognition component, the data annotation component is entrusted to start the data annotation service for the corresponding tasks, which is used for the user to perform data annotation based on the interactive interface to generate the first annotation data;

[0025] The action evaluation component receives the action evaluation index and algorithm model combination configuration set by the user to form a comprehensive video action evaluation method. For the algorithms that need to be trained by the action evaluation component, the data annotation component is entrusted to start the data annotation service for the corresponding tasks, which is used for the user to perform data annotation based on the interactive interface to generate the second annotation data; among them, the action evaluation component uses the second annotation data to train the algorithms that need to be trained in the comprehensive video action evaluation method to obtain and save the corresponding evaluation models;

[0026] The action recognition component performs inference on the video data collected in real time by the data collection component based on the comprehensive video action recognition method and outputs the comprehensive video action feature recognition result;

[0027] The action evaluation component performs inference on the video data collected in real time by the data collection component and the recognition result of the comprehensive video action feature output by the above action recognition component based on the comprehensive video action evaluation method and outputs the comprehensive video action evaluation result.

[0028] Correspondingly, in the comprehensive video action feature recognition result output by the action recognition component, at least the human body's own features and the external thing features that have related changes with the human body's own features are included.

[0029] Preferably, the step of the action recognition component performing inference on the video data collected in real time by the data collection component based on the comprehensive video action recognition method and outputting the comprehensive video action feature recognition result includes:

[0030] The data collection component inputs the collected video data into the action recognition component, and the action recognition component calls the comprehensive video action recognition method to perform inference to obtain a video frame pool and an action feature pool, and at the same time outputs the recognition result;

[0031] Correspondingly, the step of the action evaluation component performing inference on the video data collected in real time by the data collection component and the comprehensive video action feature representation output by the above action recognition component based on the comprehensive video action evaluation method and outputting the comprehensive video action evaluation result includes:

[0032] The action evaluation component analyzes the above video frame pool and action feature pool, performs action evaluation algorithm inference based on the video action comprehensive evaluation method and in combination with a preset standard action video, and outputs an evaluation result.

[0033] Preferably, the combination configuration of the action recognition scope and algorithm model set by the user, and the combination configuration of the action evaluation scope and algorithm model set by the user form an item configuration file; this item configuration file contains the path of the meta-information configuration file of all algorithms or models that the user expects to run in the algorithm model library and the path of the corresponding runtime configuration file of these algorithms or models.

[0034] Preferably, the video action comprehensive recognition method is based on single-frame images or video time series in the video, and uses a variety of feature detection algorithms and models to describe the action definitions and feature characterizations formed by the human body itself and its interaction with the environment;

[0035] Among them, the action definitions and feature types include but are not limited to:

[0036] Image features encoded by various encoding methods in a single frame or multi-frame sequence;

[0037] The coordinates of single-person human body key points in a single frame, and the coordinate sequence formed in multiple frames;

[0038] The coordinates of multi-person human body key points in a single frame, and the coordinate sequence formed in multiple frames;

[0039] The object attributes such as the category, quantity, color, etc. of the object of interest in a single frame, the position information such as the bounding box coordinates and boundary point coordinates, and the sequences formed by the above various features in multiple frames;

[0040] The comprehensive attributes defined by the features of the human body and various things in a single frame;

[0041] The comprehensive attributes defined by the feature sequences of the human body and various things in multiple frames;

[0042] The features and comprehensive attributes that may appear in the future determined by the current feature sequences of the human body and various things.

[0043] Preferably, the preferred algorithms adopted by the video action comprehensive recognition method may include object detection algorithms, human body key point detection algorithms, object tracking algorithms, skeleton modeling algorithms, and sequence classification algorithms, and there are combinations and nestings between various algorithms, which are used to describe the action definitions and feature characterizations formed by the interaction between the human body's own features and the external thing features in a single frame and video sequence.

[0044] Preferably, the action recognition algorithm model library and the action evaluation algorithm model library are a collection of code files, model files, and other related files of action recognition algorithms and models and action evaluation algorithms and models; among them, each algorithm / model must include a meta-information configuration file, a runtime configuration file, and an inference script, and a trainable algorithm must also include a training startup script.

[0045] Preferably, the meta-information configuration file of the algorithm and model specifically includes: the name of the algorithm and model; the type of the algorithm and model; the paths of the training startup script and the inference startup script of the algorithm and model; the data type input for the algorithm model inference; the data type output from the algorithm and model inference;

[0046] The data types input / output for the algorithm model inference include the following situations: the actual type of the object in memory and the parameters used to describe its attributes, as well as their nested structure. The runtime configuration file of the algorithm and model is all the parameter configuration information involved or relied on in the training process or the inference process of the algorithm and model.

[0047] The beneficial effects of the present invention are as follows: The interactive video action comprehensive recognition and evaluation system and method expand the action description scope that is traditionally only based on human skeleton modeling algorithms, can model actions generated by single-person, human-object interaction, and multi-person interaction, and provide a rich comprehensive evaluation index system, forming a general video action recognition and evaluation method solution, which can describe the differences between the actions to be evaluated and the standards from multiple qualitative and quantitative levels, and can integrate the latest data, data processing technologies, and the most advanced algorithm models with the development of technology to achieve the best action recognition and evaluation effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a schematic diagram of the components of the interactive video action comprehensive recognition and evaluation system according to an embodiment of the present invention;

[0049] Figure 2 is a schematic diagram of the training process of the action recognition model according to an embodiment of the present invention;

[0050] Figure 3 is a schematic diagram of the inference process of the action recognition model according to an embodiment of the present invention;

[0051] Figure 4 is a schematic diagram of the training process of the action evaluation model according to an embodiment of the present invention;

[0052] Figure 5 is a schematic diagram of the inference process of the action evaluation model according to an embodiment of the present invention;

[0053] Figure 6It is a schematic diagram of the physical deployment of the interactive video action comprehensive recognition and evaluation system according to an embodiment of the present invention. Detailed implementation manners

[0054] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the convenience of description, only parts related to the present invention rather than all structures are shown in the drawings.

[0055] Embodiments of the interactive video action comprehensive recognition and evaluation system and method

[0056] First, the interactive video action comprehensive recognition and evaluation system proposed according to an embodiment of the present invention will be described with reference to the accompanying drawings. Figure 1 An interactive video action comprehensive recognition and evaluation system according to an embodiment of the present invention.

[0057] As Figure 1 shown, the system 10 includes: a data acquisition component 100, a data annotation component 200, an action recognition component 300, and an action evaluation component 400.

[0058] Among them, the data acquisition component 100 is used for collecting original video data; on the one hand, the original video data collected by the data acquisition component 100 is used as the input of the data annotation component during the model training stage to generate annotation data, and on the other hand, it is used to generate the original video data to be recognized and evaluated during the inference stage.

[0059] The action recognition component 300 is used to receive a user's request for adding a new action recognition algorithm model component and add it to the action recognition algorithm model library.

[0060] The action evaluation component 400 is used to receive a user's request for adding a new action evaluation algorithm model component and add it to the action evaluation algorithm model library.

[0061] It should be noted that the request for the newly added action recognition algorithm model component has a potential association with the original video data collected by the data acquisition component 100, and the request for the newly added action evaluation algorithm model component has a potential association with the original video data collected by the data acquisition component 100.

[0062] That is, for the system to identify a certain person in the video data collected by the data acquisition component 100, the action recognition component 300 requires an algorithm model capable of detecting people. However, in this embodiment, the relationship between the newly added action recognition algorithm model component request / action evaluation algorithm model component request and the collected original video data will not be verified because the same algorithm model can be used for many different data types, and the same data type can also adopt different algorithm models. Therefore, the newly added action recognition algorithm model and the newly added action evaluation algorithm model component in this embodiment can be used for many different data types. In a specific implementation, taking the action recognition of a certain railway signal flag language video as an example, this recognition task requires identifying all human bodies in a real-time video, identifying whether the actions they perform are railway signal flag languages, and classifying specific flag language categories. Then the original video data is the above-mentioned certain railway signal flag language video. The requests for the newly added action recognition algorithm model component and the action evaluation algorithm model component by the user are both potentially related to the above-mentioned certain railway signal flag language video, that is, the relevant algorithms for the newly added action recognition and action evaluation are both potentially relevant to the actions of railway signals. The newly added action recognition algorithm model component and the action evaluation algorithm model component by the user refer to the collection of the code files and other relevant files of the above algorithm model. These files are added to their respective algorithm model libraries by the action recognition algorithm model component or the action evaluation algorithm model component.

[0063] The action recognition component 300 is used to receive the action recognition scope and algorithm model combination configuration set by the user, form a video action comprehensive recognition method, and for the algorithm that needs to be trained by the action recognition component, entrust the data annotation component 200 to start the data annotation service for the corresponding task, which is used for the user to perform data annotation based on the interactive interface to generate the first annotation data. The action recognition component uses the first annotation data to train the algorithm that needs to be trained in the video action comprehensive recognition method, and obtains and saves the corresponding recognition model.

[0064] Among them, the preferred algorithms adopted by the video action comprehensive recognition method in this embodiment may include object detection algorithms, human key point detection algorithms, object tracking algorithms, skeleton modeling algorithms, and sequence classification algorithms, which are used to describe the action definitions and feature characterizations formed by the interaction between the characteristics of the human body itself and the characteristics of external things in single frames and video sequences.

[0065] It can be understood that the action recognition scope and algorithm model combination configuration set by the user, and the action evaluation scope and algorithm model combination configuration set by the user refer to a project configuration file. This file contains the meta-information configuration file paths of all algorithms / models that the user expects to run in the algorithm model library and the runtime configuration file paths corresponding to these algorithms / models.

[0066] The action evaluation component 400 is used to receive the configured combination of action evaluation metrics and algorithm models set by the user to form a comprehensive video action evaluation method. For the algorithms that need to be trained in the action evaluation component 400, the data annotation component 200 is entrusted to start the data annotation service for the corresponding tasks, which is used for the user to perform data annotation based on the interactive interface to generate second annotation data. Among them, the action evaluation component 400 uses the second annotation data to train the algorithms that need to be trained in the comprehensive video action evaluation method, and obtains and saves the corresponding evaluation models.

[0067] The action recognition component 300 is used to perform inference on the video data collected in real time by the data collection component based on the comprehensive video action recognition method, and output the comprehensive video action feature recognition result. Among the comprehensive video action feature recognition results output by the action recognition component 300, it includes at least the human body's own features and the features of external things that have related changes with the human body's own features.

[0068] In a specific implementation, the original video data is a certain railway signal flag language video as described above. The staff waving the flag in the video belongs to the human body's own features, while the flag waved by the staff belongs to the features of external things.

[0069] It should be noted that the comprehensive video action recognition method in this embodiment is based on single-frame images or video time series sequences in the video, and uses a variety of feature detection algorithms and models to describe the action definitions and feature characterizations formed by the human body itself and its interaction with the environment. The action definitions and feature types include but are not limited to:

[0070] Image features encoded by various encoding methods in single-frame or multi-frame sequences;

[0071] The coordinates of single-person human body key points in a single frame and the coordinate sequences formed in multiple frames;

[0072] The coordinates of multi-person human body key points in a single frame and the coordinate sequences formed in multiple frames;

[0073] The object attributes such as the category, quantity, and color of the objects of interest in a single frame, the position information such as the bounding box coordinates and boundary point coordinates, and the sequences formed by the above various features in multiple frames;

[0074] The comprehensive attributes defined by the features of the human body and various things in a single frame;

[0075] The comprehensive attributes defined by the feature sequences of the human body and various things in multiple frames;

[0076] The features and comprehensive attributes that may appear in the future determined by the current feature sequences of the human body and various things;

[0077] Based on the video action comprehensive evaluation method, the action evaluation component 400 performs inference on the video data collected in real time by the data collection component and the recognition result of the comprehensive video action features output by the above action recognition component, and outputs a comprehensive video action evaluation result.

[0078] It should be noted that in this embodiment, the video action comprehensive recognition method includes both artificial feature extraction algorithms directly used for images and optical flow, and deep learning algorithms based on image classification, object detection, key point detection, skeleton modeling algorithms, sequence classification, etc., and there are combinations and nestings between various algorithms. These deep learning algorithms need to use labeled data for supervised learning in the training stage to fit the neural network parameters, and then feature extraction and prediction can be performed in the inference stage. It can be understood that the deep learning algorithm is an algorithm that performs representation learning on data with an artificial neural network as the architecture. The learning of the deep learning algorithm is to iteratively train the neural network with a set of hyperparameters to obtain an estimated value of the neural network parameters. Different deep learning algorithms require different data types and formats. For example, the object detection algorithm requires single-frame images and the bounding box coordinates of the objects of interest, and the skeleton modeling algorithm requires the human key point coordinates of single-frame or sequences. These data may come from the original video frame sequence, and may also come from the operation output of other algorithms.

[0079] Specifically, in this embodiment, the interactive video action comprehensive recognition and evaluation process is mainly divided into two periods: training and inference. In the training period, the data collection component 100 collects or receives a large amount of image and video data. The action recognition component 300 and the action evaluation component 400 receive the newly added algorithm model components and algorithm model combination schemes from the user, and entrust the data annotation component 200 to start the data annotation service. The action recognition component 300 and the action evaluation component 400 are trained on the labeled data corresponding to their respective algorithms to generate and save the models. In the inference period, the data collection component 100 collects images or videos and inputs them into the action recognition component 300. The action recognition component 300 performs inference on the action recognition algorithm, saves a video frame pool and an action feature pool of a certain length, and outputs the recognition result at the same time. The action evaluation component 400 analyzes the above video frame pool and action feature pool, and performs action evaluation algorithm inference in combination with the evaluation criteria to output the evaluation result.

[0080] See Figure 6 , Figure 6Shows a schematic diagram of the physical deployment of the system, which consists of an image / video acquisition device and a server. The server is responsible for the operation of all system components and can perform data transmission and interaction with the data acquisition device. Specifically, when the data acquisition component 100 in the server receives a data acquisition request, it controls the data acquisition device to collect image and video data and saves them in the server. Other components in the system, such as data annotation, action recognition, and action evaluation, run in the server.

[0081] Correspondingly, based on Figure 1 the system of, the present invention corresponds to a set of method embodiments. In this embodiment, the method includes the following stage steps:

[0082] Video acquisition and manual operation stage:

[0083] Invoke the data acquisition component to collect original video data;

[0084] The action recognition component receives a user's request to add an action recognition algorithm model component and adds it to the action recognition algorithm model library;

[0085] The action evaluation component receives a user's request to add an action evaluation algorithm model component and adds it to the action evaluation algorithm model library;

[0086] Action recognition model training stage:

[0087] The action recognition component receives the action recognition scope and algorithm model combination configuration set by the user to form a comprehensive video action recognition method. For the algorithms that need to be trained by the action recognition component, it entrusts the data annotation component to start the data annotation service for the corresponding tasks, for the user to perform data annotation based on the interactive interface to generate the first annotation data; wherein, the action recognition component uses the first annotation data to train the algorithms that need to be trained in the comprehensive video action recognition method to obtain and save the corresponding recognition models; wherein, the comprehensive video action recognition method in this embodiment preferably includes at least a target detection algorithm, a human key point detection algorithm, a target tracking algorithm, a skeleton modeling algorithm, and a sequence classification algorithm, which are used to describe the action definitions and feature characterizations formed by the interaction between the human body's own features and the external thing features in single frames and video sequences;

[0088] Action evaluation model training stage:

[0089] The action evaluation component receives the action evaluation metrics and algorithm model combination configuration set by the user to form a comprehensive video action evaluation method. For the algorithms that need to be trained by the action evaluation component, the data annotation component is entrusted to start the data annotation service for the corresponding tasks, enabling the user to perform data annotation based on the interactive interface to generate second annotation data. Among them, the action evaluation component uses the second annotation data to train the algorithms that need to be trained in the comprehensive video action evaluation method, and obtains and saves the corresponding evaluation models.

[0090] Action recognition inference process stage:

[0091] Based on the comprehensive video action recognition method, the action recognition component performs inference on the video data collected in real time by the data collection component and outputs the comprehensive video action feature recognition result.

[0092] Action evaluation inference process stage:

[0093] Based on the comprehensive video action evaluation method, the action evaluation component performs inference on the video data collected in real time by the data collection component and the recognition result of the comprehensive video action features output by the above action recognition component, and outputs the comprehensive video action evaluation result.

[0094] Furthermore, for each stage of the above interactive video action comprehensive recognition and evaluation method of the present invention, different method embodiments are introduced in detail respectively:

[0095] Method Embodiment 1 <Action Recognition Model Training Stage>

[0096] In this embodiment, an action recognition model training embodiment is provided, which can be executed by the action recognition component 300 in this embodiment. As a specific example, in this embodiment, the action recognition of a certain railway signal flag language video is taken as an example. This recognition task requires identifying all human bodies in the real-time video, identifying whether the actions they perform are railway signal flag languages, and classifying specific flag language categories.

[0097] In this embodiment, the action recognition algorithm model component includes: an object detection algorithm for detecting human bodies and flags (distinguishing colors), an object tracking algorithm for tracking human bodies and flags, a key point detection algorithm for detecting key points of human body parts, a human skeleton modeling algorithm for skeleton modeling, a feature encoder for general image feature encoding, and a sequence feature classification algorithm for classifying action categories. Except for the tracking algorithm, the skeleton modeling algorithm, and the general feature encoder, the rest of the algorithms need to be trained. The algorithms listed here are designed according to the recognition requirements of this embodiment. The present invention does not limit the specific action recognition scenarios nor the specific algorithms adopted.

[0098] According toFigure 2 As shown in the figure, the action recognition model training process in this embodiment includes the following steps:

[0099] In step S101, the action recognition component 300 entrusts the data annotation component 200 to enable the data annotation function for the above algorithms.

[0100] In step S102, the data annotation component 200 enables the data annotation service of the corresponding algorithm, and the user performs specific annotation work. In this embodiment, the specific content to be annotated is as follows:

[0101] The category labels and bounding box coordinates of the human body and the flag (distinguishing colors) in a single-frame image;

[0102] The category and coordinates of the human body key points in a single-frame image;

[0103] In the continuous frame sequence, the action category labels corresponding to the above information of each person;

[0104] In step S103, the action recognition component 300 trains each algorithm, and the specific data used by each algorithm is as follows:

[0105] The human body and flag object detection algorithm uses a single-frame image and the corresponding category and bounding box coordinate annotations;

[0106] The human body key point detection algorithm uses a single-person image slice defined by the human body bounding box in a single-frame image and the human body key point coordinates and category annotations relative to the slice; (The following formula is not newly added:)

[0107] The call sequence feature classification algorithm performs certain processing on the original annotation: First, use the object tracking algorithm to extract the key point coordinate sequence of the same human body ID (human body own feature) in any continuous frame sequence, and use the skeleton modeling algorithm to convert it into a skeleton feature sequence, denoted as the actual human body skeleton feature sequence where L is the sequence length, D sk is the feature dimension output by the skeleton modeling algorithm. Use the object tracking algorithm to extract the flags (external object features) existing in these frames, and only consider the flags (target external object features) whose distance from the person (the human body own feature) is less than the preset threshold. If there is a target flag (target external object feature) that meets the requirements, convert the absolute coordinates of the center point of the target flag coordinate box into relative coordinates relative to the center point of the human body coordinate box of the person and normalize it using the width and height of the human body coordinate box, normalize the width and height of the coordinate box of the target flag using the width and height of the human body coordinate box, and denote the processed coordinate box sequence of the target flag as Denote the color category sequence of the one-hot encoded target flag as C flagis the number of color categories. Encode the depth feature vector corresponding to the image slice defined by the original bounding box of the target flag using a general feature encoder and compress it to a fixed length. Denote this feature sequence as D fe is the feature dimension output by the general feature encoder. Concatenate BBox flag,i , Cls flag,i , FE flag,i along the feature dimension to form F flag,i = Concat(BBox flag,i , Cls flag,i , FE flag,i ) and fill in the feature values corresponding to the frames that do not appear with all zeros. Concatenate the feature sequences of all target flags that meet the requirements along the feature dimension to form F flag = Concat(F flag,i ). Truncate the number of flags according to the preset maximum number threshold n, and fill in the features corresponding to the missing values with all zeros. Finally, concatenate F sk and F flag along the feature dimension to form Truncate the sequence length to F' according to the preset sequence window length L max . If the sequence exceeds this length, generate several subsequence data based on a sliding window Use the action category label of the complete sequence as the action category label for each truncated subsequence. The labeled data required by the sequence feature classification algorithm is the sequence feature and action category label {F', Y} of each subsequence, where Y is the action category label index.

[0108] Step S104, after the action recognition component 300 is trained, save each model in the computer storage.

[0109] Method Embodiment 2 <Action Recognition Inference Process Stage>

[0110] In this embodiment, a data inference method for the action recognition model is provided, which can be executed by the action recognition component 300. Still taking the specific scenario in <Method Embodiment 1> as an example, for the output of each model, all outputs will be saved in the action feature pool. Since this scenario only cares about the recognized action category, only the sequence feature classification result needs to be displayed as the action recognition result.

[0111] According to Figure 3 shown, the action recognition model inference process in this embodiment includes the following steps:

[0112] Step S201, the action recognition component 300 initializes the video frame pool Pool video , the action feature pool Pool rec_feat and the recognition result pool Pool rec_result ;

[0113] In step S202, the action recognition component 300 receives the video stream provided by the data acquisition component 100;

[0114] In step S203, each frame of the video stream is traversed in sequence. If the traversal ends, it stops; otherwise, step S204 is executed;

[0115] In step S204, the target frame image Frame i is added to the video frame pool, and the frame sequence number i is marked;

[0116] In step S205, taking Frame i as the input, the forward propagation Detect(·) of the target detection algorithm is run to obtain all the human body bounding box coordinates bbox human , flag category cls flag and the bounding box coordinates bbox flag in the current frame, and they are temporarily stored in the action feature pool;

[0117] In step S206, all the image slices of the human body regions bbox i are extracted from the frame image Frame human , and the forward propagation KPDetect(·) of the human body key point detection algorithm is run to obtain all the human body key point information s kp in the current frame, and they are temporarily stored in the action feature pool;

[0118] In step S207, for all the above human body key point detection results, the forward propagation SK(·) of the skeleton modeling algorithm is run, and the skeleton feature result f sk is temporarily stored in the action feature pool;

[0119] In step S208, all the image slices of the flag regions bbox i are extracted from the frame image Frame flag , and the forward propagation FE(·) of the general feature extraction model is run to obtain the flag image feature fe flag compressed to a fixed length, and it is temporarily stored in the action feature pool;

[0120] In step S209, taking all the above detected bounding boxes bbox human and bbox flag in the current frame as the input, the target tracking algorithm Track(·) is run to determine the identity indices id human and id flag corresponding to all the above target detection results in the current frame, and the identity index information is added to the recognition results of the corresponding target features in the action feature pool, and the frame sequence number i is marked;

[0121] Step S210, when there is an identity index id of any person in the action feature pool human , and its continuous existence exceeds the number of frames of the sequence window length L described in <Method Embodiment 1> max . Obtain all flag detection results bbox flag,i , cls flag,i that appear in these frame numbers. Retain the flags whose distance from this person is less than a certain threshold as described in <Method Embodiment 1>, and then retain the n flags with the most occurrences according to the preset maximum number threshold n of the flags described in <Method Embodiment 1>. For these flags, convert the absolute coordinates of the center point of the flag coordinate box to relative coordinates relative to the center point of the person's body coordinate box and normalize them using the width and height of the person's body coordinate box. Normalize the width and height of the flag coordinate box using the width and height of the person's body coordinate box. Then, according to the feature dimension, concatenate the processed flag bounding box sequence BBox flag , the flag color category sequence Cls flag , and the flag image feature sequence FE flag into the overall flag feature sequence F flag = Concat(BBox flag , Cls flag , FE flag ). For the feature vectors corresponding to the frame numbers where a certain flag does not appear, or the spare feature vectors at the position where the number of flags is less than the maximum truncation value n, fill them with all 0s. Concatenate the actual human skeleton feature sequence F sk formed by the skeleton feature f sk of this person output in Step S207 and the above overall flag feature sequence F flag together according to the feature dimension F = Concat(F sk , F flag ), and then input it into the sequence feature classification model Action(·). Write the action classification act result into the action feature pool and the recognition result pool at the same time, and mark the frame number range {...} included in this sequence.

[0122] Method Embodiment 3 <Action Evaluation Model Training Stage>

[0123] In this embodiment, an action evaluation model training method is provided, and this method can be executed by the action evaluation component 400. Still taking the specific scenario in <Method Embodiment 1> as an example, in addition to detecting the action category, this embodiment introduces a new task, that is, on the basis of recognizing that the action performed by the human body is the railway flag signal action, evaluating the difference degree between this action and the standard action.

[0124] In this embodiment, the action evaluation algorithm model component includes: a numerical difference calculation method and a sequence execution difference prediction algorithm. Among them, the numerical difference calculation method outputs the feature numerical differences at the same time points of two sequences based on fixed rules without training; the sequence execution difference prediction algorithm is based on deep learning and is used to predict the frame rate difference and delay of the execution of two sequences, which requires training. In addition, during the data processing, the human target detection algorithm, the target tracking model, and the human skeleton modeling algorithm in <Method Embodiment 1> are also used. The algorithms listed here are designed according to the action evaluation requirements of this embodiment. The present invention does not limit the specific action evaluation scenario nor the specific algorithms used.

[0125] According to Figure 4 As shown, the action evaluation model training process in this embodiment includes the following steps:

[0126] Step S301, the action evaluation component 400 entrusts the data annotation component 200 to enable the data annotation function for the above algorithms.

[0127] Step S302, the data annotation component 200 enables the data annotation service for the corresponding algorithm, and the user performs specific annotation work. In this embodiment, the specific content to be annotated is as follows:

[0128] The human body label and bounding box coordinates in a single-frame image;

[0129] The human body key point category and coordinates in a single-frame image;

[0130] The above annotation content can also be obtained through the same annotation content in the data annotation step of the action recognition model training in <Method Embodiment 1> of the present invention.

[0131] Step S303, the action evaluation 400 trains each algorithm. The sequence execution difference prediction algorithm requires a special training data generation method, which is specifically as follows:

[0132] First, use the target tracking algorithm to extract the key point coordinate sequence of the same (human body) human body own feature ID in any continuous frame sequence in the original video data. In each iteration of training, use the preset sequence window length L in <Method Embodiment 1> max , first randomly intercept a subsequence of this sequence and denote it as C kp is the number of key point categories. Then, sample a frame number interval δ relative to the starting frame of the sequence according to a random distribution , and sample a frame rate scaling factor γ according to another random distribution . Using the position δ frames away from the starting frame of the sequence as the starting position, and using int(γL As the end position, sample a sub-sequence from the original sequence with a length of int(γL) starting from the starting position. The sub-sequence is used as the training data for the sequence execution difference prediction algorithm. As the end position, sample a sub-sequence from the original sequence with a length of int(γL) starting from the starting position. The sub-sequence is used as the training data for the sequence execution difference prediction algorithm.max ) As the sequence length, intercept the sequence Then for the sequence Apply interpolation (such as bilinear interpolation) to adjust its length to L max , to obtain Sample the jitter distance of the key point coordinates according to a random distribution And adjust to Finally, convert the human key point sequence and and into skeleton features using the described skeleton modeling algorithm SK(·) and And splice them in the time dimension into As the training input of the model, using δ and γ as the supervision values of the model output, train the model.

[0133] Step S304, after the action evaluation component 400 is trained, save each model in the computer storage.

[0134] Method Embodiment 4 <Action Recognition Inference Process Stage>

[0135] In this embodiment, a data inference method for an action evaluation model is provided, and this method can be executed by the action evaluation component 400. This embodiment uses the algorithm model proposed in the <Method Embodiment 3> to evaluate and compare the video to be evaluated with a preset standard action video. Among them, the standard action video is a preset video or some videos, including all categories of the railway flag signals in the <Method Embodiment 1>, and there can be multiple standard videos for the same action category.

[0136] According to Figure 5 shown, the action evaluation model inference process in this embodiment includes the following steps:

[0137] Step S401: The action evaluation component 400 initializes the standard action feature pool Pool eval_std_feat and the action evaluation output pool Pool eval_result ;

[0138] Step S402: The action evaluation component 400 uses the object detection algorithm Detect(·), the object tracking algorithm Track(·) and the skeleton modeling algorithm SK(·) to extract the human key points s std_kp (including the sequence S std_kp ) and the human skeleton features f std_sk (including the sequence F std_sk ) corresponding to all standard actions in the above-mentioned labeled action video, and stores them in the standard action feature pool.

[0139] Step S403: The action evaluation component 400 receives the action feature pool Pool output by the action recognition component 300 in real time rec_feat ;

[0140] Step S404: The action evaluation component 400 uses the numerical difference calculation method ValueDiff(·) to calculate the difference diff between the actual detected human key points sp in the action feature pool and the standard human key points ss in the standard action feature pool frame by frame k and writes the difference diff std_kp into the action evaluation output pool, marking the corresponding frame number i kp kp

[0141] Step S405: If there is a human skeleton feature sequence F with a length of the preset sequence window length L of <Method Embodiment 3> in the action feature pool max sk then the action evaluation component 400 splices the actual human skeleton feature sequence F sk and the standard human skeleton feature sequence F of the corresponding frame in the standard action feature pool std_sk in the time series dimension to obtain F = Concat(F std_sk , F sk ), inputs the sequence into the difference prediction model ExeDiff(·) to obtain the action delay prediction value and the action frame rate difference prediction value Writes the results into the action evaluation output pool and marks the corresponding frame number range {...}

[0142] The beneficial effects of the embodiments of the present invention are as follows:

[0143] The action comprehensive recognition and evaluation method constructs a set of general video action recognition and evaluation method systems. This method greatly expands the scope of action description of existing methods, can recognize actions and features in various situations such as single person, multiple people, human interaction, static, and dynamic, also expands the action evaluation indicators, and has good scalability at the same time

[0144] ​​​Through the design of action recognition / evaluation algorithm meta-information and runtime configuration, users can add any specific algorithm implementation to the preset algorithm library and complete data annotation and model training. During the inference period, starting from the original video frames, according to the input data type and output data type in the algorithm meta-information configuration adopted by the user, each algorithm model can implicitly and orderly execute in sequence, take the data of the required input type from the feature pool as input, and the output features are put back into the feature pool to be used by other algorithm models later. The "expert rules", "human skeleton sequence feature classification model", and "background feature fusion" described in the background art can all be regarded as special cases or specific implementations of an algorithm model combination in a certain application of this method. However, this method also has the ability to be extended to scenarios such as human interaction action recognition and evaluation, which other methods do not have.

[0145] In addition, the action comprehensive evaluation and recognition system proposed by the present invention has the ability to dynamically expand data and algorithm models, and can add, change configurations, or add new data for retraining according to needs to adapt to demand changes. Therefore, it can integrate the latest data, data processing technologies, and the most advanced algorithm models with the development of technology to achieve the best action recognition and evaluation effect.

[0146] In addition, it should be noted that in the specific implementation of the above embodiments, the action recognition algorithm model library and the action evaluation algorithm model library refer to the collection of code files / model files and other related files of the action recognition algorithm / model and the action evaluation algorithm / model. Each algorithm / model must include a meta-information configuration file, a runtime configuration file, and an inference script. The trainable algorithms also need to include a training startup script.

[0147] The meta-information configuration file of the above algorithms / models specifically includes: algorithm / model name; algorithm / model type; the path of the training startup script (if any) and the inference startup script of the algorithm / model; the data type of the input for algorithm model inference; the data type of the output for algorithm / model inference; the data types of the input / output for algorithm / model inference include the following situations: the real type of the object in memory (such as: integer, string, list, dictionary, numpy array, etc.) and the parameters used to describe its attributes (such as: list length, array shape, and element type, etc.), as well as their nested structure. Manually defined data type tags, such as using the string "obj_id" to represent the target identity index output by the target tracking algorithm (its real data type is integer or integer list).

[0148] The runtime configuration file of the above algorithm / model refers to all parameter configuration information involved or relied on in the training process or inference process of the algorithm / model. For example: the dataset path, hyperparameters, etc. during model training. The same algorithm / model can have multiple different runtime configuration files.

[0149] Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, it can also include more other equivalent embodiments, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. An interactive video action comprehensive recognition and evaluation system, characterized in that, the system includes a data acquisition component 100, a data annotation component 200, an action recognition component 300, and an action evaluation component 400: the data acquisition component 100 is used to collect original video data; the action recognition component 300 is used to receive a user's request for a new action recognition algorithm model component and add it to the action recognition algorithm model library; the action evaluation component 400 is used to receive a user's request for a new action evaluation algorithm model component and add it to the action evaluation algorithm model library; the action recognition component 300 is used to receive the configured action recognition scope and algorithm model combination set by the user to form a comprehensive video action recognition method. For the algorithms that need to be trained by the action recognition component, it entrusts the data annotation component 200 to start the data annotation service for the corresponding task, for the user to perform data annotation based on the interactive interface to generate first annotation data; wherein, the action recognition component uses the first annotation data to train the algorithms that need to be trained in the comprehensive video action recognition method, and obtains and saves the corresponding recognition models; wherein, the comprehensive video action recognition method is based on single-frame images or video time series in the video, and uses a variety of feature detection algorithms and models to describe the action definitions and feature characterizations formed by the human body itself and its interaction with the environment; the action evaluation component 400 receives the configured action evaluation indicators and algorithm model combination set by the user to form a comprehensive video action evaluation method. For the algorithms that need to be trained by the action evaluation component 400, it entrusts the data annotation component 200 to start the data annotation service for the corresponding task, for the user to perform data annotation based on the interactive interface to generate second annotation data; wherein, the action evaluation component 400 uses the second annotation data to train the algorithms that need to be trained in the comprehensive video action evaluation method, and obtains and saves the corresponding evaluation models; the action recognition component 300 is used to perform inference on the video data collected in real time by the data acquisition component based on the comprehensive video action recognition method, and output the comprehensive video action feature recognition result; the action evaluation component 400 is used to perform inference on the video data collected in real time by the data acquisition component and the recognition result of the comprehensive video action feature output by the above action recognition component based on the comprehensive video action evaluation method, and output the comprehensive video action evaluation result.

2. The interactive video action comprehensive recognition and evaluation system according to claim 1, characterized in that, in the comprehensive video action feature recognition result output by the action recognition component 300, it at least includes the human body's own features and the features of external things that have relevant changes with the human body's own features.

3. An interactive video action comprehensive recognition and evaluation method, characterized in that, the method includes the following steps: call the data acquisition component to collect original video data; the action recognition component receives a user's request for a new action recognition algorithm model component and adds it to the action recognition algorithm model library; The action evaluation component receives a user's request to add a new action evaluation algorithm model component and adds it to the action evaluation algorithm model library; The action recognition component receives the action recognition scope and algorithm model combination configuration set by the user to form a video action comprehensive recognition method. For the algorithms that need to be trained by the action recognition component, it entrusts the data annotation component to start the data annotation service for the corresponding task, which is used for the user to perform data annotation based on the interactive interface to generate first annotation data. Among them, the action recognition component uses the first annotation data to train the algorithms that need to be trained in the video action comprehensive recognition method and obtains and saves the corresponding recognition models. Among them, the video action comprehensive recognition method is based on single-frame images or video time series in the video and uses a variety of feature detection algorithms and models to describe the action definitions and feature characterizations formed by the human body itself and its interaction with the environment; The action evaluation component receives the action evaluation indicators and algorithm model combination configuration set by the user to form a video action comprehensive evaluation method. For the algorithms that need to be trained by the action evaluation component, it entrusts the data annotation component to start the data annotation service for the corresponding task, which is used for the user to perform data annotation based on the interactive interface to generate second annotation data. Among them, the action evaluation component uses the second annotation data to train the algorithms that need to be trained in the video action comprehensive evaluation method and obtains and saves the corresponding evaluation models; The action recognition component performs inference on the video data collected in real time by the data collection component based on the video action comprehensive recognition method and outputs the video action comprehensive feature recognition result; The action evaluation component performs inference on the video data collected in real time by the data collection component and the recognition result of the video action comprehensive feature output by the above action recognition component based on the video action comprehensive evaluation method and outputs the video action comprehensive evaluation result.

4. The method according to claim 3, wherein, in the video action comprehensive feature recognition result output by the action recognition component, it at least includes the human body's own features and the external thing features that have related changes with the human body's own features.

5. The method according to claim 3, wherein, the step of the action recognition component performing inference on the video data collected in real time by the data collection component based on the video action comprehensive recognition method and outputting the recognition result of the video action comprehensive feature includes: the data collection component inputs the collected video data into the action recognition component, and the action recognition component calls the video action comprehensive recognition method to perform inference, obtains a video frame pool and an action feature pool, and outputs the recognition result at the same time; correspondingly, the step of the action evaluation component performing inference on the video data collected in real time by the data collection component and the video action comprehensive feature representation output by the above action recognition component based on the video action comprehensive evaluation method and outputting the video action comprehensive evaluation result includes: The action evaluation component analyzes the above video frame pool and action feature pool, performs action evaluation algorithm reasoning based on the video action comprehensive evaluation method and in combination with a preset standard action video, and outputs an evaluation result.

6. The method according to any one of claims 3-5, wherein, the combination configuration of the action recognition scope and algorithm model set by the user, and the combination configuration of the action evaluation scope and algorithm model set by the user form an item configuration file; this item configuration file contains the path of the meta-information configuration file of all algorithms or models that the user expects to run in the algorithm model library and the path of the runtime configuration file corresponding to these algorithms or models.

7. The method according to any one of claims 3-5, wherein, the action definitions and feature types include but are not limited to: image features encoded by various coding methods in a single frame or multi-frame sequence; the coordinates of single-person human key points in a single frame and the coordinate sequence formed in multiple frames; the coordinates of multi-person human key points in a single frame and the coordinate sequence formed in multiple frames; the category, quantity, color object attributes of the object of interest in a single frame, the boundary box coordinates, the position information of the boundary point coordinates, and the sequences formed by the above various features in multiple frames; the comprehensive attributes defined by the features of the human body and various things in a single frame; the comprehensive attributes defined by the feature sequences of the human body and various things in multiple frames; the features and comprehensive attributes that may appear in the future determined by the current feature sequences of the human body and various things.

8. The method according to claim 7, wherein, the preferred algorithms adopted by the video action comprehensive recognition method include object detection algorithms, human key point detection algorithms, object tracking algorithms, skeleton modeling algorithms, and sequence classification algorithms, and there are combinations and nestings between various algorithms, which are used to describe the action definitions and feature characterizations formed by the interaction of the human body's own features and external thing features in a single frame and video sequence.

9. The method according to any one of claims 3-5, wherein, the action recognition algorithm model library and action evaluation algorithm model library are a collection of code files, model files, and other related files of action recognition algorithms and models and action evaluation algorithms and models; among them, each algorithm / model must contain a meta-information configuration file, a runtime configuration file, and an inference script, and a trainable algorithm must also contain a training start script.

10. The method according to any one of claims 3-5, wherein, the meta-information configuration file of the algorithm and model specifically includes: the name of the algorithm and model; the type of the algorithm and model; the paths of the training start script and inference start script of the algorithm and model; the data type input for the algorithm model inference; the data type output by the algorithm and model inference; the data types of the input / output of the algorithm and model inference include the following situations: the actual types of objects in memory and the parameters used to describe their attributes, and the nested structure between them; the runtime configuration file of the algorithm and model is all the parameter configuration information involved or relied on in the training process or inference process of the algorithm and model.

Citation Information

Patent Citations

  • Human motion posture recognition and evaluation method and system

    CN111401270A

  • Self-weight fitness auxiliary coach system, method and terminal based on human body posture recognition

    CN113762133A