Training method for discriminator model and action recognition method
Through the discriminator model trained by the adversarial training method, the problems of insufficient recognition capabilities and inflexible support for new actions under changes in external conditions are solved, and higher robustness and flexibility are achieved.
Patent Information
- Application Number
- CN202110939838.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-13
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-08-13
AI Technical Summary
The existing action recognition technology is not flexible enough under external conditions such as environment, perspective, and clothing, and has poor flexibility in supporting new actions, resulting in poor generalization.
By determining positive and negative sample pairs, the discriminator model is trained using an adversarial training method to enable it to identify actions in different scenarios and support flexible expansion of action categories.
It improves the robustness and flexibility of action recognition, reduces the time cost of data acquisition and labeling, and supports the rapid identification and support of new actions.
Smart Images

Figure CN113642472B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and further relates to the fields of deep learning and computer vision. In particular, it relates to a training method for a discriminator model, an action recognition method, a device, an electronic device, and a storage medium. Background Art
[0002] With the popularization of video recording devices, the increasing number of video software, the improvement of network speed, etc., a large number of videos are spread on the Internet and increase exponentially. These video information are of various types and huge in quantity, far exceeding the capacity of human manual processing. Therefore, it is very necessary to invent an action recognition method in videos suitable for various applications such as video recommendation, human behavior analysis, video surveillance, etc. Summary of the Invention
[0003] The present disclosure provides a training method for a discriminator model, an action recognition method, a device, an electronic device, and a storage medium.
[0004] According to one aspect of the present disclosure, there is provided a training method for a discriminator model, including: determining a first positive sample pair, the positive sample pair including first template video data and second template video data, both the first template video data and the second template video data including a first template action; determining a first negative sample pair, the negative sample pair including at least one of the first template video data and the second template video data and third template video data, the third template video data including a second template action different from the first template action; and training the discriminator model by using the first positive sample pair and the first negative sample pair.
[0005] According to another aspect of the present disclosure, there is provided an action recognition method, including: inputting a first temporal feature vector of a to-be-recognized video sequence including a to-be-recognized action and a second temporal feature vector of a template video sequence including a template action into a discriminator model to obtain a target template action belonging to the same category as at least part of the action characterized by the first temporal feature vector; and determining an action category of the to-be-recognized action according to the target template action; wherein the discriminator model is trained based on the above training method.
[0006] According to another aspect of the present disclosure, there is provided a training device for a discriminator model, including: a first determination module configured to determine a first positive sample pair, the positive sample pair including first template video data and second template video data, both the first template video data and the second template video data including a first template action; a second determination module configured to determine a first negative sample pair, the negative sample pair including at least one of the first template video data and the second template video data and third template video data, the third template video data including a second template action different from the first template action; and a first training module configured to train the discriminator model by using the first positive sample pair and the first negative sample pair.
[0007] According to another aspect of the present disclosure, there is provided an action recognition device, including: an input module configured to input a first temporal feature vector of a to-be-recognized video sequence including a to-be-recognized action and a second temporal feature vector of a template video sequence including a template action into a discriminator model to obtain a target template action belonging to the same category as at least part of the action characterized by the first temporal feature vector; and a determination module configured to determine the action category of the to-be-recognized action according to the target template action; wherein the discriminator model is trained based on the above-mentioned training device.
[0008] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method as described above.
[0009] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method as described above.
[0010] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, which implements the method as described above when executed by a processor.
[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0012] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0013] Figure 1Schematically shows an exemplary system architecture to which an action recognition method and apparatus according to an embodiment of the present disclosure can be applied;
[0014] Figure 2 Schematically shows a flowchart of a method for training a discriminator model according to an embodiment of the present disclosure;
[0015] Figure 3 Schematically shows a flowchart of an action recognition method according to an embodiment of the present disclosure;
[0016] Figure 4A Schematically shows a schematic diagram of data collection and discriminator training according to an embodiment of the present disclosure;
[0017] Figure 4B Schematically shows a schematic diagram of feature extraction according to an embodiment of the present disclosure;
[0018] Figure 4C Schematically shows a schematic diagram of discriminator processing according to an embodiment of the present disclosure;
[0019] Figure 4D Schematically shows a schematic diagram of post-processing output results according to an embodiment of the present disclosure;
[0020] Figure 5 Schematically shows a block diagram of a training apparatus for a discriminator model according to an embodiment of the present disclosure;
[0021] Figure 6 Schematically shows a block diagram of an action recognition apparatus according to an embodiment of the present disclosure; and
[0022] Figure 7 Shows a schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure. Detailed implementation
[0023] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0024] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved all comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good customs.
[0025] Real-time action recognition based on a monocular RGB (RGB color model, a color standard in the industrial field) camera is a hot research direction in the field of artificial intelligence and is the technical cornerstone of business scenarios such as virtual humans, live broadcast auditing, AR (Augmented Reality), and VR (Virtual Reality), with broad application prospects and great commercial value.
[0026] To achieve action recognition, a multi-classifier is usually directly trained end-to-end using a deep neural network, which generally consists of five steps: data collection, collecting data for the application scenario; data annotation, giving annotation rules according to the target action types and annotating the data accordingly; network training, sending the annotated data into the deep neural network for training; network deployment, performing engineering deployment on the trained network model; and network inference, inputting an image and the network inferring the output category.
[0027] The inventors found in the process of implementing the concept of the present disclosure that end-to-end training of a multi-classifier is strictly bound to the application scenario. For each scenario, data needs to be collected and corresponding annotation rules need to be formulated to determine which actions need to be recognized. Then the data is manually annotated. These two steps require a large amount of human and time costs, usually measured in days. Moreover, the network training step also takes a lot of time. In addition, the action categories supported by the multi-classifier obtained from one training are fixed. If new actions need to be supported subsequently, the three steps of data collection, data annotation, and network training all need to be repeated, which is costly and has poor flexibility and scalability. In addition, when performing action recognition based on the multi-classifier trained end-to-end, it is very sensitive to external conditions such as the environment, perspective, and clothing, and has poor generalization. If the actual data and the training data have a large difference in distribution, the effect will deteriorate severely.
[0028] Therefore, the current dynamic action recognition technology is not mature enough and there is no stable and reliable general solution. In the digital human live broadcast scenario, a highly stable and accurate dynamic action recognition ability is required as a support.
[0029] Figure 1 An exemplary system architecture to which the action recognition method and apparatus according to embodiments of the present disclosure can be applied is schematically shown.
[0030] It should be noted that Figure 1The figure shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios. For example, in another embodiment, the exemplary system architecture to which the action recognition method and apparatus can be applied may include a terminal device, but the terminal device can implement the action recognition method and apparatus provided by the embodiments of the present disclosure without interacting with the server.
[0031] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0032] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only for example).
[0033] The terminal devices 101, 102, 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0034] The server 105 may be a server providing various services, such as a background management server that provides support for the content browsed by users using the terminal devices 101, 102, 103 (only for example). The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, to solve the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server may also be a server of a distributed system, or a server combined with a blockchain.
[0035] It should be noted that the action recognition method provided by the embodiments of the present disclosure can generally be executed by the terminal devices 101, 102, or 103. Correspondingly, the action recognition apparatus provided by the embodiments of the present disclosure can also be disposed in the terminal devices 101, 102, or 103.
[0036] Alternatively, the action recognition method provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the action recognition device provided by the embodiments of the present disclosure can generally be disposed in the server 105. The action recognition method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103, and / or the server 105. Correspondingly, the action recognition device provided by the embodiments of the present disclosure can also be disposed in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103, and / or the server 105.
[0037] For example, when it is necessary to recognize an action to be recognized, the terminal devices 101, 102, 103 can input a first temporal feature vector of a video sequence to be recognized including the action to be recognized and a second temporal feature vector of a template video sequence including a template action into a discriminator model, and obtain a target template action belonging to the same category as at least part of the actions represented by the first temporal feature vector. Then, the action category of the action to be recognized is determined according to the target template action. Or a server or a server cluster capable of communicating with the terminal devices 101, 102, 103, and / or the server 105 analyzes the target content and realizes determining the action category of the action to be recognized.
[0038] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the servers in
[0039] According to an embodiment of the present disclosure, the action recognition method can be implemented by using a discriminator model trained based on adversarial training.
[0040] Figure 2 Schematically shows a flowchart of a training method of a discriminator model according to an embodiment of the present disclosure.
[0041] As Figure 2 shown, the method includes operations S210 to S230.
[0042] In operation S210, a first positive sample pair is determined, the positive sample pair includes first template video data and second template video data, and both the first template video data and the second template video data include a first template action.
[0043] In operation S220, a first negative sample pair is determined, the negative sample pair includes at least one of the first template video data and the second template video data and third template video data, and the third template video data includes a second template action different from the first template action.
[0044] In operation S230, a discriminator model is trained using the first positive sample pair and the first negative sample pair.
[0045] According to an embodiment of the present disclosure, multiple segments of template videos can be collected in advance, and the first template video data, the second template video data, and the third template video data can be determined from the multiple segments of template videos. The template videos include all required actions, i.e., template actions. The first template video data and the second template video data can be determined based on two template videos including the same template action, and then the first positive sample pair can be determined. The first template video data and the third template video data, and at least one of the second template video data and the third template video data can be determined based on two template videos including different template actions, and then the first negative sample pair can be determined. The discriminator model can be trained based on the first positive sample pair and the first negative sample pair, so that the discriminator model can identify actions of the same category as the first template action.
[0046] According to an embodiment of the present disclosure, each of the first template video data, the second template video data, and the third template video data includes an annotation accurate to video frames, and the annotation can represent the current action category of each video frame.
[0047] According to an embodiment of the present disclosure, the discriminator is trained based on the positive sample pairs and negative sample pairs formed by the annotated template video data. By inputting two samples and outputting the result of whether they are of the same category, a discriminator model that can be used for action recognition can be trained.
[0048] Through the above embodiments of the present disclosure, the discriminator model is trained using positive sample pairs and negative sample pairs, which can effectively improve the robustness of the discriminator model.
[0049] The following further describes the Figure 2 method shown with specific embodiments.
[0050] According to an embodiment of the present disclosure, determining the first positive sample pair includes: generating first template video data including a first template action based on a first generation scenario. Generating second template video data including the first template action based on a second generation scenario. Wherein, the first generation scenario is different from the second generation scenario.
[0051] According to an embodiment of the present disclosure, the first generation scenario and the second generation scenario may include at least one of different background sounds, different background colors, different recording angles, different human objects performing the first template action, etc. The generation scenarios of the two template video data in the positive sample pair can be different. For example, the first template video data can be video data generated by object A in an outdoor scene, and the second template video data can be video data generated by object B in an indoor scene, etc.
[0052] Through the above embodiments of the present disclosure, by generating a first template video and a second template video with the same template actions based on different scenarios and performing adversarial training on the discriminator model, the influence of external factors on action recognition can be effectively removed, and good robustness is achieved for the environment, identity, and perspective.
[0053] According to an embodiment of the present disclosure, when it is necessary to add actions that can be recognized by the discriminator model, one or more template videos can be shot for the added actions and annotated, and the annotation can represent the current action category of each video frame. On this basis, the training method of the discriminator model may further include: in response to an indication to add a third template action different from the first template action, generating fourth template video data including the third template action and fifth template video data including the third template action as a second positive sample pair. Generating sixth template video data including a fourth template action different from the third template action. Using at least one of the fourth template video data and the fifth template video data and the sixth template video data as a second negative sample pair. Training the discriminator model using the second positive sample pair and the second negative sample pair.
[0054] According to an embodiment of the present disclosure, the third template action may be a template action that needs to be added. The third template action may be the same as the second template action. The fourth template video data and the fifth template video data may be determined from the additionally shot template videos. The sixth template video data may be obtained by additional shooting or determined according to the first template video data and the second template video data. When the third template action is different from the second template action, the sixth template video data may also be determined according to the third template video data.
[0055] According to an embodiment of the present disclosure, the fourth template video data and the fifth template video data can be determined based on the additionally shot template videos including the same added template action, and then the second positive sample pair can be determined. The fourth template video data and the sixth template video data, and at least one of the fifth template video data and the sixth template video data can be determined based on two template videos including different template actions, and then the second negative sample pair can be determined. The discriminator model can be trained according to the second positive sample pair and the second negative sample pair, so that the discriminator model can recognize actions of the same category as the third template action, that is, the added template action.
[0056] Through the above embodiments of the present disclosure, new positive sample pairs and negative sample pairs can be determined by adding and shooting template videos. By training the discriminator model, the types of actions that the discriminator model can support for recognition are increased, and the flexible scalability of the discriminator model is improved in a relatively simple implementation manner. In addition, the template acquisition and annotation steps only require single-segment video shooting and simple annotation, which can usually be completed in a few minutes, reducing the time cost caused by data acquisition, data annotation, etc.
[0057] According to an embodiment of the present disclosure, generating the fourth template video data including the third template action and the fifth template video data including the third template action includes: generating the fourth template video data including the third template action based on the third generation scenario. Generating the fifth template video data including the third template action based on the fourth generation scenario. Wherein, the third generation scenario and the fourth generation scenario are different.
[0058] According to an embodiment of the present disclosure, the third generation scenario and the fourth generation scenario may also include scenarios corresponding to at least one of different background sounds, different background colors, different recording angles, different human objects performing the first template action, etc. And it is not limited thereto.
[0059] Through the above embodiments of the present disclosure, by generating the first template video and the second template video with the same template action based on different scenarios and performing adversarial training on the discriminator model, the influence of external factors on action recognition can be effectively removed, and it has good robustness for the environment, identity, and perspective.
[0060] According to an embodiment of the present disclosure, the action recognition method can be implemented by using the discriminator model trained based on the above adversarial training.
[0061] Figure 3 A flowchart of an action recognition method according to an embodiment of the present disclosure is schematically shown.
[0062] As Figure 3 shown, the method includes operations S310~S320.
[0063] In operation S310, input the first temporal feature vector of the video sequence to be recognized including the action to be recognized and the second temporal feature vector of the template video sequence including the template action into the discriminator model to obtain the target template action that belongs to the same category as at least part of the actions represented by the first temporal feature vector.
[0064] In operation S320, determine the action category of the action to be recognized according to the target template action.
[0065] According to an embodiment of the present disclosure, the first temporal feature vector may be a depth feature vector corresponding to a part of the video sequence in the video sequence to be recognized. The second temporal feature vector may be a depth feature vector corresponding to a part of the video sequence in the template video sequence. The template video may include at least one of the first template video data, the second template video data, the fourth template video data, and the fifth template video data in the foregoing training process, or may be other template videos newly added in a new training. At least part of the action may represent a partial action corresponding to any incomplete action segment in the action to be recognized.
[0066] According to an embodiment of the present disclosure, the extraction of each depth feature vector may be based on a human feature extraction network, such as frankmocap (a 3D human pose and shape estimation algorithm) and openpose (a human pose estimation algorithm), and the video to be recognized and the template video are extracted frame by frame.
[0067] According to an embodiment of the present disclosure, in human action recognition, when the depth features to be extracted include body part features and gesture part features, since the feature dimensions of the body part features and the gesture part features are different, the two can be normalized separately and then concatenated to obtain a depth feature vector representing the entire human action.
[0068] According to an embodiment of the present disclosure, the depth feature vector extracted for the template video can be saved as a separate file to support reuse.
[0069] According to an embodiment of the present disclosure, for the action to be recognized, it can be input into a trained discriminator model together with the template action. Through this discriminator model, the target template action belonging to the same category as each part of the action to be recognized can be determined. For each partial action, one or more target template actions belonging to the same category as it can be obtained. According to the multiple target template actions belonging to the same category as the multiple partial actions, the target template action belonging to the same category as the entire action to be recognized corresponding to the multiple partial actions can be further determined, and then the action category of the action to be recognized can be determined.
[0070] Through the above embodiments of the present disclosure, using the discriminator model obtained based on adversarial training for action recognition can improve the robustness of the action recognition process.
[0071] The following combines specific embodiments to further illustrate the Figure 3 method shown.
[0072] According to an embodiment of the present disclosure, before performing the above-mentioned action recognition method, for example, it is necessary to first obtain a first temporal feature vector. The determination process of each first temporal feature vector may include: for a target video frame in a video sequence to be recognized, obtaining a first frame sequence including the target video frame and at least one video frame adjacent to the target video frame. Taking the feature vector of the first frame sequence as the first temporal feature vector.
[0073] According to an embodiment of the present disclosure, the target video frame may include one or more determined video frames in the video sequence to be recognized, or may include each video frame in the video sequence to be recognized. Considering the temporal continuity of the action, for example, in the form of a window sequence, the feature vectors of each video frame in the video to be recognized and the feature vectors of each video frame in one or more adjacent video frames thereto are spliced as the dynamic sequence feature of this video frame, that is, the first temporal feature vector. For example, for each video frame, the feature vectors of 30 frames before and after can be taken for splicing as the dynamic sequence feature of this video frame. The 30 frames before and after can form the first frame sequence.
[0074] Through the above embodiments of the present disclosure, considering the temporal continuity of the action and determining the temporal feature vector of each video frame can effectively enhance the accuracy of the recognition result.
[0075] According to an embodiment of the present disclosure, before performing the above-mentioned action recognition method, for example, it is also necessary to first obtain a second temporal feature vector. The determination process of each second temporal feature vector may include: for a target template video frame in a template video sequence, obtaining a second frame sequence including the target template video frame and template video frames adjacent to the target template video frame. Taking the feature vector of the second frame sequence as the second temporal feature vector.
[0076] According to an embodiment of the present disclosure, the target template video frame may include one or more determined template video frames in the template video sequence, or may include each template video frame in the template video sequence. For example, for each video frame in the template video, in the form of a window sequence, the feature vectors of each video frame and the feature vectors of each video frame in one or more adjacent video frames thereto are spliced as the dynamic sequence feature of this video frame, that is, the second temporal feature vector. To achieve better discrimination, for example, for each video frame in the template video, the feature vectors of 30 frames before and after can be taken for splicing as the dynamic sequence feature of this video frame. The 30 frames before and after can form the second frame sequence.
[0077] According to an embodiment of the present disclosure, since the template video includes multiple template actions, to further improve the working efficiency of the discriminator model, for example, according to the annotations in the template video, the template video segments corresponding to each category of actions can be determined, and then for each template video segment, in combination with the window sequence method, the second temporal feature vector can be determined.
[0078] Through the above embodiments of the present disclosure, considering the temporal continuity of actions and determining the temporal feature vector of each video frame can effectively enhance the accuracy of the recognition result.
[0079] According to an embodiment of the present disclosure, inputting the multiple first temporal feature vectors of the video sequence to be recognized including the action to be recognized and the second temporal feature vectors of the template video including the template actions into the discriminator model, the target template actions belonging to the same category as a partial action of the action to be recognized are obtained as follows: inputting the first temporal feature vectors and at least one second temporal feature vector into the discriminator model to obtain the target second temporal feature vectors whose similarity with the first temporal feature vectors is greater than or equal to a preset threshold. Taking the template actions represented by the target second temporal feature vectors as the target template actions.
[0080] According to an embodiment of the present disclosure, whether two actions belong to the same category can be determined, for example, according to the similarity of the temporal feature vectors corresponding to the two actions respectively, and a preset threshold can be constructed as the criterion for judging whether they are of the same category. For example, when the similarity of the temporal feature vectors corresponding to the two actions respectively is greater than or equal to the preset threshold, it can be determined that the two actions belong to the same category; when the similarity of the temporal feature vectors corresponding to the two actions respectively is less than the preset threshold, it can be determined that the two actions do not belong to the same category.
[0081] According to an embodiment of the present disclosure, the two actions may include a partial action in the action to be recognized and a labeled template action, or a partial action in the action to be recognized and a partial action in the labeled template action. When it is determined that the similarity between the first temporal feature vector corresponding to a certain partial action in the action to be recognized and the second temporal feature vector corresponding to a certain partial action in a certain labeled template action is greater than or equal to the preset threshold, the labeled template action or the partial action thereof can be taken as the target template action belonging to the same category as the partial action in the action to be recognized.
[0082] Through the above embodiments of the present disclosure, determining whether the action to be recognized and the template action belong to the same category based on the similarity of the temporal feature vectors can effectively enhance the accuracy of the discriminator model in the process of action recognition.
[0083] According to an embodiment of the present disclosure, the target template actions include multiple target template actions. Determining the action category of the action to be recognized based on the target template actions includes: determining the occurrence times of each target template action among the multiple target template actions. Taking the action category of the target template action with the most occurrence times as the action category of the action to be recognized.
[0084] According to an embodiment of the present disclosure, multiple partial actions can be correspondingly obtained according to the action to be recognized, and there may be one or more target template actions belonging to the same category as each partial action. The target template action for determining the action category of the action to be recognized can be the target template action with the most occurrence times among the multiple target template actions belonging to the same category as all the partial actions in the action to be recognized.
[0085] For example, the template actions include raising one hand, raising both hands, waving, putting both hands on the head, tying hair, etc. The first partial action in the action to be recognized, for example, is manifested as the hand stretching upward. Then, this first partial action has a high similarity with the preliminary actions of raising one hand, raising both hands, waving, putting both hands on the head, and tying hair. If this similarity is greater than a preset threshold, raising one hand, raising both hands, waving, putting both hands on the head, and tying hair can all be used as the target template actions belonging to the same category as this first partial action. Further, it is recognized that the second partial action in the action to be recognized, for example, is manifested as both hands stretching upward. Then, this second partial action has a high similarity with some partial actions of raising both hands, putting both hands on the head, and tying hair. If this similarity is greater than a preset threshold, raising both hands, putting both hands on the head, and tying hair can be used as the target template actions belonging to the same category as this second partial action. Further, it is recognized that the third partial action in the action to be recognized, for example, is manifested as both hands over the head. Then, this third partial action has a high similarity with the later action of raising both hands. If this similarity is greater than a preset threshold, raising both hands can be used as the target template action belonging to the same category as this second partial action. Since raising both hands belongs to the same category as at least three partial actions in the action to be recognized, that is, when discriminating the action to be recognized and the template, raising both hands has a high occurrence times, it can be determined based on this that the action category of the action to be recognized includes raising both hands.
[0086] Through the above embodiments of the present disclosure, based on the occurrence times of the target template actions belonging to the same category as each partial action in the action to be recognized, determining the target template action that can determine the action category of the action to be recognized, and further determining the action category of the action to be recognized can further enhance the accuracy of action recognition.
[0087] The following refers to Figures 4A to 4D , and further illustrates the above action recognition method in combination with specific embodiments.
[0088] According to an embodiment of the present disclosure, the above-mentioned action recognition method implemented based on a discriminator model mainly includes the following processes: data collection and discriminator training, feature extraction, discriminator processing, and post-processing to output results.
[0089] Figure 4A FIG. schematically shows a schematic diagram of data collection and discriminator training according to an embodiment of the present disclosure.
[0090] As Figure 4A shown, the data collection process can be implemented by collecting multiple segments of template videos 401 including the same or different template actions in the same or different scenarios. Different template actions in the template video 401 can be distinguished by the annotation 403. The action category of each part of the template action can be determined through the annotation 403. The deep feature extraction network 402 can extract the template features 404 of the template video 401. The positive sample pairs constructed by using the template features representing the same template action and the negative sample pairs constructed by using the template features representing different template actions can perform adversarial training 405 on the discriminator model 406, so that the discriminator model 406 can complete the action recognition of the action to be recognized based on the template features.
[0091] It should be noted that the relevant process of the annotation 403 can be completed in the template video 401. The template features 404 extracted by the deep feature extraction network 402 can be saved as independent files for subsequent reuse in the discriminator processing process.
[0092] Figure 4B FIG. schematically shows a schematic diagram of feature extraction according to an embodiment of the present disclosure.
[0093] As Figure 4B shown, since the template features 404 have been saved as independent files that support reuse, this part of the feature extraction process can only extract the video features 408 of the video 407 to be recognized, and this part of the extraction process can also be implemented by the deep feature extraction network 402.
[0094] It should be noted that the video features 408 can represent the first temporal feature vectors corresponding to each part of the actions in the video 407 to be recognized. The template features 404 can represent the second temporal feature vectors corresponding to each part of the actions in the template video data.
[0095] Figure 4C FIG. schematically shows a schematic diagram of discriminator processing according to an embodiment of the present disclosure.
[0096] As Figure 4CAs shown, the discriminator model 409 can take the template feature 404 and the video feature 408 as inputs, and output the similarity that the template feature and the video feature belong to the same category as the discrimination result 410. Combining with a preset threshold, one or more template features belonging to the same category as the video feature can be determined based on this discrimination result.
[0097] Figure 4D Schematically shows a schematic diagram of the post - processing output result according to an embodiment of the present disclosure.
[0098] As Figure 4D shown, the discrimination result 410 can reflect one or more template features belonging to the same category as each video feature. The post - processing module 411 can determine the action category of the action in the video to be recognized according to the occurrence times of each template feature, and output it through the output category 412.
[0099] Through the above - mentioned embodiments of the present disclosure, a lightweight, high - precision, and low - cost action recognition method is provided. Compared with the existing mainstream methods, the process time consumption is greatly reduced, and it has good scalability, supporting the flexible expansion of any newly added actions. On the premise of ensuring the recognition effect, the R & D and deployment costs are greatly reduced. This method can be applied to products related to 3D virtual digital humans, including virtual anchors, virtual customer service, virtual assistants, virtual teachers, virtual idols, etc., and can also be applied to fields such as label recognition and agenda detection, supporting the rapid iteration of products with good scalability and excellent performance.
[0100] Figure 5 Schematically shows a block diagram of a training device for a discriminator model according to an embodiment of the present disclosure.
[0101] As Figure 5 shown, the training device 500 of the discriminator model includes a first determination module 510, a second determination module 520, and a first training module 530.
[0102] The first determination module 510 is used to determine a first positive sample pair. The positive sample pair includes first template video data and second template video data. Both the first template video data and the second template video data include a first template action.
[0103] The second determination module 520 is used to determine a first negative sample pair. The negative sample pair includes at least one of the first template video data and the second template video data and third template video data. The third template video data includes a second template action different from the first template action.
[0104] The first training module 530 is used to train the discriminator model by using the first positive sample pair and the first negative sample pair.
[0105] According to an embodiment of the present disclosure, the training device of the discriminator model further includes a first generation module, a second generation module, a first definition module, and a second training module.
[0106] The first generation module is configured to generate, in response to an indication to add a third template action different from the first template action, a fourth template video data including the third template action and a fifth template video data including the third template action as a second positive sample pair.
[0107] The second generation module is configured to generate a sixth template video data including a fourth template action different from the third template action.
[0108] The first definition module is configured to use at least one of the fourth template video data and the fifth template video data and the sixth template video data as a second negative sample pair.
[0109] The second training module is configured to train the discriminator model by using the second positive sample pair and the second negative sample pair.
[0110] According to an embodiment of the present disclosure, the first determination module includes a first generation unit and a second generation unit.
[0111] The first generation unit is configured to generate a first template video data including a first template action based on a first generation scenario.
[0112] The second generation unit is configured to generate a second template video data including a first template action based on a second generation scenario.
[0113] Wherein, the first generation scenario is different from the second generation scenario.
[0114] According to an embodiment of the present disclosure, the first generation module includes a third generation unit and a fourth generation unit.
[0115] The third generation unit is configured to generate a fourth template video data including a third template action based on a third generation scenario.
[0116] The fourth generation unit is configured to generate a fifth template video data including a third template action based on a fourth generation scenario.
[0117] Wherein, the third generation scenario is different from the fourth generation scenario.
[0118] Figure 6 The block diagram of the action recognition device according to an embodiment of the present disclosure is schematically shown.
[0119] As Figure 6 shown, the action recognition device 600 includes an input module 610 and a determination module 620.
[0120] An input module 610 is configured to input a first temporal feature vector of a video sequence to be recognized including an action to be recognized and a second temporal feature vector of a template video sequence including a template action into a discriminator model, so as to obtain a target template action that belongs to the same category as at least part of the actions represented by the first temporal feature vector.
[0121] A determination module 620 is configured to determine the action category of the action to be recognized according to the target template action.
[0122] According to an embodiment of the present disclosure, the input module includes an input unit and a first definition unit.
[0123] The input unit is configured to input the first temporal feature vector and at least one second temporal feature vector into the discriminator model, so as to obtain a target second temporal feature vector whose similarity to the first temporal feature vector is greater than or equal to a preset threshold.
[0124] The first definition unit is configured to use the template action represented by the target second temporal feature vector as the target template action.
[0125] According to an embodiment of the present disclosure, the target template action includes multiple target template actions, and the determination module includes a determination unit and a second definition unit.
[0126] The determination unit is configured to determine the occurrence times of each target template action among the multiple target template actions.
[0127] The second definition unit is configured to use the action category of the target template action with the most occurrence times as the action category of the action to be recognized.
[0128] According to an embodiment of the present disclosure, the action recognition device further includes a first acquisition module and a second definition module.
[0129] The first acquisition module is configured to, for a template video frame in the video sequence to be recognized, acquire a first frame sequence including the template video frame and at least one video frame adjacent to the template video frame.
[0130] The second definition module is configured to use the feature vector of the first frame sequence as the first temporal feature vector.
[0131] According to an embodiment of the present disclosure, the action recognition device further includes a second acquisition module and a third definition module.
[0132] The second acquisition module is configured to, for a template video frame in the template video sequence, acquire a second frame sequence including the template video frame and template video frames adjacent to the template video frame.
[0133] The third definition module is configured to use the feature vector of the second frame sequence as the second temporal feature vector.
[0134] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0135] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.
[0136] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.
[0137] According to an embodiment of the present disclosure, a computer program product includes a computer program, and the computer program implements the method as described above when executed by a processor.
[0138] Figure 7 A schematic block diagram of an exemplary electronic device 700 that may be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0139] As Figure 7 shown, the device 700 includes a computing unit 701, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0140] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, such as a keyboard, mouse, etc.; output unit 707, such as various types of displays, speakers, etc.; storage unit 708, such as a disk, optical disc, etc.; and communication unit 709, such as a network card, modem, wireless communication transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0141] Computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 701 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 701 executes the various methods and processes described above, such as the training method of the discriminator model and the action recognition method. For example, in some embodiments, the training method of the discriminator model and the action recognition method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by computing unit 701, one or more steps of the training method of the discriminator model and the action recognition method described above can be executed. Alternatively, in other embodiments, computing unit 701 can be configured to execute the training method of the discriminator model and the action recognition method in any other suitable way (e.g., by means of firmware).
[0142] Various embodiments of the systems and techniques described above in this article can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs, the one or more computer programs can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a special or general-purpose programmable processor, can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0143] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0144] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronics, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0145] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer that has: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, speech input, or tactile input).
[0146] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0147] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs that run on the respective computers and have a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0148] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0149] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. An action recognition method, comprising: Inputting a first temporal feature vector of a to-be-recognized video sequence including an action to be recognized and a second temporal feature vector of a template video sequence including a template action into a discriminator model to obtain an output result of the discriminator model, where the output result includes a result of whether the action to be recognized and the template action belong to the same category; Obtaining a target template action belonging to the same category as at least part of the actions characterized by the first temporal feature vector according to the output result of the discriminator model; And Determining the action category of the action to be recognized according to the target template action; Wherein, the discriminator model is obtained through adversarial training; the method further includes: in response to the need to add a template action that the discriminator model can support for recognition, taking additional template videos for the added template action, forming a positive sample pair with two template videos including the added template action, forming a negative sample pair with a template video including the added template action and a template video not including the added template action, and performing adversarial training on the discriminator model so that the discriminator model can recognize actions of the same category as the added template action.
2. The method according to claim 1, wherein The inputting a first temporal feature vector of a to-be-recognized video sequence including an action to be recognized and a second temporal feature vector of a template video sequence including a template action into a discriminator model to obtain the output result of the discriminator model includes: Inputting the first temporal feature vector and at least one of the second temporal feature vectors into the discriminator model to obtain a target second temporal feature vector whose similarity to the first temporal feature vector is greater than or equal to a preset threshold; The obtaining a target template action belonging to the same category as at least part of the actions characterized by the first temporal feature vector according to the output result of the discriminator model includes: Taking the template action characterized by the target second temporal feature vector as the target template action.
3. The method according to claim 1, wherein The target template action includes multiple target template actions. Determining the action category of the action to be recognized according to the target template action includes: Determining the occurrence times of each of the multiple target template actions; and Taking the action category of the target template action with the most occurrences as the action category of the action to be recognized.
4. The method according to claim 1, further comprising: For a target video frame in the to-be-recognized video sequence, obtaining a first frame sequence including the target video frame and at least one video frame adjacent to the target video frame; And Taking the feature vector of the first frame sequence as the first temporal feature vector.
5. The method according to claim 1, further comprising: For a target template video frame in the template video sequence, obtaining a second frame sequence including the target template video frame and template video frames adjacent to the target template video frame; And Taking the feature vector of the second frame sequence as the second temporal feature vector.
6. The method according to claim 1, wherein, The discriminator model is obtained through the following operations for training: Determine a first positive sample pair, where the positive sample pair includes first template video data and second template video data, and both the first template video data and the second template video data include a first template action; Determine a first negative sample pair, where the negative sample pair includes at least one of the first template video data and the second template video data and third template video data, and the third template video data includes a second template action different from the first template action; and Use the first positive sample pair and the first negative sample pair to train the discriminator model.
7. The method according to claim 6, further comprising: In response to an indication to add a third template action different from the first template action, generate fourth template video data including the third template action and fifth template video data including the third template action as a second positive sample pair; Generate sixth template video data including a fourth template action different from the third template action; Use at least one of the fourth template video data and the fifth template video data and the sixth template video data as a second negative sample pair; and Use the second positive sample pair and the second negative sample pair to train the discriminator model.
8. The method according to claim 6, wherein The determining of the first positive sample pair includes: Based on a first generation scenario, generate first template video data including the first template action; and Based on a second generation scenario, generate second template video data including the first template action, where the first generation scenario is different from the second generation scenario.
9. The method according to claim 7, wherein, The generating of the fourth template video data including the third template action and the fifth template video data including the third template action includes: Based on a third generation scenario, generate fourth template video data including the third template action; and Based on a fourth generation scenario, generate fifth template video data including the third template action, where the third generation scenario is different from the fourth generation scenario.
10. An action recognition device, comprising: A model processing module, configured to input a first temporal feature vector of a to-be-recognized video sequence including a to-be-recognized action and a second temporal feature vector of a template video sequence including multiple template actions into a discriminator model, and obtain an output result of the discriminator model, where the output result includes a result of whether the to-be-recognized action and the template action are of the same category; A target template action determination module, configured to obtain a target template action that is of the same category as at least part of the actions characterized by the first temporal feature vector according to the output result of the discriminator model; and An action category determination module, configured to determine the action category of the to-be-recognized action according to the target template action; where the discriminator model is obtained through adversarial training; the device further includes: An expansion module, which is used to respond to the need to add template actions that the discriminator model can support for recognition, capture template videos for the added template actions, form positive sample pairs with two template videos including the added template actions, form negative sample pairs with the template video including the added template action and the template video not including the added template action, and perform adversarial training on the discriminator model so that the discriminator model can recognize actions of the same category as the added template action.
11. The apparatus according to claim 10, wherein The model processing module is configured to input the first temporal feature vector and at least one of the second temporal feature vectors into the discriminator model to obtain a target second temporal feature vector whose similarity to the first temporal feature vector is greater than or equal to a preset threshold; and The target template action determining module is configured to use the template action represented by the target second temporal feature vector as the target template action.
12. The device according to claim 10, wherein, The target template action includes multiple target template actions, and the determining module includes: a determining unit configured to determine the occurrence times of each of the multiple target template actions; and a second defining unit configured to use the action category of the target template action with the most occurrences as the action category of the action to be recognized.
13. The apparatus according to claim 10, further comprising: a first obtaining module configured to obtain a first frame sequence including the target video frame and at least one video frame adjacent to the target video frame for the target video frame in the video sequence to be recognized; and a second defining module configured to use the feature vector of the first frame sequence as the first temporal feature vector.
14. The apparatus according to claim 10, further comprising: a second obtaining module configured to obtain a second frame sequence including the target template video frame and the template video frames adjacent to the target template video frame for the target template video frame in the template video sequence; and a third defining module configured to use the feature vector of the second frame sequence as the second temporal feature vector.
15. The apparatus according to claim 10, further comprising: a first determining module configured to determine a first positive sample pair, the positive sample pair including first template video data and second template video data, both the first template video data and the second template video data including a first template action; a second determining module configured to determine a first negative sample pair, the negative sample pair including at least one of the first template video data and the second template video data and third template video data, the third template video data including a second template action different from the first template action; and a first training module configured to train the discriminator model by using the first positive sample pair and the first negative sample pair.
16. The apparatus according to claim 15, further comprising: A first generation module, configured to generate, in response to an indication of adding a third template action different from the first template action, fourth template video data including the third template action and fifth template video data including the third template action as a second positive sample pair; A second generation module, configured to generate sixth template video data including a fourth template action different from the third template action; A first definition module, configured to use at least one of the fourth template video data and the fifth template video data and the sixth template video data as a second negative sample pair; And A second training module, configured to train the discriminator model by using the second positive sample pair and the second negative sample pair.
17. The device according to claim 15, wherein, The first determination module includes: A first generation unit, configured to generate first template video data including the first template action based on a first generation scenario; and A second generation unit, configured to generate second template video data including the first template action based on a second generation scenario, wherein the first generation scenario is different from the second generation scenario.
18. The apparatus according to claim 16, wherein, The first generation module includes: A third generation unit, configured to generate fourth template video data including the third template action based on a third generation scenario; and A fourth generation unit, configured to generate fifth template video data including the third template action based on a fourth generation scenario, wherein the third generation scenario is different from the fourth generation scenario.
19. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-9.
20. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-9.
21. A computer program product, comprising a computer program, where the computer program, when executed by a processor, implements the method according to any one of claims 1-9.
Citation Information
Patent Citations
Video processing method and device, electronic equipment and storage medium
CN110472531A
Face anti-counterfeiting detection method and system
CN112733760A