Training method, device, computing device and storage medium for video action recognition model
By converting video action recognition into multi-label classification problem, using R(2+1)D network and multi-label classification module to process video features, the problem of poor multi-category action recognition and long video recognition in the prior art is solved, and efficient recognition of multiple actions in the video is achieved.
Patent Information
- Application Number
- CN202111591869.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-23
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-12-23
AI Technical Summary
Existing video action recognition technology is difficult to effectively process action recognition in multiple categories in the same input video, and the recognition effect of longer videos is poor.
By converting video action recognition into multi-label classification problem, the video sequence is extracted using the R(2+1)D network, and multi-label classification of the video feature sequence is combined with the multi-label classification module, the loss function is calculated and the network parameters are adjusted to train the target video action recognition model.
It realizes effective identification of multiple actions in the video, solves the problem that multiple categories of actions in the same input video are difficult to classify, and can also effectively identify longer video inputs.
Smart Images

Figure CN114266997B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of neural network technology, and in particular to a training method and device for a video action recognition model, a video action recognition method and device, a computing device and a storage medium. Background Art
[0002] Video action recognition is a hot topic in artificial intelligence today. This technology allows computers to identify and judge certain specific actions in videos, and also allows computers to better perceive and understand the story content in videos, so it is widely used in scenarios such as video classification, electronic monitoring, and advertising. Compared with images, video content is more complex and changeable, and there may be occlusion, jitter, and changes in perspective when shooting videos, which brings more difficulties to action recognition. In addition, unlike static images, videos also contain time series information. How to effectively utilize the time series information and organically combine it with spatial information is the only way to achieve video action recognition. Summary of the invention
[0003] A series of simplified concepts are introduced in the summary of the invention, which will be further described in detail in the specific embodiments. The summary of the invention does not mean to attempt to define the key features and essential technical features of the technical solution claimed for protection, nor does it mean to attempt to determine the scope of protection of the technical solution claimed for protection.
[0004] In view of the above technical problems, the present invention provides a training method, device, computing equipment and storage medium for a video action recognition model, which can effectively solve the problem of difficult classification caused by the simultaneous appearance of multiple categories in the same input video of the action recognition task, and can also effectively recognize longer video inputs.
[0005] According to one aspect of the present invention, a method for training a video action recognition model is provided, which comprises:
[0006] Sampling the sample video according to a preset sampling strategy to obtain at least two picture sequences, each of the picture sequences comprising a plurality of frames of pictures collected from the sample video and arranged in time sequence;
[0007] Extracting features of the image sequence through an R(2+1)D network to obtain video sequence features of the sample video;
[0008] Inputting the video sequence features into a multi-label classification module for processing to obtain a video action classification result, and calculating a loss function based on the video action classification result;
[0009] The R(2+1)D network and the multi-label classification module are adjusted according to the calculation result of the loss function to obtain a target video action recognition model.
[0010] According to another aspect of the present invention, a video action recognition method is provided, comprising:
[0011] Obtain the target video to be identified;
[0012] Sampling the target video according to a preset sampling strategy to obtain at least two picture sequences, each of which includes multiple frames of pictures collected from the target video and arranged in time sequence;
[0013] The picture sequence is input into a target video action recognition model trained by the training method according to the present invention to perform video action recognition.
[0014] According to another aspect of the present invention, there is provided a training device for a video action recognition model, comprising:
[0015] A picture sampling module, used for sampling the sample video according to a preset sampling strategy to obtain at least two picture sequences, each of which includes a plurality of frames of pictures collected from the sample video and arranged in time sequence;
[0016] A feature extraction module, used to extract features from the image sequence through an R(2+1)D network to obtain video sequence features of the sample video;
[0017] An action recognition module is used to input the video sequence features into a multi-label classification module for processing to obtain a video action classification result, and calculate a loss function based on the video action classification result;
[0018] An adjustment module is used to adjust the R(2+1)D network and the multi-label classification module according to the calculation result of the loss function to obtain a target video action recognition model.
[0019] According to another aspect of the present invention, there is provided a video action recognition device, comprising:
[0020] An acquisition module, used to acquire a target video to be identified;
[0021] A sampling module, used for sampling the target video according to a preset sampling strategy to obtain at least two picture sequences, each of which includes a plurality of frames of pictures collected from the target video and arranged in time sequence;
[0022] The recognition module is used to input the picture sequence into a target video action recognition model trained by the training method according to the present invention to perform video action recognition.
[0023] According to another aspect of the present invention, a computing device is provided, comprising: a processor, and a memory storing computer program instructions; the processor reads and executes the computer program instructions to implement the video action recognition model training method or video action recognition method according to one aspect of the present invention.
[0024] According to another aspect of the present invention, a computer storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, a training method for a video action recognition model or a video action recognition method according to one aspect of the present invention is implemented.
[0025] According to the training method and apparatus, computing device and storage medium of the video action recognition model of the present invention, by converting video action recognition into a multi-label classification problem, a multi-label classification module is used to perform multi-label classification on the video feature sequence, so that multiple actions contained in the video can be recognized, and longer video inputs can also be effectively recognized. According to the video action recognition method and apparatus of the embodiment of the present invention, since the video action recognition model is obtained by the training method of the present invention, the problem that multiple categories appear simultaneously in the same input video of the action recognition task, resulting in difficulty in classification, can be effectively solved, and longer video inputs can also be effectively recognized. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0027] Figure 1 is a schematic flow chart of a method for training a video action recognition model according to an embodiment of the present invention;
[0028] Figure 2 An example of a training process of a video action recognition model according to an embodiment of the present invention;
[0029] Figure 3 is a schematic flow chart of a video action recognition method according to an embodiment of the present invention;
[0030] Figure 4 Schematic structural block diagram of a training device for a video action recognition model according to an embodiment of the present invention
[0031] Figure 5 is a schematic structural block diagram of a video action recognition device according to an embodiment of the present invention; and
[0032] Figure 6 It is a structural schematic diagram of a computing device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0033] The technical scheme in the embodiment of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiment of the present invention. In the following description, a large number of specific details are given to provide a more thorough understanding of the present invention. However, it is obvious to those skilled in the art that the embodiment of the present invention can be implemented without one or more of these details. In other examples, in order to avoid confusion with the embodiment of the present invention, some technical features well known in the art are not described.
[0034] Unless otherwise specifically stated, the relative arrangement, numerical expressions and numerical values of the parts and steps described in these embodiments do not limit the scope of the present invention. Meanwhile, it should be understood that for ease of description, the size of each part shown in the accompanying drawings is not drawn according to the actual proportional relationship.
[0035] Techniques, methods, and apparatus known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such techniques, methods, and apparatus should be considered part of the specification. In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limiting. Therefore, other examples of the exemplary embodiments may have different values. It should be noted that similar reference numerals and letters represent similar items in the following figures, and therefore, once an item is defined in one figure, it does not need to be further discussed in subsequent figures.
[0036] In order to make the purpose, technical scheme and advantages of the present invention more obvious, the exemplary embodiments according to the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments of the present invention, and it should be understood that the present invention is not limited to the exemplary embodiments described herein. Based on the embodiments of the present invention described in the present invention, all other embodiments obtained by those skilled in the art without creative work should fall within the protection scope of the present invention.
[0037] The action recognition method based on 3D convolutional neural network uses 3D convolution module to replace the original 2D convolution, and uses the extra dimension to process timing information. Although the action recognition model designed based on the open source action data set (usually only one or two seconds) can better complete the action recognition task under the conditions set by the data set, the recognition effect for longer videos is not good, and in reality, some actions are often completed in more than one or two seconds. Therefore, the present invention sets the completion action duration of the action recognition task in the actual environment to 10s.
[0038] When defining action categories, some categories often co-occur, such as a fight video often contains running and chasing actions. In addition, after lengthening the video to be identified, some videos often contain multiple different actions. Therefore, the present invention defines it as a multi-label classification problem.
[0039] Based on the above description, the embodiment of the present invention provides a training method, apparatus, computing device and storage medium for a video action recognition model, which are described in detail below with reference to the accompanying drawings.
[0040] First, the training method of the video action recognition model provided by the embodiment of the present invention is introduced.
[0041] Figure 1 4 is a schematic flowchart of a method 100 for training a video action recognition model according to an embodiment of the present invention. Figure 2 This is an example of a training process of a video action recognition model according to an embodiment of the present invention.
[0042] Please refer to Figure 1 and Figure 2 The video action recognition model training method 100 disclosed in the embodiment of the present invention includes:
[0043] Step S101 : sampling a sample video according to a preset sampling strategy to obtain at least two picture sequences, each of which includes a plurality of frames of pictures collected from the sample video and arranged in time sequence.
[0044] Exemplarily, in an embodiment of the present invention, the sampling strategy includes a continuous sampling strategy and a frame skipping sampling strategy. The continuous sampling strategy is to randomly collect multiple frames of pictures arranged in time sequence at different starting points of the sample video. The frame skipping sampling strategy is to randomly collect multiple frames of pictures arranged in time sequence at different starting points of the sample video, with n frames of pictures between adjacent pictures, where n is a natural number greater than 0.
[0045] In this embodiment, the same video is sampled multiple times, for example, twice, three times or more. As an example, taking the continuous sampling strategy as an example, a sequence of 64 consecutive frames of images with different starting points is randomly sampled twice for the same sample video to obtain two sets of sequences of 64 consecutive frames of images with different starting points.
[0046] For example, in this embodiment, the sample video is a video with a length of less than 10 seconds. In order to enable the trained model to recognize actions of a longer duration, the sample video includes a video with a length of about 10 seconds.
[0047] Step S102: extract features from the image sequence through an R(2+1)D network to obtain video sequence features of the sample video.
[0048] After the sample video is sampled to obtain a picture sequence in S101, the picture sequence is input into the R(2+1)D network for feature extraction to obtain the video sequence features of the sample video.
[0049] It should be known that the R(2+1)D network is a basic model in the field of video recognition. Its basic structure is consistent with the resnet network using 2D convolution. The difference is that its convolution operation has an additional dimension to represent the time series. Therefore, the convolution in this network is (2+1)D, and R stands for resnet.
[0050] It should be understood that the video sequence features include feature vectors corresponding to each picture in each picture sequence, and these feature vectors are arranged in a corresponding time sequence.
[0051] For example, in one embodiment of the present invention, in S101, a sequence of 64 consecutive frames of a sample video at different starting points is randomly collected twice to obtain 128 pictures. Figure 2 The number of corresponding video feature sequences F1 to Fn is 128. The 128 feature vectors are two groups of feature vectors arranged in time sequence.
[0052] Step S103: Input the video sequence features into a multi-label classification module for processing to obtain a video action classification result, and calculate a loss function based on the video action classification result.
[0053] Exemplarily, in an embodiment of the present invention, the multi-label classification module adopts a BERT network. The multi-label classification module includes a position encoding network, a multi-head attention and a position-aware feedforward network. The position encoding network is used to encode the video feature sequence F1 to Fn and reduce the dimension of the video sequence features so that the video sequence features are suitable for processing by the multi-label classification module. The multi-head attention and position-aware feedforward network are used to obtain the video action classification result based on the video sequence features after the dimension reduction and the embedded classification feature vector.
[0054] like Figure 2 As shown, the position encoding network processes the input video feature sequence F1 to Fn to obtain corresponding vectors X1 to Xn, and the multi-head attention and position-aware feedforward network processes the video sequence features X1 to Xn after reducing the dimension to obtain outputs Y1 to Yn.
[0055] Furthermore, since the BERT network does not have a downsampling operation, there are as many outputs as there are inputs. In order to apply it to classification tasks, an additional classification embedded classification feature vector Xcls is added, and its corresponding output is the video action classification result Ycls.
[0056] Exemplarily, in an embodiment of the present invention, an improved one-hot encoding is used to represent category information, that is, a one-dimensional vector is used to represent category labels, such as [1,0,1,...,0] indicates that the category labels of the current video action are 1 and 3, that is, the video contains two actions represented by labels 1 and 3.
[0057] Furthermore, in an embodiment of the present invention, the loss function is a modified binary cross entropy (BCE) loss function. This is because the binary cross entropy loss function is widely used in multi-label classification, and BCE regards multi-label classification as a series of binary classification tasks. For multi-label tasks, simply resampling different samples by label cannot ensure that the positive and negative samples of a single label in the final training set are uniformly distributed due to the label co-occurrence problem, and will also cause a sharp increase in the number of negative samples. Therefore, the BCE loss function is modified as follows:
[0058] Assume that for the sample x k , whose corresponding true label is y k ,y k =1 means y k Contains the tag i.
[0059] For label i, the sampling frequency can be expressed as Where C represents the total number of categories, n i Represents the total number of samples belonging to label i.
[0060] For the sample x k For example, it will be k The corresponding label is repeatedly sampled, so its sampling frequency can be expressed as
[0061] So define the rebalancing weight factor In order to prevent this factor from being 0 under certain conditions, it is improved to obtain
[0062] At this point, the rebalancing BCE function is as follows:
[0063]
[0064] in, Represents the classification result
[0065] Furthermore, in order to prevent the situation where the number of negative samples in some categories is still much larger than the number of positive samples after sample rebalancing, a corresponding bias term v is added to the prediction results of each category. i , so the final loss function is:
[0066] L(x k ,yk )
[0067] Step S104, adjusting the R(2+1)D network and the multi-label classification module according to the calculation result of the loss function to obtain a target video action recognition model.
[0068] After the loss function is calculated in S103, the network parameters of the R(2+1)D network and the multi-label classification module, such as the weight size, etc., can be adjusted according to the calculation result of the loss function. The training is continued until the calculation result of the loss function reaches the set threshold, thereby obtaining the target video action recognition model, which includes the R(2+1)D network and the multi-label classification module.
[0069] Furthermore, in an embodiment of the present invention, for the sample videos of the same batch, after completing one training, a new training adopts a sampling strategy different from that of the previous training. Exemplarily, for the sample videos of the same batch, the sample videos are sampled by first adopting a continuous sampling strategy and then a frame skipping sampling strategy. This is because the original sampling method may not be able to capture the complete action cycle. Therefore, after the first version of the model is trained using the initial sampling strategy, the original strategy of collecting continuous frames is modified to frame skipping (the number of frame skipping is n) to continue fine-tuning the obtained model, that is, when n=1, the model parameters of n=0 are inherited to continue training, and when n=2, the model parameters of n=1 are inherited to continue training, and so on until an optimal model is obtained.
[0070] According to the training method of the video action recognition model of the present invention, by converting video action recognition into a multi-label classification problem, a multi-label classification module is used to perform multi-label classification on the video feature sequence, so that multiple actions contained in the video can be identified, and longer video inputs can also be effectively recognized.
[0071] Figure 3 A schematic flow chart of a video action recognition method provided by an embodiment of the present invention.
[0072] like Figure 3 The video action recognition method 300 provided by the embodiment of the present invention includes:
[0073] Step S301, obtaining a target video to be identified.
[0074] Exemplarily, the target video is a video with a length of less than 10 seconds.
[0075] Step S302: sampling the target video according to a preset sampling strategy to obtain at least two picture sequences, each of which includes a plurality of frames of pictures collected from the target video and arranged in time sequence.
[0076] Exemplarily, the sampling strategy may adopt the aforementioned continuous sampling strategy or frame skipping sampling strategy. The number of pictures included in each picture sequence may be, for example, the aforementioned 64 pictures, or may be another number determined as required.
[0077] Step S303: input the picture sequence into a target video action recognition model trained by the training method according to the present invention to perform video action recognition.
[0078] Exemplarily, the target video action recognition model includes:
[0079] R(2+1)D network, used for extracting features from the picture sequence to obtain video sequence features of the target video;
[0080] The multi-label classification module is used to obtain a video action classification result based on the video sequence features and the embedded classification feature vector.
[0081] The video action recognition method according to the embodiment of the present invention has better recognition performance because the target video action recognition model trained by the training method of the present invention is used to perform video action recognition.
[0082] Figure 4 FIG. 4 is a schematic structural block diagram of a training device 400 for a video action recognition model according to an embodiment of the present invention. Figure 4 A video action recognition model training device 400 according to an embodiment of the present invention is described.
[0083] Please refer to Figure 4 The training device 400 of the video action recognition model according to the embodiment of the present invention includes a picture sampling module 410, a feature extraction module 420, an action recognition module 430 and an adjustment module 440.
[0084] The picture sampling module 410 is used to sample the sample video according to a preset sampling strategy to obtain at least two picture sequences, each of which includes multiple frames of pictures collected from the sample video and arranged in time sequence. Figure 1 Step S101 in the training method of the video action recognition model described above, the detailed description of this process can be found in the above-mentioned combination Figure 1-Figure 2 The description will not be repeated here.
[0085] The feature extraction module 420 is used to extract features from the image sequence through the R(2+1)D network to obtain the video sequence features of the sample video. Figure 1 Step S102 in the training method of the video action recognition model described above, the detailed description of this process can be found in the above-mentioned combination Figure 1-Figure 2The description will not be repeated here.
[0086] The action recognition module 430 is used to input the video sequence features into the multi-label classification module for processing to obtain a video action classification result, and calculate the loss function based on the video action classification result. Figure 1 Step S103 in the training method of the video action recognition model described above, the detailed description of this process can be found in the above-mentioned combination Figure 1-Figure 2 The description will not be repeated here.
[0087] The adjustment module 440 is used to adjust the R(2+1)D network and the multi-label classification module according to the calculation result of the loss function to obtain the target video action recognition model. Figure 1 Step S104 in the training method of the video action recognition model described above, the detailed description of this process can be found in the above-mentioned combination Figure 1-Figure 2 The description will not be repeated here.
[0088] Figure 4 Each module / unit in the training device 400 of the video action recognition model has the following functions: Figure 1 The functions of each step in the process can achieve the corresponding technical effects, which will not be described in detail here for the sake of brevity.
[0089] Figure 5 FIG. 5 is a schematic structural block diagram of a video action recognition device 500 according to an embodiment of the present invention. Figure 5 A video action recognition device 500 according to an embodiment of the present invention is described.
[0090] Please refer to Figure 5 , the video action recognition device 500 according to the embodiment of the present invention includes an acquisition module 510 , a sampling module 520 and a recognition module 530 .
[0091] The acquisition module 510 is used to acquire the target video to be identified. Figure 3 Step S301 in the video action recognition method described above, the detailed description of the process can be found in the above-mentioned combination Figure 3 The description will not be repeated here.
[0092] The sampling module 520 is used to sample the target video according to a preset sampling strategy to obtain at least two picture sequences, each of which includes multiple frames of pictures collected from the target video and arranged in time sequence. Figure 3 Step S302 in the training method of the video action recognition model described above, the detailed description of this process can be found in the above-mentioned combination Figure 3 The description will not be repeated here.
[0093] The recognition module 530 is used to input the picture sequence into the target video action recognition model trained by the training method described in the embodiment of the present invention to perform video action recognition. Figure 3 Step S303 in the training method of the video action recognition model described above, the detailed description of this process can be found in the above-mentioned combination Figure 3 The description will not be repeated here.
[0094] Figure 5 Each module / unit in the video action recognition device 500 has the following functions: Figure 3 The functions of each step in the process can achieve the corresponding technical effects, which will not be described in detail here for the sake of brevity.
[0095] Figure 6 A schematic diagram of the hardware structure of a computing device provided by an embodiment of the present invention is shown.
[0096] The computing device 600 may include a processor 601 and a memory 602 storing computer program instructions.
[0097] Specifically, the processor 601 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiment of the present invention.
[0098] The memory 602 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 602 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In one example, the memory 602 may include a removable or non-removable (or fixed) medium, or the memory 602 is a non-volatile solid-state memory. The memory 602 may be inside or outside the integrated gateway disaster recovery device.
[0099] In one example, the memory 602 may be a read-only memory (ROM). In one example, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.
[0100] The memory 602 may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Thus, in general, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.
[0101] The processor 601 reads and executes the computer program instructions stored in the memory 602 to implement Figure 1 The method / steps S101 to S104 in the illustrated embodiment, and Figure 3 The method / steps S301 to S303 in the embodiment shown in the figure are achieved Figure 1 and Figure 3 The corresponding technical effects achieved by executing the methods / steps in the example shown are not repeated here for the sake of brevity.
[0102] The processor 601 reads and executes the computer program instructions stored in the memory 602 to implement Figure 4 The training device 400 of the video action recognition model in the embodiment shown, as well as the picture sampling module 410, the feature extraction module 420, the action recognition module 430 and the adjustment module 440, achieve Figure 4 The corresponding technical effects achieved by the device in the example shown, as well as the video action recognition device 500, the acquisition module 510, the sampling module 520 and the recognition module 530, and achieving Figure 5 The corresponding technical effects achieved by the device in the illustrated example will not be described in detail here for the sake of brevity.
[0103] In one example, the computing device 600 may further include a communication interface 603 and a bus 610. Figure 6 As shown, the processor 601, the memory 602, and the communication interface 603 are connected via a bus 610 and communicate with each other.
[0104] The communication interface 603 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiment of the present invention.
[0105] Bus 610 includes hardware, software or both, and couples the components of online data traffic billing equipment to each other. For example, but not limitation, the bus may include Accelerated Graphics Port (AGP) or other graphics bus, Enhanced Industry Standard Architecture (EISA) bus, Front Side Bus (FSB), Hyper Transport (HT) interconnection, Industry Standard Architecture (ISA) bus, InfiniBand interconnection, Low Pin Count (LPC) bus, Memory bus, Micro Channel Architecture (MCA) bus, Peripheral Component Interconnect (PCI) bus, PCI-Express (PCI-X) bus, Serial Advanced Technology Attachment (SATA) bus, Video Electronics Standards Association Local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 610 may include one or more buses. Although the embodiments of the present invention describe and illustrate specific buses, the present invention considers any suitable bus or interconnection.
[0106] The computing device 600 can execute the training method of the video action recognition model in the embodiment of the present invention, thereby realizing the combination of Figure 1 The computing device 600 can also execute the video action recognition method in the embodiment of the present invention, thereby realizing the combination of Figure 3 Video action recognition method described
[0107] In addition, according to an embodiment of the present invention, a storage medium is also provided, on which program instructions are stored, and when the program instructions are run by a computer or a processor, the training method of the video action recognition model of the embodiment of the present invention and the corresponding steps of the video action recognition method are used to execute, and the training device of the video action recognition model and the corresponding unit or module of the video action recognition device according to the embodiment of the present invention are used to implement. The storage medium may, for example, include a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a USB memory, or any combination of the above storage media. The computer-readable storage medium may be any combination of one or more computer-readable storage media.
[0108] In one embodiment, when the computer program instructions are executed by a computer, they can implement the various functional modules in the training device of the video action recognition model and the video action recognition device according to the embodiment of the present invention, and / or can execute the training method of the video action recognition model and the video action recognition method according to the embodiment of the present invention.
[0109] In one embodiment, the computer program instructions perform the following steps when executed by a computer: sampling a sample video according to a preset sampling strategy to obtain at least two image sequences, each of which includes multiple frames of images collected from the sample video and arranged in time sequence; extracting features of the image sequence through an R(2+1)D network to obtain video sequence features of the sample video; inputting the video sequence features into a multi-label classification module for processing to obtain a video action classification result, and calculating a loss function based on the video action classification result; adjusting the R(2+1)D network and the multi-label classification module according to the calculation result of the loss function to obtain a target video action recognition model.
[0110] According to the training method and apparatus, computing device and storage medium of the video action recognition model of the present invention, by converting video action recognition into a multi-label classification problem, a multi-label classification module is used to perform multi-label classification on the video feature sequence, so that multiple actions contained in the video can be recognized, and longer video inputs can also be effectively recognized. According to the video action recognition method and apparatus of the embodiment of the present invention, since the video action recognition model is obtained by the training method of the present invention, the problem that multiple categories appear simultaneously in the same input video of the action recognition task, resulting in difficulty in classification, can be effectively solved, and longer video inputs can also be effectively recognized.
[0111] According to the training method, device, computing device and storage medium of the video action recognition model of the present invention, a plurality of different open source data sets are used to train the video action recognition model by adding a supervision network to provide supervision information, so as to expand the number of training samples without increasing the complexity of the recognition network, and effectively improve the recognition performance of the video action recognition network. According to the video action recognition method and device of the embodiment of the present invention, the video action recognition model is obtained by adopting the training method of the present invention, so it has better recognition performance.
[0112] Although example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above example embodiments are merely exemplary and are not intended to limit the scope of the present invention thereto. Various changes and modifications may be made therein by one of ordinary skill in the art without departing from the scope and spirit of the present invention. All such changes and modifications are intended to be included within the scope of the present invention as required by the appended claims.
[0113] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0114] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.
[0115] In the description provided herein, a large number of specific details are described. However, it is understood that embodiments of the present invention can be practiced without these specific details. In some instances, well-known methods, structures and techniques are not shown in detail so as not to obscure the understanding of this description.
[0116] Similarly, it should be understood that in order to streamline the present invention and help understand one or more of the various inventive aspects, in the description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, the method of the present invention should not be interpreted as reflecting the following intention: the claimed invention requires more features than the features explicitly stated in each claim. More specifically, as reflected in the corresponding claims, the inventive point is that the corresponding technical problem can be solved with less than all the features of a single disclosed embodiment. Therefore, the claims following the specific embodiment are hereby expressly incorporated into the specific embodiment, wherein each claim itself serves as a separate embodiment of the present invention.
[0117] It will be understood by those skilled in the art that, except for mutually exclusive features, all features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device disclosed in this specification may be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) may be replaced by an alternative feature that provides the same, equivalent or similar purpose.
[0118] In addition, those skilled in the art will appreciate that, although some embodiments described herein include certain features included in other embodiments but not other features, the combination of features of different embodiments is meant to be within the scope of the present invention and form different embodiments. For example, in the claims, any one of the claimed embodiments may be used in any combination.
[0119] The various component embodiments of the present invention may be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. It should be understood by those skilled in the art that a microprocessor or digital signal processor (DSP) may be used in practice to implement some or all of the functions of some modules in the article analysis device according to an embodiment of the present invention. The present invention may also be implemented as a device program (e.g., a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present invention may be stored on a computer-readable medium, or may be in the form of one or more signals. Such a signal may be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0120] It should be noted that the above embodiments illustrate the present invention rather than limit it, and that those skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbol between brackets shall not be construed as a limitation on the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "one" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention may be implemented by means of hardware comprising a number of different elements and by means of a suitably programmed computer. In a unit claim enumerating a number of devices, several of these devices may be embodied by the same hardware item. The use of the words first, second, and third, etc., does not indicate any order. These words may be interpreted as names.
[0121] The above is only a specific embodiment of the present invention or an explanation of a specific embodiment. The protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. The protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A training method for a video action recognition model, characterized in that: include: Sampling the sample video according to a preset sampling strategy to obtain at least two picture sequences, each of the picture sequences comprising a plurality of frames of pictures collected from the sample video and arranged in time sequence; The sampling strategy includes a continuous sampling strategy and a frame skipping sampling strategy. For the sample videos of the same batch, after completing one training, a new training is performed using a sampling strategy different from that of the previous training; For the sample videos of the same batch, the sample videos are sampled by first adopting a continuous sampling strategy and then a frame skipping sampling strategy; Extracting features of the image sequence through an R(2+1)D network to obtain video sequence features of the sample video; Inputting the video sequence features into a multi-label classification module for processing to obtain a video action classification result, and calculating a loss function based on the video action classification result; The loss function is: in, x k is a sample video, y k For x k The corresponding true label, C represents the total number of categories, and represents the total number of samples belonging to label i. Represents the classification result, v i The bias item corresponding to the prediction result of each category; The R(2+1)D network and the multi-label classification module are adjusted according to the calculation result of the loss function to obtain a target video action recognition model.
2. The method according to claim 1, characterized in that: The multi-label classification module comprises: A position encoding network, used for encoding the video sequence features and reducing the dimension of the video sequence features; A multi-head attention and position-aware feedforward network is used to obtain the video action classification result based on the video sequence features after dimensionality reduction and the embedded classification feature vector.
3. The method according to any one of claims 1 to 2, characterized in that: The sampling strategy includes a continuous sampling strategy and a frame skipping sampling strategy. The continuous sampling strategy is to randomly collect multiple consecutive frames of pictures arranged in time sequence at different starting points of the sample video; The frame skipping sampling strategy is to randomly collect multiple frames of pictures arranged in time sequence from different starting points of the sample video, with n frames of pictures between adjacent pictures, where n is a natural number greater than 0.
4. A video action recognition method, characterized in that: include: Obtain the target video to be identified; Sampling the target video according to a preset sampling strategy to obtain at least two picture sequences, each of which includes multiple frames of pictures collected from the target video and arranged in time sequence; The sampling strategy includes a continuous sampling strategy and a frame skipping sampling strategy. For the sample videos of the same batch, after completing one training, a new training is performed using a sampling strategy different from that of the previous training; For the sample videos of the same batch, the sample videos are sampled by first adopting a continuous sampling strategy and then a frame skipping sampling strategy; The picture sequence is input into a target video action recognition model trained by the training method described in any one of claims 1 to 3 to perform video action recognition.
5. The video action recognition method according to claim 4, characterized in that: The target video action recognition model includes: R(2+1)D network, used for extracting features from the picture sequence to obtain video sequence features of the target video; The multi-label classification module is used to obtain a video action classification result based on the video sequence features and the embedded classification feature vector.
6. A training device for a video action recognition model, characterized in that: include: The picture sampling module is used to sample the sample video according to a preset sampling strategy to obtain at least two picture sequences, each of which includes multiple frames of pictures collected from the sample video and arranged in time sequence: The sampling strategy includes a continuous sampling strategy and a frame skipping sampling strategy. For the sample videos of the same batch, after completing one training, a new training is performed using a sampling strategy different from that of the previous training; For the sample videos of the same batch, the sample videos are sampled by first adopting a continuous sampling strategy and then a frame skipping sampling strategy; A feature extraction module, used to extract features from the image sequence through an R(2+1)D network to obtain video sequence features of the sample video; An action recognition module is used to input the video sequence features into a multi-label classification module for processing to obtain a video action classification result, and calculate a loss function based on the video action classification result; The loss function is: in, x k is a sample video, y k For x k The corresponding true label, C represents the total number of categories, and represents the total number of samples belonging to label i. Represents the classification result, v i The bias item corresponding to the prediction result of each category; An adjustment module is used to adjust the R(2+1)D network and the multi-label classification module according to the calculation result of the loss function to obtain a target video action recognition model.
7. A video action recognition device, characterized in that: include: An acquisition module, used to acquire a target video to be identified; A sampling module, used for sampling the target video according to a preset sampling strategy to obtain at least two picture sequences, each of which includes a plurality of frames of pictures collected from the target video and arranged in time sequence; The sampling strategy includes a continuous sampling strategy and a frame skipping sampling strategy. For the sample videos of the same batch, after completing one training, a new training is performed using a sampling strategy different from that of the previous training; For the sample videos of the same batch, the sample videos are sampled by first adopting a continuous sampling strategy and then a frame skipping sampling strategy; A recognition module is used to input the picture sequence into a target video action recognition model trained by the training method according to any one of claims 1 to 3 to perform video action recognition.
8. The video action recognition device according to claim 7, characterized in that: The target video action recognition model includes: The R(2+1)D network is used to extract features from the image sequence to obtain a video sequence multi-label classification module for the target video, and is used to obtain a video action classification result based on the video sequence features and the embedded classification feature vector.
9. A computing device, characterized in that The device includes: a processor, and a memory storing computer program instructions: the processor reads and executes the computer program instructions to implement the training method of the video action recognition model as described in any one of claims 1-3, or the video action recognition method as described in claim 4 or 5.
10. A computer storage medium, characterized in that: The computer storage medium stores computer program instructions, which, when executed by a processor, implement the video action recognition model training method as described in any one of claims 1 to 3, or the video action recognition method as described in claim 4 or 5.