A method for action recognition, a terminal, and a storage medium
By acquiring the video feature vector and using attention dynamic threshold and weighting processing, the problem of insufficient recognition of fuzzy video clips in the weakly supervised action positioning method is solved, and action recognition with higher accuracy is achieved.
Patent Information
- Application Number
- CN202111534245.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-15
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2041-12-15
AI Technical Summary
The existing weakly supervised action positioning methods can only identify the most recognizable action and background clips in the video, and cannot correctly identify a large number of fuzzy video clips.
By obtaining the feature vector representation of the target video, the attention weight of the foreground image frame is determined based on the preset attention dynamic threshold, and combined with the attention mechanism and weighting processing, the probability of the action contained in the video is identified and inputted to the action recognition network for action recognition.
The reliability of image action recognition and the accuracy of video action recognition results are improved, and the actions in the video can be more accurately identified and the foreground and background can be distinguished.
Smart Images

Figure CN114360053B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of machine vision, and particularly relates to an action recognition method, a terminal, and a storage medium. Background Art
[0002] With the rapid growth of the number of videos, computer vision technology has received extensive attention from researchers. The technology of extracting frames with actions from uncropped videos to achieve action localization tasks has been widely applied in the fields of video surveillance, autonomous driving, video description, video search, human-computer interaction, automatic sports commentary, patient monitoring, etc.
[0003] For the existing technology, the strongly supervised action localization method has made great progress, but it requires manual annotation of each action and the time when each action occurs, which is very time-consuming and laborious, and it is difficult to be applicable to most real-world scenarios. Therefore, the weakly supervised method has emerged. Weakly supervised action localization only requires video-level labels, which reduces the annotation cost of such video data a lot and avoids manual annotation deviation.
[0004] However, in the weakly supervised action localization method, for each video, usually a certain number of video frames are selected, and an effective classification task is used to determine whether a video segment is an action segment to achieve the localization task. It can only identify the most distinguishable actions and background segments in the video. However, in the large number of videos created in real life, in addition to the most distinguishable video segments, there are also a large number of blurred and interfering video segments that cannot be correctly recognized. Summary of the Invention
[0005] The embodiments of this application provide an action recognition method, a terminal, and a storage medium to solve the problem that the existing weakly supervised action localization method can only identify the most distinguishable actions and background segments in the video, and a large number of blurred video segments cannot be correctly recognized.
[0006] The first aspect of the embodiments of this application provides an action recognition method, including:
[0007] Obtain the feature vector representation of the target video;
[0008] Determine the attention weight of the foreground image frames in the target video according to a preset attention dynamic threshold, where the attention weight is used to indicate the probability that an image frame contains an action;
[0009] Obtain a target feature vector according to the attention weight of the foreground image frames and the feature vector representation;
[0010] Input the target feature vector into an action recognition network for action recognition to obtain the action recognition result in the target video.
[0011] In a second aspect of the embodiments of the present application, an action recognition device is provided, including:
[0012] A first acquisition module, configured to acquire a feature vector representation of a target video;
[0013] A determination module, configured to determine an attention weight of a foreground image frame in the target video according to a preset attention dynamic threshold, where the attention weight is used to indicate the probability of an action in the image frame;
[0014] A second acquisition module, configured to obtain a target feature vector according to the attention weight of the foreground image frame and the feature vector representation;
[0015] An action recognition module, configured to input the target feature vector into an action recognition network for action recognition to obtain an action recognition result in the target video.
[0016] In a third aspect of the embodiments of the present application, a terminal is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method described in the first aspect are implemented.
[0017] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0018] In a fifth aspect of the present application, a computer program product is provided. When the computer program product runs on a terminal, the terminal is enabled to execute the steps of the method described in the first aspect above.
[0019] As can be seen from the above, in the embodiments of the present application, by acquiring a feature vector representation of a target video, determining an attention weight of a foreground image frame in the target video according to a preset attention dynamic threshold, and obtaining a target feature vector according to the attention weight of the foreground image frame and the feature vector representation, and inputting the target feature vector into an action recognition network for action recognition to obtain an action recognition result in the target video. This process combines an attention mechanism, introduces an attention dynamic threshold to identify foreground image frames in the video, realizes the distinction between foreground and background, and combines the attention weight of the foreground image frame with the feature vector representation to implement attention weighting processing, so as to obtain a target feature vector after attention weighting processing and realize the final action recognition. This process fuses the attention mechanism and the attention weighting mechanism on the basis of the feature sequence. After excluding irrelevant image frames, it highlights the image features, improves the reliability of image action recognition, and ensures the accuracy of the video action recognition result. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 is the flow of an action recognition method provided by an embodiment of the present application Figure 1 ;
[0022] Figure 2 is the framework structure diagram for performing action recognition based on a target video provided by an embodiment of the present application;
[0023] Figure 3 is the flow of an action recognition method provided by an embodiment of the present application Figure 2 ;
[0024] Figure 4 is the structure diagram of an action recognition device provided by an embodiment of the present application;
[0025] Figure 5 is the structure diagram of a terminal provided by an embodiment of the present application. Detailed implementation manners
[0026] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0027] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0028] It should also be understood that the terms used in this specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0029] It should also be further understood that the term "and / or" used in the specification and appended claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0030] As used in this specification and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrases "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.
[0031] In a specific implementation, the terminal described in the embodiments of this application includes, but is not limited to, other portable devices such as mobile phones, laptop computers, or tablet computers having a touch-sensitive surface (e.g., a touch screen display and / or a touchpad). It should also be understood that in some embodiments, the device is not a portable communication device, but a desktop computer having a touch-sensitive surface (e.g., a touch screen display and / or a touchpad).
[0032] In the following discussion, a terminal including a display and a touch-sensitive surface is described. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, a mouse, and / or a joystick.
[0033] The terminal supports various applications, such as one or more of the following: a drawing application, a presentation application, a word processing application, a website creation application, a disc burning application, a spreadsheet application, a game application, a telephone application, a video conferencing application, an email application, an instant messaging application, an exercise support application, a photo management application, a digital camera application, a digital video camera application, a web browsing application, a digital music player application, and / or a digital video player application.
[0034] Various applications that can be executed on the terminal can use at least one common physical user interface device such as a touch-sensitive surface. One or more functions of the touch-sensitive surface and the corresponding information displayed on the terminal can be adjusted and / or changed between applications and / or within the corresponding applications. Thus, the common physical architecture of the terminal (e.g., the touch-sensitive surface) can support various applications with a user interface that is intuitive and transparent to the user.
[0035] It should be understood that the sequence numbers of the steps in this embodiment do not indicate the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0036] In order to illustrate the technical solutions described in the present application, the following will be described through specific embodiments.
[0037] See Figure 1 , Figure 1 which is a flowchart of an action recognition method provided by an embodiment of the present application Figure 1 As Figure 1 shown, an action recognition method includes the following steps:
[0038] Step 101, obtaining a feature vector representation of a target video.
[0039] For different applications, the objects to be recognized are different. Some are voices, some are images, and some are sensor data. However, they all have corresponding digital representation forms in the computer. Usually, we convert them into a feature vector and then input it into a neural network.
[0040] Each data input into the neural network is called a feature. The image represents the image feature with a vector of a set dimension to obtain a feature vector representation, and subsequent processing operations are performed based on this feature vector representation.
[0041] Specifically, when obtaining the feature vector representation of the target video, the feature vector representations of each image frame in the video can be obtained first, and the feature vector representation corresponding to the entire target video is obtained by splicing based on the feature vector representations of each image frame.
[0042] Alternatively, feature processing can also be performed separately from different video feature perspectives to obtain the feature vector representations of the target video under different feature dimensions respectively.
[0043] For example, as Figure 2 shown, in one embodiment, the feature vector representation includes a first feature vector and a second feature vector; correspondingly, obtaining the feature vector representation of the target video includes:
[0044] Dividing the target video into T non-overlapping segments; T is an integer greater than 1;
[0045] Performing feature sequence extraction on the image frames in each segment through a feature extractor to obtain the RGB feature and the optical flow feature of each segment; respectively outputting the RGB feature and the optical flow feature to a feature embedding module to obtain the first feature vector corresponding to the RGB feature and the second feature vector corresponding to the optical flow feature.
[0046] In this process, it is necessary to obtain the feature sequence at the frame level extracted from the video. Specifically, for the target video, it is divided into T non-overlapping segments to facilitate reasonable segmented processing of videos with a large number of frames. Each segment consists of, for example, 16 frames of images. T can take different values for different videos. For the T segments in the video, the RGB frames and optical flow segments contained therein are obtained, that is, the optical flow and RGB streams are obtained. Then, feature extraction is performed on the optical flow and RGB streams, and features with a dimension of 1*D are extracted. Specifically, they can be respectively input into a pre-trained I3D network to process and obtain the corresponding RGB and optical flow segment-level features.
[0047] To make it applicable to the weakly-supervised temporal action localization task, this feature can be further passed into a new convolutional layer (specifically, for example, a temporal convolutional layer) for convolutional processing to achieve feature embedding, and a new set of features x is obtained. i There are also T 1*D-dimensional x. i Respectively, the first feature vector corresponding to the RGB feature and the second feature vector corresponding to the optical flow feature are obtained. In the subsequent processing, the first feature vector of the RGB feature and the second feature vector corresponding to the optical flow feature will be sent to two different streams for independent action recognition processing, and two independent action recognition processing results will be obtained. The two action recognition processing results can be fused to obtain the action recognition result corresponding to the entire video.
[0048] Among them, optionally, the neural network structures involved in the two processing streams have the same design, but they do not share parameters.
[0049] Step 102: Determine the attention weights of the foreground image frames in the target video according to a preset attention dynamic threshold.
[0050] Among them, the attention weights are used to indicate the probability that the image frame contains an action.
[0051] An untrimmed video contains background and actions. We need to focus on the segments that may contain actions and delete the segments that may contain background.
[0052] In this step, it is necessary to introduce an attention mechanism and an attention dynamic threshold to identify the foreground and background of the image frames in the target video. Specifically, the foreground image frames that meet the threshold conditions are selected from each image frame through the attention dynamic threshold to achieve the preliminary identification and screening of the possibility of whether the image frame contains an action.
[0053] The attention dynamic threshold here is an adjustable judgment value. Using the dynamic threshold, different numbers of frames are selected when selecting foreground image frames for different videos through the attention mechanism, which has better extensibility for different videos to improve the accuracy of image localization and action recognition.
[0054] Among them, as an implementation rather than a limitation, determining the attention weight of the foreground image frame in the target video according to a preset attention dynamic threshold includes:
[0055] Input the feature vector representation into the attention mechanism module to obtain the attention weight corresponding to the image frame in the target video; use the preset attention dynamic threshold, combine the size relationship between the attention weight and the attention dynamic threshold, and select the foreground image frame from the target video and determine the attention weight of the foreground image frame.
[0056] Here, the feature vector representation is processed by the attention mechanism module. Specifically, the feature vector representation is passed into the fully connected layer network to implement the attention mechanism processing on the feature vector representation. Specifically, the following formula can be used:
[0057] A i = σ(wA·x i + bA);
[0058] Among them, σ(·) represents the sigmoid function, wA is the weight vector adopted by the attention mechanism, bA is the attention offset adopted in the attention mechanism, and A i is the generated attention weight value. The value of A i ranges between 0 and 1, indicating the possibility that the i-th image frame contains an action. The value of A i being 0 represents the background, the value of 1 represents the action, and 0-1 represents the probability of having an action. Among them, the fully connected layer network acts as the attention mechanism module to process the feature vector representation and obtain the attention weight of each image frame.
[0059] The calculation of the attention weight of the image frame in the target video can be to calculate the attention weight of each image frame in the entire target video using the attention mechanism module, or based on the selected segment from the target video, calculate the attention weight of the image frame in each segment using the attention mechanism module.
[0060] After obtaining the attention weight corresponding to the image frame in the target video, the attention weight of the foreground image frame in the target video can be determined according to the preset attention dynamic threshold, combined with the size relationship between the attention weight and the attention dynamic threshold.
[0061] Specifically, select the image frame with an attention weight greater than the attention dynamic threshold as the foreground image frame, select the image frame with an attention weight less than or equal to the attention dynamic threshold as the background image frame, and at the same time determine the attention weight corresponding to the foreground image frame to achieve accurate distinction between the foreground and the background.
[0062] Step 103: Obtain the target feature vector according to the attention weight and feature vector representation of the foreground image frame.
[0063] After selecting the foreground image frames with a greater probability of containing actions, the attention weight of the foreground image frame can be used to perform attention weighting on the feature vector representation of the target video, and the target feature vector after attention weighting is obtained. This target feature vector forms the foreground feature vector in the target video, so as to be able to more accurately focus on the foreground features and achieve the recognition and prediction of actions in the video.
[0064] Among them, the target feature vector can be obtained by multiplying the attention weight of the foreground image frame after weighted calculation with the feature vector representation.
[0065] Specifically, in an optional implementation manner, obtaining the target feature vector according to the attention weight and feature vector representation of the foreground image frame includes:
[0066] Calculate the average value of the attention weights of the foreground image frames to obtain the target attention weight corresponding to the target video; multiply the target attention weight by the feature vector representation to obtain the target feature vector.
[0067] Among them, for the acquisition of the target attention weight corresponding to the target video, for example, when the attention dynamic threshold is selected as 0.5, we take out the frames with attention weights greater than the threshold 0.5 from the image frames of the target video as the foreground image frames that are initially considered likely to contain actions. Each frame corresponds to a time, so it can be considered that the start and end times of the action are obtained. Furthermore, the average value of the possibility that the frames within the action occurrence segment defined by the start time and end time of the action contain actions is obtained as the attention weight corresponding to the entire video containing the coherent action, which improves the accuracy of subsequent action recognition.
[0068] Here, an attention weighted pooling layer can be specifically used to act on the feature vector to generate the foreground feature x fg . The calculation equation is:
[0069]
[0070] Among them, a i is the attention weight of the foreground image frame, T is the number of foreground image frames, and x i is the feature vector representation of the target video.
[0071] Corresponding to this processing process, the determination of the attention dynamic threshold preset in step 102 above can be achieved by using the target attention weight calculated therein.
[0072] Among them, the preset attention dynamic threshold can be dynamically selected according to the optimal threshold obtained during training during the model application process after the model training is completed, or it can be dynamically selected according to the optimal thresholds corresponding to different video lengths, different action objects to be recognized, etc. obtained during the model training.
[0073] During the model training process, the target video is the video in the training sample set, and the videos in the training sample set have action category labels; the preset attention dynamic threshold is obtained through the following method:
[0074] Using Calculate this attention dynamic threshold;
[0075] Among them, k n is the attention dynamic threshold, α n is the target attention weight generated in the previous model training process; M is the number of model iterative training executions, M is a positive integer, is the sum of the target attention weights generated in M executed model iterative trainings.
[0076] That is, the attention dynamic threshold adopted by the current target video in the action recognition process is calculated based on the target attention weights (i.e., the average attention value) corresponding to the videos that have undergone sample training in the model before.
[0077] Among them, the model needs to perform a cyclic iterative training process according to the training sample set during the training process. In the iterative training process of the model, the determination of the attention dynamic threshold in the current training process needs to be based on the target attention weight generated in the previous model training process and the sum of the target attention weights generated in the historical iterative training processes that have been executed (specifically M times). Specifically, the ratio of the two is used as the attention dynamic threshold in the current training process to perform the current model training operation.
[0078] In this process, through the input training of the videos as training samples in the training sample set in the model, the attention dynamic threshold in the model is adjusted and optimized accordingly.
[0079] Furthermore, in the model training process, various methods can be used to optimize the parameters of the model.
[0080] In an optional implementation manner, after inputting the feature vector representation into the attention mechanism module and obtaining the attention weights corresponding to the image frames in the target video, it further includes:
[0081] Using the attention loss function to optimize the parameters of the attention mechanism module;
[0082] Among them, the attention loss function is:
[0083]
[0084] Among them, m is a hyperparameter, a is the attention weight corresponding to the image frame in the target video, and A is the set of attention weights corresponding to the image frames in the target video. is the attention weights with values in the top m / kn in this set. is the attention weights with values in the last m / kn in this set.
[0085] In this attention loss function, from the set of attention weights corresponding to the image frames in the target video, according to the descending order of values, select the m / kn attention weights in the front (that is, the part with larger values ), and select the m / kn attention weights at the end (that is, the part with smaller values ), and the combination of the two realizes the optimization constraint of the model parameters.
[0086] The proposed attention loss function, combined with the attention dynamic threshold, enables the attention loss to force the top attention to approach 1 and the bottom attention to approach 0. Among them, the maximum and minimum attention parts dynamically adjust the loss regression according to the attention values of each video, obtaining the selection of dynamically determining the total maximum and minimum attention values, realizing better discrimination between actions and backgrounds, and improving the accuracy of action recognition of the model.
[0087] In the above processing process, in order to improve the flexibility of the attention mechanism, a dynamic attention threshold is introduced to automatically adjust the threshold part for distinguishing foreground image frames and background image frames for each video in the training sample set, which can help the attention mechanism better approach the extreme values of different videos.
[0088] Step 104: Input the target feature vector into the action recognition network for action recognition to obtain the action recognition result in the target video.
[0089] The action recognition network is a multi-classifier, which can predict the probability values of different types of actions contained in the target video based on the input attention-weighted target feature vector. For example, if there are c actions, the output 1*c prediction vector is [0.1, 0.2, 0.1, 0.7, 0.4], and each decimal between 0 and 1 represents the probability of the corresponding action, realizing the multi-action classification recognition of the target video. Or the action recognition network is a binary classifier, which predicts the probability value of a specific type of action contained in the target video based on the input attention-weighted target feature vector, realizing the recognition of a specific action.
[0090] Combined with Figure 2As shown, the action recognition network is specifically the softmax layer of the fully connected layer (FC). The target feature vector (i.e., the foreground feature x fg ) is input into the softmax layer of the fully connected layer to obtain the final video-level classification result Y.
[0091] Among them, as a specific implementation, the aforementioned target feature vector includes a first target feature vector generated based on the first feature vector and a second target feature vector generated based on the second feature vector; correspondingly, inputting the target feature vector into the action recognition network for action recognition, the action recognition result in the target video is obtained, including:
[0092] Input the first target feature vector into the action recognition network for action recognition to obtain the first recognition result;
[0093] Input the second target feature vector into the action recognition network for action recognition to obtain the second recognition result;
[0094] Fuse the first recognition result and the second recognition result according to a set ratio to obtain the action recognition result in the target video.
[0095] Among them, for the case where when obtaining the feature vector representation of the target video, the first feature vector corresponding to the RGB feature and the second feature vector corresponding to the optical flow feature are obtained, it is necessary to process and obtain the target feature vector according to the first feature vector and the attention weight of the foreground image frame, and according to the second feature vector and the attention weight of the foreground image frame in step 103, that is, obtain the first target feature vector generated based on the first feature vector and the second target feature vector generated based on the second feature vector (the specific generation process refers to the relevant description of generating the target feature vector in step 103).
[0096] And input the first target feature vector and the second target feature vector into the action recognition network (FC) respectively for action recognition to obtain their respective action recognition results Y*, and then fuse the video-level action recognition results of the RGB stream and the optical flow according to a ratio to obtain the final action recognition result Y.
[0097] When fusing according to the set ratio, specifically, the predicted probability values corresponding to the first recognition result and the predicted probability values corresponding to the second recognition result can be given different weights and then added. The set ratio is, for example, 1:1, that is, weight values of 0.5 are given to the two recognition results respectively and then added together to calculate the final action recognition result corresponding to the target video.
[0098] Furthermore, after inputting the target feature vector into the action recognition network for action recognition to obtain the action recognition result in the target video, it further includes:
[0099] Based on the action recognition result and the action category label set for the video in the training sample set, the cross-entropy loss function is used to optimize the parameters of the action recognition network.
[0100] In the embodiment of the present application, by obtaining the feature vector representation of the target video, determining the attention weight of the foreground image frame in the target video according to the preset attention dynamic threshold, and obtaining the target feature vector according to the attention weight of the foreground image frame and the feature vector representation, and inputting the target feature vector into the action recognition network for action recognition to obtain the action recognition result in the target video. This process combines the attention mechanism, introduces the attention dynamic threshold to identify the foreground image frame in the video, realizes the distinction between the foreground and the background, and combines the attention weight of the foreground image frame with the feature vector representation to realize the attention weighting process, so as to obtain the target feature vector after the attention weighting process and realize the final action recognition. This process fuses the attention mechanism and the attention weighting mechanism based on the feature sequence, highlights the image features after removing the irrelevant image frames, improves the reliability of the image action recognition, and ensures the accuracy of the video action recognition result.
[0101] In the embodiment of the present application, different implementation manners of the action recognition method are also provided.
[0102] See Figure 3 , Figure 3 is the flow of an action recognition method provided by the embodiment of the present application Figure 2 . As Figure 3 shown, an action recognition method includes the following steps:
[0103] Step 301, obtaining the feature vector representation of the target video;
[0104] The implementation process of this step is the same as that of step 101 in the foregoing embodiment, and will not be elaborated here.
[0105] Step 302, determining the attention weight of the foreground image frame in the target video according to the preset attention dynamic threshold, where the attention weight is used to indicate the probability that the image frame contains an action;
[0106] The implementation process of this step is the same as that of step 102 in the foregoing embodiment, and will not be elaborated here.
[0107] Step 303, obtaining the target feature vector according to the attention weight of the foreground image frame and the feature vector representation.
[0108] The implementation process of this step is the same as that of step 103 in the foregoing embodiment, and will not be elaborated here.
[0109] Step 304: Input the target feature vector into the action recognition network for action recognition to obtain the action recognition result in the target video.
[0110] The implementation process of this step is the same as that of step 104 in the foregoing embodiment, and will not be elaborated here.
[0111] Further, after determining the attention weight of the foreground image frames in the target video according to the preset attention dynamic threshold, it further includes:
[0112] Step 305: Take the video time corresponding to the first foreground image frame in the target video as the action start time, and take the video time corresponding to the last foreground image frame in the target video as the action end time.
[0113] Specifically, in the case where the target video is divided into T non-overlapping segments, it may specifically be to take the video time corresponding to the first foreground image frame in each segment of the target video as the action start time, and the video time corresponding to the last foreground image frame as the action end time corresponding to this action start time.
[0114] In the case where the target video is not divided into T non-overlapping segments, directly take the video time corresponding to the first foreground image frame in the entire target video as the action start time, and take the video time corresponding to the last foreground image frame in the entire target video as the action end time.
[0115] This process realizes the determination of the start and end times of the action in the video. Combining with the action recognition result in the video, it forms the action proposal {(ts, te, c, s)} in the target video, where ts is the start time of the action, te is the end time of the action, c is the action class prediction, and s is the confidence value of the action proposal, completing the complete implementation process of the temporal action localization method, and accurately determining the action that occurs and the time when the action occurs.
[0116] In the embodiment of the present application, by obtaining the feature vector representation of the target video, determining the attention weight of the foreground image frames in the target video according to the preset attention dynamic threshold, and obtaining the target feature vector according to the attention weight of the foreground image frames and the feature vector representation, inputting the target feature vector into the action recognition network for action recognition to obtain the action recognition result in the target video. This process combines the attention mechanism, introduces the attention dynamic threshold to identify the foreground image frames in the video, realizes the distinction between the foreground and the background, fuses the attention mechanism and the attention weighting mechanism on the basis of the feature sequence, highlights the image features after removing the irrelevant image frames, completes the complete implementation process of the temporal action localization method, accurately determines the action that occurs and the time when the action occurs, improves the reliability of image action recognition, and ensures the accuracy of the video action recognition result.
[0117] See Figure 4 , Figure 4 which is a structural diagram of an action recognition device provided by an embodiment of the present application. For the sake of convenience of description, only parts related to the embodiment of the present application are shown.
[0118] The action recognition device 400 includes:
[0119] A first acquisition module 401, configured to acquire a feature vector representation of a target video;
[0120] A determination module 402, configured to determine an attention weight of a foreground image frame in the target video according to a preset attention dynamic threshold, where the attention weight is used to indicate the probability that an image frame contains an action;
[0121] A second acquisition module 403, configured to obtain a target feature vector according to the attention weight of the foreground image frame and the feature vector representation;
[0122] An action recognition module 404, configured to input the target feature vector into an action recognition network for action recognition, and obtain an action recognition result in the target video.
[0123] Among them, the determination module 402 is specifically configured to:
[0124] Input the feature vector representation into an attention mechanism module to obtain an attention weight corresponding to an image frame in the target video;
[0125] Use the preset attention dynamic threshold, combine the size relationship between the attention weight and the attention dynamic threshold, select a foreground image frame from the target video, and determine the attention weight of the foreground image frame.
[0126] Among them, the second acquisition module 403 is specifically configured to:
[0127] Calculate the average value of the attention weights of the foreground image frames to obtain a target attention weight corresponding to the target video;
[0128] Multiply the target attention weight by the feature vector representation to obtain the target feature vector.
[0129] Among them, the target video is a video in a training sample set, and the videos in the training sample set have action category labels; the preset attention dynamic threshold is obtained through the following manner:
[0130] Use to calculate the attention dynamic threshold;
[0131] Among them, k n is the attention dynamic threshold, and α nis the target attention weight generated during the previous model training process; M is the number of executed model iterative trainings, and M is a positive integer. is the sum of the target attention weights generated in the M executed model iterative trainings.
[0132] Among them, the device further includes: an optimization module, configured to:
[0133] optimize the parameters of the attention mechanism module by using an attention loss function;
[0134] Among them, the attention loss function is:
[0135]
[0136] Among them, m is a hyperparameter, a is the attention weight corresponding to the image frame in the target video, and A is the set of the attention weights corresponding to the image frames in the target video. is the attention weights among the first m / kn in terms of numerical size in the set. is the attention weights among the last m / kn in terms of numerical size in the set.
[0137] Among them, the feature vector representation includes a first feature vector and a second feature vector; the first acquisition module 401 is specifically configured to:
[0138] divide the target video into T non-overlapping segments; T is an integer greater than 1;
[0139] extract a feature sequence from the image frames in each segment through a feature extractor to obtain the RGB feature and the optical flow feature of each segment;
[0140] output the RGB feature and the optical flow feature to a feature embedding module respectively to obtain the first feature vector corresponding to the RGB feature and the second feature vector corresponding to the optical flow feature.
[0141] Among them, the target feature vector includes a first target feature vector generated based on the first feature vector and a second target feature vector generated based on the second feature vector; the action recognition module 404 is specifically configured to:
[0142] input the first target feature vector into an action recognition network for action recognition to obtain a first recognition result;
[0143] input the second target feature vector into the action recognition network for action recognition to obtain a second recognition result;
[0144] Fuse the first recognition result and the second recognition result according to a set ratio to obtain the action recognition result in the target video.
[0145] The device further includes:
[0146] A time determination module, configured to use the video time corresponding to the first foreground image frame in the target video as the action start time, and use the video time corresponding to the last foreground image frame in the target video as the action end time.
[0147] The action recognition device provided by the embodiments of the present application can implement each process of the embodiments of the above action recognition method and achieve the same technical effects. To avoid repetition, details are not described herein again.
[0148] Figure 5 It is a structural diagram of a terminal provided by an embodiment of the present application. As shown in Figure 5 the figure, the terminal 5 of this embodiment includes: at least one processor 50 ( Figure 5 only one is shown in the figure), a memory 51, and a computer program 52 stored in the memory 51 and executable on the at least one processor 50. When the processor 50 executes the computer program 52, the steps in any of the above method embodiments are implemented.
[0149] The terminal 5 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal 5 may include, but is not limited to, a processor 50 and a memory 51. Those skilled in the art can understand that Figure 5 this is only an example of the terminal 5 and does not constitute a limitation on the terminal 5. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the terminal may further include input / output devices, network access devices, buses, etc.
[0150] The processor 50 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.
[0151] The memory 51 may be an internal storage unit of the terminal 5, such as the hard disk or memory of the terminal 5. The memory 51 may also be an external storage device of the terminal 5, such as a plug-in hard disk equipped on the terminal 5, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 51 may also include both the internal storage unit of the terminal 5 and the external storage device. The memory 51 is used to store the computer program and other programs and data required by the terminal. The memory 51 may also be used to temporarily store the data that has been output or will be output.
[0152] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated herein.
[0153] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0154] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by the combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraint conditions of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0155] In the embodiments provided in the present application, it should be understood that the disclosed device / terminal and method can be implemented in other ways. For example, the device / terminal embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0156] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0157] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0158] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present application, it can also be completed by a computer program instructing the relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0159] The implementation of all or part of the processes in the method of the above embodiments in this application can also be realized by a computer program product. When the computer program product runs on a terminal, it causes the terminal to execute and implement the steps in the above various method embodiments.
[0160] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for motion recognition, characterized in that: include: Obtain the feature vector representation of the target video; Determining an attention weight of a foreground image frame in the target video according to a preset attention dynamic threshold, wherein the attention weight is used to indicate a probability that an action is contained in the image frame; Obtaining a target feature vector according to the attention weight of the foreground image frame and the feature vector representation; Inputting the target feature vector into an action recognition network for action recognition to obtain an action recognition result in the target video; Optimizing parameters of an attention mechanism module using an attention loss function; the attention mechanism module is configured to obtain attention weights corresponding to image frames in the target video based on the feature vector representation; Among them, the attention loss function is: Wherein, m is a hyperparameter, a is the attention weight corresponding to the image frame in the target video, and A is the set of attention weights corresponding to the image frames in the target video. is the attention weight of the first m / kn values in the set, is the attention weight of the last m / kn values in the set.
2. The method according to claim 1, characterized in that The determining the attention weight of the foreground image frame in the target video according to a preset attention dynamic threshold comprises: Inputting the feature vector representation into the attention mechanism module to obtain the attention weight corresponding to the image frame in the target video; By using the preset attention dynamic threshold and combining the size relationship between the attention weight and the attention dynamic threshold, a foreground image frame is selected from the target video and the attention weight of the foreground image frame is determined.
3. The method according to claim 2, characterized in that Obtaining a target feature vector according to the attention weight of the foreground image frame and the feature vector representation includes: averaging the attention weights of the foreground image frames to obtain a target attention weight corresponding to the target video; The target feature vector is obtained by multiplying the target attention weight by the feature vector representation.
4. The method according to claim 3, characterized in that in, The target video is a video in a training sample set, and the video in the training sample set has an action category label; the preset attention dynamic threshold is obtained by: use Calculating the attention dynamic threshold; Among them, k n is the attention dynamic threshold, α n is the target attention weight generated in the previous model training process; M is the number of model iteration trainings that have been performed, and M is a positive integer. is the sum of the target attention weights generated in M executed model iteration trainings.
5. The method according to claim 1, wherein The feature vector representation includes a first feature vector and a second feature vector; and obtaining the feature vector representation of the target video includes: Divide the target video into T non-overlapping segments; T is an integer greater than 1; Extracting a feature sequence from the image frames in each of the segments using a feature extractor to obtain RGB features and optical flow features of each segment; The RGB feature and the optical flow feature are respectively output to a feature embedding module to obtain the first feature vector corresponding to the RGB feature and the second feature vector corresponding to the optical flow feature.
6. The method according to claim 5, characterized in that The target feature vector includes a first target feature vector generated based on the first feature vector and a second target feature vector generated based on the second feature vector; Inputting the target feature vector into an action recognition network for action recognition to obtain an action recognition result in the target video includes: Inputting the first target feature vector into an action recognition network to perform action recognition and obtain a first recognition result; Inputting the second target feature vector into the action recognition network to perform action recognition, thereby obtaining a second recognition result; The first recognition result and the second recognition result are fused according to a set ratio to obtain the action recognition result in the target video.
7. The method according to claim 1, characterized in that After determining the attention weight of the foreground image frame in the target video according to the preset attention dynamic threshold, the method further includes: The video time corresponding to the first foreground image frame in the target video is used as the action start time, and the video time corresponding to the last foreground image frame in the target video is used as the action end time.
8. A terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Method and device for recognizing motion behavior of target body
CN108629326A
Behavior recognition method and system, computer equipment and storage medium
CN110298332A
Behavior detection method based on graph structure information interaction enhancement and electronic device
CN111985333A