Action recognition method and device, and electronic device
Patent Information
- Application Number
- CN202210437437.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-22
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-04-22
AI Technical Summary
现有技术中,提取的描述视频的局部高维视觉特征通常采用密集采样的方式进行特征提取,会导致识别效率低下
[0017]本申请实施例公开的动作识别方法,通过对视频图像序列进行稀疏采样以及特征提取,获取所述视频图像序列中动作的第一特征向量,其中,所述第一特征向量携带所述视频图像序列中动作的分类信息;获取表征所述视频图像序列中动作相关性的第二特征向量;获取所述视频图像序列经稀疏采样后得到的图像帧序列的第三特征向量,其中,所述第三特征向量用于表征所述视频图像序列匹配的动作描述文本;融合所述第一特征向量,所述第二特征向量,以及,所述第三特征向量,对所述视频图像序列中的动作进行动作识别,有助于提升动作识别效率。
Smart Images

Figure CN116994327B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular to action recognition methods, devices, electronic devices, and computer-readable storage media. Background Technology
[0002] Due to the widespread application of video in security surveillance, human behavior analysis, and many other fields, understanding the behavior of objects in video (such as human behavior) has become a prominent research topic in computer vision. Most existing action recognition algorithms typically first extract local high-dimensional visual features describing the video, then fuse these densely extracted features into a fixed-size video-level descriptor, and finally train an SVM on a bag-of-visual-words dataset to predict the final result. However, in existing technologies, the extraction of local high-dimensional visual features describing the video usually employs dense sampling, leading to low recognition efficiency.
[0003] It is evident that existing action recognition methods still require improvement. Summary of the Invention
[0004] This application provides an action recognition method that helps improve action recognition efficiency.
[0005] In a first aspect, embodiments of this application provide an action recognition method, including:
[0006] By performing sparse sampling and feature extraction on the video image sequence, a first feature vector of the action in the video image sequence is obtained, wherein the first feature vector carries classification information of the action in the video image sequence;
[0007] Obtain a second feature vector characterizing the action correlation in the video image sequence;
[0008] The third feature vector of the image frame sequence obtained after sparse sampling of the video image sequence is obtained, wherein the third feature vector is used to characterize the action description text matched by the video image sequence;
[0009] The first feature vector, the second feature vector, and the third feature vector are fused to perform action recognition on the actions in the video image sequence.
[0010] Secondly, embodiments of this application provide an action recognition device, including:
[0011] The first feature vector acquisition module is used to obtain a first feature vector of the action in the video image sequence by performing sparse sampling and feature extraction on the video image sequence, wherein the first feature vector carries classification information of the action in the video image sequence;
[0012] The second feature vector acquisition module is used to acquire a second feature vector characterizing the action correlation in the video image sequence;
[0013] The third feature vector acquisition module is used to acquire the third feature vector of the image frame sequence obtained after sparse sampling of the video image sequence, wherein the third feature vector is used to characterize the action description text matched by the video image sequence;
[0014] The fusion recognition module is used to fuse the first feature vector, the second feature vector, and the third feature vector to perform action recognition on the actions in the video image sequence.
[0015] Thirdly, embodiments of this application also disclose an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the action recognition method described in embodiments of this application.
[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, represents the steps of the action recognition method disclosed in embodiments of this application.
[0017] The action recognition method disclosed in this application obtains a first feature vector of actions in the video image sequence by sparse sampling and feature extraction, wherein the first feature vector carries classification information of actions in the video image sequence; obtains a second feature vector characterizing the correlation of actions in the video image sequence; obtains a third feature vector of the image frame sequence obtained after sparse sampling of the video image sequence, wherein the third feature vector is used to characterize the action description text matched by the video image sequence; and fuses the first feature vector, the second feature vector, and the third feature vector to perform action recognition on actions in the video image sequence, which helps to improve action recognition efficiency.
[0018] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] Figure 1 This is a flowchart of the action recognition method according to Embodiment 1 of this application;
[0021] Figure 2 This is a schematic diagram of the network structure applied to action recognition in Embodiment 1 of this application;
[0022] Figure 3 This is a schematic diagram of the action recognition device structure according to Embodiment 2 of this application;
[0023] Figure 4 A block diagram schematically illustrates an electronic device for performing the method according to this application; and
[0024] Figure 5 A storage unit for holding or carrying program code implementing the method according to this application is illustrated schematically. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0026] Example 1
[0027] This application discloses an action recognition method, such as... Figure 1 As shown, the method includes steps 110 to 140.
[0028] Step 110: Obtain the first feature vector of the action in the video image sequence by sparse sampling and feature extraction.
[0029] The first feature vector carries classification information of the actions in the video image sequence.
[0030] In some embodiments of this application, a Temporal Segment Network (TSN) structure is used to extract short snippets from a long video image sequence through sparse sampling. These short snippets follow a uniform distribution in the time dimension. Therefore, information can be collected from the sampled short snippets using the segment structure. In some embodiments of this application, by downsampling the video image sequence, information about different actions can be collected from the sampled image sequence. For example, information about whether the task action in the image belongs to a preset action such as skating, smoking, or drinking water can be obtained.
[0031] In some embodiments of this application, obtaining the first feature vector of actions in the video image sequence by sparse sampling and feature extraction includes: segmenting the video image sequence into segments with equal time intervals to determine several video segments; randomly downsampling each video segment to obtain sample segments of each video segment; performing classification mapping on each sample segment to obtain action classification results corresponding to each sample segment; obtaining consensus on the action classification results corresponding to each sample segment; and predicting the action category in the video image sequence based on the consensus to obtain the first feature vector.
[0032] For example, for a video V (i.e., a sequence of video images), it can first be divided into M video segments of equal duration, for example, represented as {S1, S2, ..., S...}. k} where K is an integer greater than 2. Then, for each video segment obtained from the division, random downsampling is performed, and multiple frames of video images are randomly sampled to form a sampled segment. The multiple frames of video images constituting the sampled segment can be consecutive image frames or interval image frames arranged in chronological order according to timestamps. Preferably, the multiple frames of video images constituting the sampled segment are consecutive image frames. For example, for video segment S... k Downsampling is performed to obtain sample segment T. k Next, the action category in the video image sequence is predicted from the image segment sequence composed of all the sampled segments obtained above through the TSN network.
[0033] In some embodiments of this application, a pre-trained TSN network model can be used to obtain the action classification results of the downsampled sample segment sequence. The TSN network model can be represented as:
[0034] TSN(T1,T2,...,T k )=H(G(F(T1;W),F(T2;W),...,F(Tk ;W)));
[0035] Where (T1,T2,...,T) k ) represents the sequence of sampled segments; F(;W) represents the convolution function with parameter W, which runs on the input sampled segments and generates category scores for all preset categories; G(,,...,) represents the consensus extraction function, which combines the outputs of multiple sampled segments to obtain the classification consensus; H() is the action classification result prediction function, which predicts the probability that the human action in the video image sequence matches each action category to be identified based on the consensus extracted by the consensus extraction function.
[0036] Taking the action categories to be identified as: skating, smoking, drinking water, cycling, and applying lipstick as an example, the category score generation function F(T) k ;W) in sampling segment T k After running, the output sample segment T will be executed. k The probability score of matching the character's actions with the above-mentioned action categories to be identified. Consensus extraction function G(F(T1;W),F(T2;W),...,F(T) k ;W)) is used to combine the function F(;W) for the sample segments T1, T2, ..., T respectively. k The classification output is used to obtain the sampled segments T1, T2, ..., T k The classification results reach a consensus. Next, the action classification result prediction function H() extracts sampling fragments T1, T2, ..., T based on the consensus extraction function. k The classification results reach a consensus, predicting the sampled segments T1, T2, ..., T k The probability of matching each action category to the actions of the person in the sampled segment. Sampling segments T1, T2, ..., T k The probability that a person's actions match each action category to be identified is the probability that a person's actions match each action category to be identified in the video image sequence.
[0037] In some embodiments of this application, the aforementioned TSN network model can be obtained by supervised training using labeled image segment sequences as training samples.
[0038] Each labeled image segment sequence is obtained through the aforementioned method of segmenting and sparsely sampling the video images, and the sample label corresponding to each image segment sequence can be manually assigned. The sample label is used to indicate the true value of the action category of the human action contained in the corresponding image segment sequence.
[0039] During the training of the TSN network model, the function H() can be implemented using the widely used Softmax function. The loss of the TSN network model is expressed using the classification cross-entropy loss. In some embodiments of this application, the loss function of the TSN network model can be defined as: Where C is the number of action categories to be identified, and y i It is the label for action category i.
[0040] The iterative training process of the TSN network model is similar to the supervised training process of multi-classification networks in the prior art, and will not be repeated in the embodiments of this application.
[0041] In the process of action recognition of video image sequences, the hidden layer vector output by the last layer of the TSN network model can be used as the first feature vector of the video image sequence to characterize the action category information in the video image sequence.
[0042] Step 120: Obtain a second feature vector characterizing the action correlation in the video image sequence.
[0043] The second feature vector represents the proposal box information of the action sequence in the video image sequence.
[0044] In some embodiments of this application, the correlation of actions in the video sequence can be measured from the temporal order dimension and / or distance dimension of the actions.
[0045] In some embodiments of this application, obtaining a second feature vector characterizing the action correlation in the video image sequence includes: obtaining at least one set of action proposals describing the actions in the video image sequence; instantiating nodes of a graph with the action proposals, and constructing edges connecting the nodes according to the correlation between the action proposals to obtain an action proposal graph describing the video image sequence; and performing feature extraction and mapping on the action proposal graph through a pre-trained graph convolutional network to obtain a second feature vector carrying proposal box information in the video image sequence. In some embodiments of this application, the proposal box includes: key point location information describing a complete action. Each set of action proposals includes: a sequence of human key point location information describing the entire action.
[0046] Each human movement is expressed by different positional relationships of limbs, that is, each movement corresponds to different positions of human key points. These human key points include, but are not limited to, one or more of the following: head, upper arm, forearm, hand, nose, thigh, calf, foot, and knee joint. The definition of human key points can be found in existing human key point detection models. In some embodiments of this application, a pre-trained human key point detection model can detect the positional information of human key points for each movement within each frame of the video image sequence, and generate a proposal based on the positional information of the human key points for each movement in each frame. Thus, taking a video image containing continuous movements of a person as an example, a set of movement proposals for the video image sequence can be obtained based on the positional information of the human key points in each frame of the video image detected by the human key point detection model. A corresponding proposal box can be determined based on each movement proposal.
[0047] For details on generating action proposals and proposal boxes based on the detection results of human key points in the image, please refer to the prior art; these details will not be repeated in the embodiments of this application.
[0048] Next, a graph structure corresponding to the video image sequence can be constructed based on a set of action proposals. In this embodiment, this graph is referred to as an "action proposal graph". In some embodiments of this application, nodes of the action proposal graph can be instantiated for each action proposal, and then edges connecting the nodes can be constructed based on the correlation between different action proposals.
[0049] In some embodiments of this application, the edges between nodes in the action proposal graph include two types: context edges and neighboring edges. Constructing edges connecting the nodes based on the correlation between the action proposals includes: constructing edges connecting the corresponding nodes based on the temporal correlation between the action proposals; and constructing edges connecting the corresponding nodes based on the distance correlation between the action proposals.
[0050] The construction methods for different types of edges are described below.
[0051] (I) Context Edge
[0052] p i and p j Taking the proposal boxes representing the action proposals corresponding to nodes i and j in the action proposal graph as an example, if r(p i ,p j )>θ ctx Then, a context edge is established between node i and node j, where r(p i ,p j θ represents the temporal correlation between the proposal boxes of the action proposals corresponding to nodes i and j, respectively. ctxθ is a defined time correlation threshold. In some embodiments of this application, θ ctx The value of r(p) is determined based on the test results of the trained graph network model. i ,p j The tIoU metric (a text detection evaluation index) can be defined as follows:
[0053]
[0054] Among them, I(p) i ,p j U(p) is used to calculate the temporal intersection between the proposal frames of the two action proposals corresponding to nodes i and j, respectively; i ,p j This is used to calculate the temporal union between the proposal frames of the two action proposals corresponding to nodes i and j, respectively. In some embodiments of this application, the time information of each proposal frame is matched with the timestamp of the video image frame that generates the corresponding action proposal. The calculation methods for the temporal intersection and temporal union between two proposal frames are as described in the prior art for calculating the union and intersection, and will not be repeated here.
[0055] As can be seen from the construction method of context edges, the closer the timestamps of the video image frames matched by the action proposals of two nodes are, the greater the temporal correlation r(p) between the action proposal bounding boxes. i ,p j The larger the value, the greater the likelihood of establishing a context edge between the two nodes.
[0056] (ii) Adjacent edges
[0057] As can be seen from the construction method of context edges, context edges typically connect overlapping proposals corresponding to the same action. In practice, the inventors have found that adjacent actions may also be related, and the information passed between them can help with mutual detection. To handle this correlation, in some embodiments of this application, a corresponding action proposal box is determined based on the action proposal corresponding to each video image frame, using r(p i ,p j If d(p) = 0, query different suggestion boxes, then calculate the distance between suggestion boxes. i ,p j )<θ sur Then, add a neighboring edge between node i and node j, where d(p i ,p j θ represents the distance correlation between the proposal boxes of the action proposals corresponding to nodes i and j, respectively. sur For a defined distance correlation threshold θ sur In some embodiments of this application, θsur The value of d(p) is determined based on the test results of the trained graph network model. i ,p j Calculated using the following formula:
[0058]
[0059] Where, c represents i and c j These represent the center positions of the proposal boxes for the action proposals corresponding to nodes i and j, respectively. i -c j | represents the distance between the proposal boxes corresponding to the action proposals of nodes i and j, respectively; U(p i ,p j ) represents the union of the regions of the proposal boxes corresponding to the action proposals of nodes i and j, respectively.
[0060] As can be seen from the method of constructing adjacent edges, the closer the proposal boxes of two different actions are, the greater the distance correlation d(p) between the proposal boxes of these two action proposals. i ,p j The larger the value, the greater the likelihood of establishing a neighboring edge between the two nodes.
[0061] By instantiating the nodes of the graph based on the detected action proposal boxes according to the above method, and constructing the edges between the nodes based on the temporal and distance correlations between the proposal boxes, the action proposal graph describing the video image sequence can be obtained.
[0062] Next, a pre-trained graph convolutional network is used to extract and map features from the action proposal graph, obtaining all proposal box information in the video image sequence. In the action detection task, each proposal refers to the generated human body's position information on the image. Human actions of the same type are represented in the image with correlation, and actions of different types are also correlated. For example, when a person drinks water, they first need to take the cup, and then they need to drink. Taking the cup and drinking are different types of actions, but they are correlated. Therefore, the contextual information of the action proposals can improve the accuracy of human action detection. The correlation between different actions is beneficial for action classification. The graph convolutional network can accurately classify correlated proposal boxes, thus obtaining proposal box information based on all action proposals in the video image sequence.
[0063] In some embodiments of this application, the hidden layer vector carrying proposal box information output by the last layer of the graph convolutional network can be used as the feature vector of the input action proposal graph, that is, the second feature vector of the video image sequence.
[0064] In some embodiments of this application, the graph convolutional network is pre-trained based on an action proposal graph labeled with proposal boxes. For example, one or more labeled samples can be provided to each type of node, wherein the sample data of the labeled sample is an action proposal for an action in a video image sequence, and the corresponding sample label is the proposal box for that action.
[0065] In other embodiments of this application, the graph convolutional network can also be jointly trained with the fusion network described below, which will not be repeated here.
[0066] In some embodiments of this application, when obtaining at least one set of action proposals describing actions in the video image sequence, at least one set of action proposals for actions in the original video image sequence and at least one set of action proposals in the sampled segment sequence obtained after sparse sampling of the video image sequence can also be obtained simultaneously. Then, a first action proposal map is constructed based on the action proposals obtained from the original video image sequence, and a second action proposal map is constructed based on the action proposals obtained from the sampled segment sequence. Next, the first action proposal map is used for feature extraction and mapping through a pre-trained first graph convolutional network to obtain all proposal box information in the original video image sequence, and the second action proposal map is used for feature extraction and mapping through a pre-trained second graph convolutional network to obtain all proposal box information in the sampled segment sequence. Finally, the proposal box information output by the first graph convolutional network and the second graph convolutional network is aggregated, and the aggregation result is used as the proposal box information of the video image sequence.
[0067] The first graph convolutional network is trained based on labeled samples generated from an uncropped video image sequence, and the second graph convolutional network is trained based on labeled samples generated from a sequence of sampled segments obtained without sparse sampling of the video image sequence.
[0068] The first and second graph convolutional networks use the same network structure. The training process of the first and second graph convolutional networks is described in the previous section on the training process of graph convolutional networks, and will not be repeated here.
[0069] By combining the original video image sequence and the sampled segment sequence obtained by sparse sampling to extract a second feature vector for action recognition, the accuracy of action recognition can be improved.
[0070] Step 130: Obtain the third feature vector of the image frame sequence obtained after sparse sampling of the video image sequence.
[0071] In some embodiments of this application, the third feature vector is used to characterize the action description text matched by the video image sequence.
[0072] In some embodiments of this application, obtaining the third feature vector of the image frame sequence obtained after sparse sampling of the video image sequence includes: obtaining the image frame sequence obtained after sparse sampling of the video image sequence; superimposing the temporal information of each image frame in the image frame sequence onto the visual information of the corresponding image frame through the Transformer encoding module of a pre-trained visual-language model to obtain the third feature vector of the image frame sequence; wherein, the image frame sequence is obtained by the following method: segmenting the video image sequence into segments with equal time intervals to determine several video segments of the video image sequence; randomly downsampling each video segment to obtain a sampled segment of each video segment; selecting several image frames from the sampled segments and arranging them into an image frame sequence according to the chronological order of the video timestamps of the image frames.
[0073] The method for obtaining each image frame in the image frame sequence is the same as the specific method for sparse sampling of the video image sequence in the preceding steps, and will not be repeated here.
[0074] For each sampled segment obtained through sparse sampling, one or more image frames are randomly sampled from each segment. Then, the sampled multiple image frames are arranged into an image frame sequence according to the video timestamp order of the image frames. The image frames in the image frame sequence are arranged according to the video image frame timestamp order; that is, the sequence position of the image frames in the image frame sequence can be used to represent the occurrence time information of the action in the corresponding video image frame. Therefore, in some embodiments of this application, each image frame in the image frame sequence can be encoded using the Transformer encoding module in the vision-language model, so that the action occurrence time information corresponding to the image frame is stacked into the visual features of the image frame, thereby generating the feature vector of the image frame sequence.
[0075] Visual-language models can be used to predict textual representations from input images. In some embodiments of this application, the visual-language model applied to video action recognition can adopt the VL-BERT architecture. The VL-BERT model, based on BERT, embeds a new visual feature into the input to adapt to visually relevant content. Similar to BERT, the VL-BERT model mainly consists of a multi-layer bidirectional Transformer encoder. However, unlike BERT, which only processes sentence words, the VL-BERT model takes both visual and linguistic elements as input, defining corresponding features on the regions of interest (RoIs) of the image and the words in the input sentence, respectively. In the prior art, visual-language models can be trained using a static image set with captions (i.e., descriptive text describing actions in the image).
[0076] In some embodiments of this application, the visual-language model is trained based on several image-text pairs, wherein the images in the image-text pairs include image sequences, and the text in the image-text pairs is descriptive text of actions in the image sequences; during the training of the visual-language model, the two encoding modules of the visual-language model are respectively used to calculate the feature vectors of the image sequences and text in the image-text pairs; the training objective of the visual-language model is to maximize the similarity of correctly matched candidate image-text pairs and minimize the similarity of incorrectly matched candidate image-text pairs, wherein the candidate image-text pairs are generated based on the combination of image sequences and text in the several image-text pairs, the similarity of correctly matched candidate image-text pairs is obtained by calculating a dense similarity matrix based on the feature vectors of the combined image sequences and text of the candidate image-text pairs, and the similarity of incorrectly matched candidate image-text pairs is obtained by calculating a symmetric cross-entropy based on the combination of the image sequences and text of the candidate image-text pairs.
[0077] The encoding module is a Transformer encoding module. When calculating the feature vector of the image sequence in the image-text pair, the Transformer encoding module superimposes the temporal information of each image frame in the image sequence onto the visual information of the corresponding image frame to generate the third feature vector corresponding to the image sequence.
[0078] Taking N image-text pairs (e.g., represented as (image sequence, text)) in a training sample batch as an example, two encoding modules are used to calculate the feature embeddings of the image sequence and the text, respectively, and a dense similarity matrix is calculated among all N possible (image sequence, text) pairs. The elements in the calculated dense similarity matrix represent the relationship between different actions (i.e., image sequences) and text. For example, an image sequence of skating actions has a high similarity to the text describing skating, while an image sequence of skating actions has a low similarity to the text describing "I pick up my water glass to drink water". The training objective of the vision-language model is to optimize the image sequence encoding module and the text encoding module by maximizing the similarity between the N correctly matched (image sequence, text) pairs and minimizing the similarity between the N×(N-1) incorrectly matched (image sequence, text) pairs by calculating symmetric cross-entropy.
[0079] In some embodiments of this application, when training a visual-language model, optimization training can be performed based on existing visual-language models, optimizing only the network parameters of the image sequence encoding module to reduce model training parameters and improve model training efficiency.
[0080] In the process of action recognition of video image sequences, the image sequence obtained after sparse sampling is input into the trained visual-language model. The Transformer encoding module of the visual-language model is used to encode the image sequence after sparse sampling. Then, the vector output by the last hidden layer of the Transformer encoding module can be used as the third feature vector of the input image sequence.
[0081] Step 140: Merge the first feature vector, the second feature vector, and the third feature vector to perform action recognition on the actions in the video image sequence.
[0082] Next, the feature vectors output by the three models are fused together, and action recognition is performed based on the fused vector.
[0083] In some embodiments of this application, fusing the first feature vector, the second feature vector, and the third feature vector to perform action recognition on actions in the video image sequence includes: performing convolution operations on the first feature vector, the second feature vector, and the third feature vector respectively to obtain corresponding vectors of a specified number of dimensions, wherein the specified number of dimensions is equal to the number of action categories to be identified; concatenating the corresponding vectors; performing classification mapping on the concatenated vectors to output the probability that the actions in the video image sequence match the action categories to be identified.
[0084] For example, it can be adopted as follows Figure 2 The network structure shown executes the action recognition method disclosed in this embodiment. The TSN network 210, graph convolutional network 220, and visual-language network 230 are three parallel network branches. The fusion recognition network 240 is used to perform feature fusion and classification mapping based on the hidden layer vectors output by the aforementioned three network branches. The structure of the TSN network 210 and its encoding process for the input video image sequence are described in the aforementioned TSN network model section and will not be repeated here. Similarly, the structure of the graph convolutional network 220 and its encoding process for the input video image sequence are described in the aforementioned graph convolutional network section and will not be repeated here. Likewise, the structure of the visual-language network 230 and its encoding process for the input video image sequence are described in the visual-language model section and will not be repeated here.
[0085] In some embodiments of this application, such as Figure 2As shown, the fusion recognition network 240 includes three parallel convolutional layers 2401, a feature concatenation layer 2402, and an activation function 2403. The TSN network 210, graph convolutional network 220, and visual-language network 230 are decomposed and connected to one convolutional layer. The fusion recognition network 240 maps the output of the last hidden layer of each network branch to a feature space of a specified number of dimensions through the corresponding convolutional layers, obtaining three feature vectors of the same dimension. Then, the feature concatenation layer concatenates the three feature vectors of the same dimension to obtain a concatenated vector. Finally, the activation function performs a classification mapping on the concatenated vector output by the feature concatenation layer, outputting the probability that the concatenated vector matches each action category to be recognized, i.e., the probability that the action in the input video image sequence matches each action category to be recognized. In some embodiments of this application, the action category corresponding to the highest probability can be used as the action recognition result in the input video image sequence.
[0086] In some embodiments of this application, the convolutional layers of the fusion recognition network 240 may use 1×1 convolutional kernels, the feature splicing layer may be implemented using the concat() function, and the activation function may be the softmax activation function.
[0087] In some embodiments of this application, the TSN network 210, graph convolutional network 220, and visual-language network 230 can be jointly trained with the fusion recognition network 240, or the TSN network model, graph convolutional network, and visual-language model can be trained separately.
[0088] The action recognition method disclosed in this application obtains a first feature vector of actions in the video image sequence by sparse sampling and feature extraction, wherein the first feature vector carries classification information of actions in the video image sequence; obtains a second feature vector characterizing the correlation of actions in the video image sequence; obtains a third feature vector of the image frame sequence obtained after sparse sampling of the video image sequence, wherein the third feature vector is used to characterize the action description text matched by the video image sequence; and fuses the first feature vector, the second feature vector, and the third feature vector to perform action recognition on actions in the video image sequence, which helps to improve action recognition efficiency.
[0089] The action recognition method disclosed in this application downsamples a video image sequence, extracts features from multiple aspects based on the downsampled video image sequence, and performs feature fusion recognition. This reduces the number of video image frames processed when performing action recognition on video, thereby improving action recognition efficiency. Simultaneously, because features are extracted from multiple aspects for feature fusion and recognition, the accuracy of action recognition is ensured.
[0090] Furthermore, compared to the application of the TSN network model in the field of behavior recognition in the existing technology, this application obtains sampled segments by randomly downsampling the segmented video segments, and then extracts features based on the sampled segments. It further integrates the features extracted by the graph convolutional network based on the original video image sequence, thereby removing redundant information, improving the speed of action recognition, avoiding the loss of action information, and ensuring the accuracy of action recognition.
[0091] On the other hand, by generating motion proposal graphs based on the actions in the video and then extracting the correlations between motion proposals through a graph convolutional network, the ability of the extracted correlations to express the correlations between motion proposals is enhanced.
[0092] On the other hand, by transferring the visual-language model to the action recognition task in video, the training parameters of the visual-language model applied to the action recognition task are reduced, thereby improving the model training efficiency.
[0093] Example 2
[0094] This application discloses an action recognition device, such as... Figure 3 As shown, the device includes:
[0095] The first feature vector acquisition module 310 is used to acquire a first feature vector of the action in the video image sequence by performing sparse sampling and feature extraction on the video image sequence, wherein the first feature vector carries classification information of the action in the video image sequence;
[0096] The second feature vector acquisition module 320 is used to acquire a second feature vector characterizing the action correlation in the video image sequence;
[0097] The third feature vector acquisition module 330 is used to acquire the third feature vector of the image frame sequence obtained after sparse sampling of the video image sequence, wherein the third feature vector is used to characterize the action description text matched by the video image sequence;
[0098] The fusion recognition module 340 is used to fuse the first feature vector, the second feature vector, and the third feature vector to perform action recognition on the actions in the video image sequence.
[0099] In some embodiments of this application, the first feature vector acquisition module 310 is further configured to:
[0100] The video image sequence is segmented into segments with equal time intervals to determine several video segments of the video image sequence;
[0101] Each video segment is randomly downsampled to obtain a sampled segment of each video segment;
[0102] Each of the sampled segments is classified and mapped to obtain the action classification result corresponding to each of the sampled segments;
[0103] A consensus is reached on the action classification results corresponding to each of the sampled segments;
[0104] Based on the consensus, the action category in the video image sequence is predicted to obtain a first feature vector.
[0105] In some embodiments of this application, the second feature vector acquisition module 320 is further configured to:
[0106] Obtain at least one set of action proposals describing the actions in the video image sequence;
[0107] The nodes of the graph are instantiated with the action proposals, and edges connecting the nodes are constructed according to the correlation between the action proposals to obtain an action proposal graph describing the video image sequence.
[0108] The action proposal graph is feature extracted and mapped by a pre-trained graph convolutional network to obtain a second feature vector carrying proposal box information from the video image sequence.
[0109] In some embodiments of this application, constructing the edges connecting the nodes based on the correlation between the action proposals includes:
[0110] Based on the temporal correlation between the action proposals, construct edges connecting the corresponding nodes; and,
[0111] Based on the distance correlation between the proposed actions, construct edges connecting the corresponding nodes.
[0112] In some embodiments of this application, the third feature vector acquisition module 330 is further configured to:
[0113] Obtain the image frame sequence obtained by sparse sampling of the video image sequence;
[0114] The Transformer encoding module of the pre-trained visual-language model superimposes the temporal information of each image frame in the image frame sequence onto the visual information of the corresponding image frame to obtain the third feature vector of the image frame sequence.
[0115] In some embodiments of this application, the visual-language model is trained based on several image-text pairs, wherein the images in the image-text pairs include image sequences, and the text in the image-text pairs is descriptive text of actions in the image sequences;
[0116] During the training of the visual-language model, the two encoding modules of the visual-language model are used to calculate the feature vectors of the image sequence and the text in the image-text pair, respectively. The training objective of the visual-language model is to maximize the similarity of correctly matched candidate image-text pairs and minimize the similarity of incorrectly matched candidate image-text pairs. The candidate image-text pairs are generated based on the combination of image sequences and text in several image-text pairs. The similarity of correctly matched candidate image-text pairs is obtained by calculating a dense similarity matrix based on the feature vectors of the combined image sequence and text of the candidate image-text pair. The similarity of incorrectly matched candidate image-text pairs is obtained by calculating a symmetric cross-entropy based on the combination of the image sequence and text of the candidate image-text pair.
[0117] In some embodiments of this application, the fusion recognition module 340 is further configured to:
[0118] The first feature vector, the second feature vector, and the third feature vector are convolved to obtain corresponding vectors with a specified number of dimensions, wherein the specified number of dimensions is equal to the number of action categories to be identified.
[0119] Concatenate the corresponding vectors;
[0120] The spliced vector obtained after splicing is classified and mapped, and the probability of the action in the video image sequence matching the action category to be identified is output.
[0121] The action recognition device disclosed in this application is used to implement the action recognition method described in Embodiment 1 of this application. The specific implementation methods of each module of the device will not be repeated here. Please refer to the specific implementation methods of the corresponding steps in the method embodiment.
[0122] The action recognition device disclosed in this application performs sparse sampling and feature extraction on a video image sequence to obtain a first feature vector of actions in the video image sequence, wherein the first feature vector carries classification information of actions in the video image sequence; obtains a second feature vector characterizing the correlation of actions in the video image sequence; obtains a third feature vector of the image frame sequence obtained after sparse sampling of the video image sequence, wherein the third feature vector is used to characterize the action description text matched by the video image sequence; and fuses the first feature vector, the second feature vector, and the third feature vector to perform action recognition on actions in the video image sequence, which helps to improve action recognition efficiency.
[0123] The action recognition device disclosed in this application downsamples a video image sequence, extracts features from multiple aspects based on the downsampled video image sequence, and performs feature fusion recognition. This reduces the number of video image frames processed when performing action recognition on video, thereby improving action recognition efficiency. Simultaneously, because features are extracted from multiple aspects for feature fusion and recognition, the accuracy of action recognition is ensured.
[0124] Furthermore, compared to the application of the TSN network model in the field of behavior recognition in the existing technology, this application obtains sampled segments by randomly downsampling the segmented video segments, and then extracts features based on the sampled segments. It further integrates the features extracted by the graph convolutional network based on the original video image sequence, thereby removing redundant information, improving the speed of action recognition, avoiding the loss of action information, and ensuring the accuracy of action recognition.
[0125] On the other hand, by generating motion proposal graphs based on actions in the video and then extracting the correlations between motion proposals through a graph convolutional network, the ability of the extracted correlations to express the correlations between motion proposals is enhanced.
[0126] On the other hand, by transferring the visual-language model to the action recognition task in video, the training parameters of the visual-language model applied to the action recognition task are reduced, thereby improving the model training efficiency.
[0127] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus embodiments, since they are fundamentally similar to the method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0128] The above provides a detailed description of the action recognition method and apparatus provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and its core idea. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the idea of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
[0129] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0130] The various component embodiments of this application can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components in the electronic device according to the embodiments of this application. This application can also be implemented as a device or apparatus program (e.g., a computer program and computer program product) for performing part or all of the methods described herein. Such a program implementing this application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.
[0131] For example, Figure 4 An electronic device is shown that can implement the methods according to this application. The electronic device may be a PC, mobile terminal, personal digital assistant, tablet computer, etc. The electronic device conventionally includes a processor 410 and a memory 420, and program code 430 stored in the memory 420 and executable on the processor 410, which, when executing the program code 430, implements the methods described in the above embodiments. The memory 420 may be a computer program product or a computer-readable medium. The memory 420 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. The memory 420 has a storage space 4201 for the program code 430 of a computer program for performing any of the method steps described above. For example, the storage space 4201 for the program code 430 may include various computer programs for implementing the various steps in the above methods. The program code 430 is computer-readable code. These computer programs can be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, CDs, memory cards, or floppy disks. The computer program includes computer-readable code that, when executed on an electronic device, causes the electronic device to perform the method according to the above embodiments.
[0132] This application also discloses a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the action recognition method as described in Embodiment 1 of this application.
[0133] Such a computer program product can be a computer-readable storage medium, which can have the same characteristics as... Figure 4 The memory 420 in the illustrated electronic device is similarly arranged with storage segments, storage spaces, etc. Program code can be stored, for example, in a compressed form on the computer-readable storage medium. The computer-readable storage medium is typically as shown in the reference... Figure 5 The portable or fixed storage unit is described above. Typically, the storage unit includes computer-readable code 430', which is code read by a processor and, when executed by the processor, implements the various steps in the method described above.
[0134] The terms "an embodiment," "embodiment," or "one or more embodiments" as used herein mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of this application. Furthermore, please note that the examples of the phrase "in one embodiment" do not necessarily all refer to the same embodiment.
[0135] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this application may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0136] In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.
[0137] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. An action recognition method, characterized in that, include: By performing sparse sampling and feature extraction on the video image sequence, a first feature vector of the action in the video image sequence is obtained, wherein the first feature vector carries classification information of the action in the video image sequence; Obtain a second feature vector characterizing the action correlation in the video image sequence; The third feature vector of the image frame sequence obtained after sparse sampling of the video image sequence is obtained, wherein the third feature vector is used to characterize the action description text matched by the video image sequence; By fusing the first feature vector, the second feature vector, and the third feature vector, action recognition is performed on the actions in the video image sequence; The step of obtaining the second feature vector characterizing the action correlation in the video image sequence includes: Obtain at least one set of action proposals describing the actions in the video image sequence; The nodes of the graph are instantiated with the action proposals, and edges connecting the nodes are constructed according to the correlation between the action proposals to obtain an action proposal graph describing the video image sequence. The action proposal graph is feature extracted and mapped by a pre-trained graph convolutional network to obtain a second feature vector carrying proposal box information from the video image sequence; The step of obtaining the first feature vector of the action in the video image sequence by performing sparse sampling and feature extraction includes: The video image sequence is segmented into segments with equal time intervals to determine several video segments of the video image sequence; Each video segment is randomly downsampled to obtain a sampled segment of each video segment; Each of the sampled segments is classified and mapped to obtain the action classification result corresponding to each of the sampled segments; A consensus is reached on the action classification results corresponding to each of the sampled segments; Based on the consensus, the action category in the video image sequence is predicted to obtain a first feature vector; The step of obtaining the third feature vector of the image frame sequence obtained after sparse sampling of the video image sequence includes: Obtain the image frame sequence obtained by sparse sampling of the video image sequence; The temporal information of each image frame in the image frame sequence is superimposed onto the visual information of the corresponding image frame by the Transformer encoding module of the pre-trained visual-language model to obtain the third feature vector of the image frame sequence. The proposal box includes: key point location information describing a complete action; Each set of motion proposals includes a sequence of key human body location information describing the entire motion.
2. The method according to claim 1, characterized in that, The step of constructing edges connecting the nodes based on the correlation between the action proposals includes: Based on the temporal correlation between the action proposals, construct edges connecting the corresponding nodes; and, Based on the distance correlation between the proposed actions, construct edges connecting the corresponding nodes.
3. The method according to claim 1, characterized in that, The visual-language model is trained based on several image-text pairs, wherein the images in the image-text pairs include image sequences, and the text in the image-text pairs is descriptive text of actions in the image sequences; During the training of the visual-language model, the two encoding modules of the visual-language model are used to calculate the feature vectors of the image sequence and the text in the image-text pair, respectively. The training objective of the visual-language model is to maximize the similarity of correctly matched candidate image-text pairs and minimize the similarity of incorrectly matched candidate image-text pairs. The candidate image-text pairs are generated based on the combination of image sequences and text in several image-text pairs. The similarity of correctly matched candidate image-text pairs is obtained by calculating a dense similarity matrix based on the feature vectors of the combined image sequence and text of the candidate image-text pair. The similarity of incorrectly matched candidate image-text pairs is obtained by calculating a symmetric cross-entropy based on the combination of the image sequence and text of the candidate image-text pair.
4. The method according to any one of claims 1 to 3, characterized in that, The step of fusing the first feature vector, the second feature vector, and the third feature vector to perform action recognition on actions in the video image sequence includes: The first feature vector, the second feature vector, and the third feature vector are convolved to obtain corresponding vectors with a specified number of dimensions, wherein the specified number of dimensions is equal to the number of action categories to be identified. Concatenate the corresponding vectors; The spliced vector obtained after splicing is classified and mapped, and the probability of the action in the video image sequence matching the action category to be identified is output.
5. A motion recognition device, characterized in that, include: The first feature vector acquisition module is used to obtain a first feature vector of the action in the video image sequence by performing sparse sampling and feature extraction on the video image sequence, wherein the first feature vector carries classification information of the action in the video image sequence; The second feature vector acquisition module is used to acquire a second feature vector characterizing the action correlation in the video image sequence; The third feature vector acquisition module is used to acquire the third feature vector of the image frame sequence obtained after sparse sampling of the video image sequence, wherein the third feature vector is used to characterize the action description text matched by the video image sequence; The fusion recognition module is used to fuse the first feature vector, the second feature vector, and the third feature vector to perform action recognition on the actions in the video image sequence. The second feature vector acquisition module is specifically used for: acquiring at least one set of action proposals describing the actions in the video image sequence; instantiating nodes of a graph with the action proposals, and constructing edges connecting the nodes according to the correlation between the action proposals to obtain an action proposal graph describing the video image sequence; and performing feature extraction and mapping on the action proposal graph through a pre-trained graph convolutional network to obtain a second feature vector carrying proposal box information in the video image sequence. The first feature vector acquisition module is specifically used for: acquiring at least one set of action proposals describing the actions in the video image sequence; instantiating nodes of a graph with the action proposals, and constructing edges connecting the nodes according to the correlation between the action proposals to obtain an action proposal graph describing the video image sequence; and performing feature extraction and mapping on the action proposal graph through a pre-trained graph convolutional network to obtain a second feature vector carrying proposal box information in the video image sequence. The third feature vector acquisition module is specifically used for: Obtain the image frame sequence obtained by sparse sampling of the video image sequence; The temporal information of each image frame in the image frame sequence is superimposed onto the visual information of the corresponding image frame by the Transformer encoding module of the pre-trained visual-language model to obtain the third feature vector of the image frame sequence. The proposal box includes: key point location information describing a complete action; Each set of motion proposals includes a sequence of key human body location information describing the entire motion.
6. An electronic device, comprising a memory, a processor, and program code stored in the memory and executable on the processor, characterized in that, When the processor executes the program code, it implements the action recognition method according to any one of claims 1 to 4.
7. A computer-readable storage medium having program code stored thereon, characterized in that, When the program code is executed by the processor, it implements the steps of the action recognition method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Unedited video action time sequence positioning method based on graph convolution network
CN110362715A
Video action detection method based on time sequence convolution modeling
CN110688927A
Visual behavior recognition method and system based on text semantic supervision and computer readable medium
CN112580362A
Video behavior recognition method and system based on action knowledge base and ensemble learning
CN113313039A