Video recognition method and device

By performing feature elements mixing and cross-channel perception feature fusion processing in video recognition technology, the problem of difficulty in balancing accuracy and computing efficiency in the prior art is solved, and more efficient and accurate video recognition is achieved.

CN114973096BActive Publication Date: 2025-05-23JINGDONG TECH HLDG CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210655418.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-10
Publication Date
2025-05-23
Estimated Expiration
2042-06-10

AI Technical Summary

Technical Problem

Existing video recognition technologies are difficult to achieve a good balance between accuracy and computing efficiency, resulting in excessive computational complexity and number of parameters.

Method used

A video recognition method is proposed. By obtaining the initial feature map of the video frame, N-sequential feature fusion processing is performed. The feature fusion processing includes mixing feature elements and cross-channel perception on the feature dimension, reducing the calculation complexity and number of parameters.

Benefits of technology

It achieves a better balance on video recognition problems, reduces the computational complexity and number of parameters, and improves the recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114973096B_ABST
    Figure CN114973096B_ABST
Patent Text Reader

Abstract

The present application proposes a video recognition method and device thereof, which relates to the field of image processing, by obtaining an initial feature map of a video frame in a video to be recognized; starting from the initial feature map, N feature fusion processes are performed in sequence, and the feature map i processed by the i-th feature fusion process is the target feature map i-1 output by the i-th feature fusion process, i and N are both positive integers, 1<i≤N; the feature fusion process includes: mixing feature elements of feature map i in feature dimensions on each feature extraction channel to obtain mixed feature elements, and fusing the mixed feature elements in all feature dimensions, performing cross-channel perception on the fused feature map i to obtain the target feature map i; obtaining the target category of the video to be recognized based on the target feature map N output by the N-th feature fusion process. The present application performs feature element mixing in three independent feature dimensions, reduces computational complexity and the number of parameters, and achieves a balance between accuracy and computational efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and in particular to a video recognition method and device thereof. Background Art

[0002] Deep neural networks have recently achieved widespread success in the field of video recognition, and have made great improvements in previous performance in applications including video action recognition, event detection, indexing, and retrieval. Designing high-performance deep neural networks has also become the key to improving the effects of video-related applications and implementing technologies. Among related technologies, the main drawback of the neural network structure for video recognition is that it is difficult to achieve a good balance between accuracy and computational efficiency. Summary of the invention

[0003] The present application aims to solve one of the technical problems in the related art at least to some extent.

[0004] To this end, the first aspect of the present application proposes a video recognition method, which obtains an initial feature map of a video frame in a video to be recognized; starting from the initial feature map, performing N feature fusion processes in sequence, wherein the feature map i processed by the i-th feature fusion process is the target feature map i-1 output by the i-1-th feature fusion process, and i and N are both positive integers, 1<i≤N; the feature fusion process includes: mixing the feature elements of the feature map i in the feature dimension on each feature extraction channel to obtain a mixed feature element, and fusing the mixed feature elements in all feature dimensions to obtain a fused feature map i, performing cross-channel perception on the fused feature map i to obtain a target feature map i; based on the target feature map N output by the N-th feature fusion process, performing category recognition on the video to be recognized to obtain a target category of the video to be recognized.

[0005] The video recognition method proposed in the embodiment of the present application decomposes the mixture between the feature elements in the video to be identified into the interaction of three independent feature dimensions, and then uses linear projection to combine the interaction results of different feature dimensions, thereby reducing the computational complexity and the number of parameters, and achieving a better balance between accuracy and computational efficiency in the video recognition problem.

[0006] According to an embodiment of the present application, feature elements of feature graph i are mixed in feature dimension to obtain mixed feature elements, including: mixing elements of height feature elements in feature graph i in height dimension to obtain mixed height feature elements; mixing elements of width feature elements in feature graph i in width dimension to obtain mixed width feature elements; mixing elements of time feature elements in feature graph i in time dimension to obtain mixed time feature elements.

[0007] According to an embodiment of the present application, the time feature elements in the feature graph i are mixed in the time dimension to obtain the mixed time feature elements, including: obtaining the time feature elements of the feature graph i in the time dimension and grouping the time feature elements; mixing the time feature elements within each group to obtain the mixed time feature elements.

[0008] According to an embodiment of the present application, grouping the time feature elements includes: uniformly grouping the time feature elements to obtain a first group of the time feature elements.

[0009] According to an embodiment of the present application, grouping the time feature elements includes: discretely sampling the time feature elements to obtain a second grouping of the time feature elements.

[0010] According to an embodiment of the present application, grouping the time feature elements includes: on the basis of uniformly grouping the time feature elements, performing window translation on the time feature elements to obtain a third grouping of the time feature elements.

[0011] According to an embodiment of the present application, grouping time feature elements includes: starting from the first time feature element, determining each time feature element and a preset number of time feature elements consecutively following the time feature element as a group to obtain a fourth group of time feature elements.

[0012] According to an embodiment of the present application, obtaining an initial feature map of a video frame in a video to be identified includes: projecting the video frame in the video to be identified to a feature space to obtain the initial feature map.

[0013] According to an embodiment of the present application, a video recognition method includes: inputting a video frame in a video to be recognized into a classification recognition model, projecting the video frame by a three-dimensional projection layer in the classification recognition model to obtain an initial feature map; performing feature fusion processing N times in sequence starting from the initial feature map by N three-dimensional multilayer perceptron networks in the classification recognition model to output a target feature map N; wherein the N three-dimensional multilayer perceptron networks are connected in series, the input of the first three-dimensional multilayer perceptron network is the initial feature map, the feature map i input by the i-th three-dimensional multilayer perceptron network is the target feature map i-1 output by the i-1-th three-dimensional multilayer perceptron network, i and N are both positive integers, 1<i≤N; inputting the target feature map N into the average pooling layer in the classification recognition model to perform an average pooling operation on the target feature map N, and inputting the mean feature map obtained after the average pooling operation into the fully connected layer in the classification recognition model to obtain the target category of the video to be recognized output by the fully connected layer.

[0014] According to an embodiment of the present application, a three-dimensional multilayer perceptron network includes a feature element mixing unit and a cross-channel perception unit, wherein: the feature element mixing unit mixes the feature elements of a feature map i in a feature dimension to obtain a mixed feature element, and fuses the mixed feature elements in all feature dimensions to obtain a fused feature map i; the cross-channel perception unit performs cross-channel perception on the fused feature map i and outputs a target feature map i.

[0015] According to an embodiment of the present application, a feature element mixing unit includes: a height feature element mixing subunit, a width feature element mixing subunit and a time feature element mixing subunit; the method also includes: the height feature element mixing subunit performs element mixing on the height feature elements in the feature graph i in the height dimension to obtain a mixed height feature element; the width feature element mixing subunit performs element mixing on the width feature elements in the feature graph i in the width dimension to obtain a mixed width feature element; the time feature element mixing subunit performs element mixing on the time feature elements in the feature graph i in the time dimension to obtain a mixed time feature element.

[0016] According to an embodiment of the present application, a transition layer is included between each two adjacent three-dimensional multi-layer perceptron networks, and the target feature map i is input into the transition layer. The transition layer increases the number of feature extraction channels for the target feature map i and reduces the resolution of the target feature map i.

[0017] The second aspect of the present application proposes a video recognition device, including: an acquisition module, used to acquire an initial feature map of a video frame in a video to be recognized; a processing module, used to perform N feature fusion processes in sequence starting from the initial feature map, wherein the feature map i processed by the i-th feature fusion process is the target feature map i-1 output by the i-1-th feature fusion process, i and N are both positive integers, 1<i≤N; the feature fusion process includes: mixing feature elements of the feature map i in the feature dimension on each feature extraction channel to obtain mixed feature elements, and fusing the mixed feature elements in all feature dimensions to obtain a fused feature map i, performing cross-channel perception on the fused feature map i to obtain a target feature map i; an identification module, used to perform category recognition on the video to be recognized based on the target feature map N output by the N-th feature fusion process to obtain a target category of the video to be recognized.

[0018] According to an embodiment of the present application, the processing module is also used to: mix the feature elements of the feature graph i in the feature dimension to obtain a mixed feature element, including: mixing the elements of the height feature elements in the feature graph i in the height dimension to obtain a mixed height feature element; mixing the elements of the width feature elements in the feature graph i in the width dimension to obtain a mixed width feature element; mixing the elements of the time feature elements in the feature graph i in the time dimension to obtain a mixed time feature element.

[0019] According to an embodiment of the present application, the processing module is also used to: mix the time feature elements in the feature graph i in the time dimension to obtain a mixed time feature element, including: obtaining the time feature elements of the feature graph i in the time dimension and grouping the time feature elements; mixing the time feature elements within each group to obtain a mixed time feature element.

[0020] According to an embodiment of the present application, the processing module is further used to: group the time feature elements, including: uniformly grouping the time feature elements to obtain a first group of the time feature elements.

[0021] According to an embodiment of the present application, the processing module is further used to: group the time feature elements, including: discretely sampling the time feature elements to obtain a second group of the time feature elements.

[0022] According to an embodiment of the present application, the processing module is further used to: group the time feature elements, including: on the basis of uniformly grouping the time feature elements, performing window translation on the time feature elements to obtain a third grouping of the time feature elements.

[0023] According to an embodiment of the present application, the processing module is also used to: group the time feature elements, including: starting from the first time feature element, determining each time feature element and a preset number of time feature elements consecutively following the time feature element as a group to obtain a fourth group of time feature elements.

[0024] According to an embodiment of the present application, the acquisition module is further used to: obtain an initial feature map of a video frame in the video to be identified, including: projecting the video frame in the video to be identified to a feature space to obtain the initial feature map.

[0025] According to an embodiment of the present application, in a video recognition device: an acquisition module is used to input a video frame in a video to be recognized into a classification recognition model, and the three-dimensional projection layer in the classification recognition model projects the video frame to obtain an initial feature map; a processing module is used to perform N feature fusion processes in sequence starting from the initial feature map by N three-dimensional multilayer perceptron networks in the classification recognition model to output a target feature map N; wherein the N three-dimensional multilayer perceptron networks are connected in series, the input of the first three-dimensional multilayer perceptron network is the initial feature map, the feature map i input by the i-th three-dimensional multilayer perceptron network is the target feature map i-1 output by the i-1-th three-dimensional multilayer perceptron network, i and N are both positive integers, 1<i≤N; an identification module is used to input the target feature map N into the average pooling layer in the classification recognition model to perform an average pooling operation on the target feature map N, and input the mean feature map obtained after the average pooling operation into the fully connected layer in the classification recognition model to obtain the target category of the video to be recognized output by the fully connected layer.

[0026] According to an embodiment of the present application, in a processing module, a three-dimensional multilayer perceptron network includes a feature element mixing unit and a cross-channel perception unit, wherein: the feature element mixing unit mixes the feature elements of a feature map i in a feature dimension to obtain a mixed feature element, and fuses the mixed feature elements in all feature dimensions to obtain a fused feature map i; the cross-channel perception unit performs cross-channel perception on the fused feature map i and outputs a target feature map i.

[0027] According to an embodiment of the present application, in the processing module, the feature element mixing unit includes: a height feature element mixing subunit, a width feature element mixing subunit and a time feature element mixing subunit; the method also includes: the height feature element mixing subunit performs element mixing on the height feature elements in the feature graph i in the height dimension to obtain a mixed height feature element; the width feature element mixing subunit performs element mixing on the width feature elements in the feature graph i in the width dimension to obtain a mixed width feature element; the time feature element mixing subunit performs element mixing on the time feature elements in the feature graph i in the time dimension to obtain a mixed time feature element.

[0028] According to an embodiment of the present application, in a processing module, a transition layer is included between each two adjacent three-dimensional multi-layer perceptron networks, and the target feature map i is input into the transition layer, and the transition layer increases the number of feature extraction channels for the target feature map i and reduces the resolution of the target feature map i.

[0029] The third aspect embodiment of the present application proposes an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to implement a video recognition method as in the first aspect embodiment of the present application.

[0030] The fourth aspect embodiment of the present application proposes a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to implement the video recognition method as in the first aspect embodiment of the present application.

[0031] The fifth aspect embodiment of the present application proposes a computer program product, including a computer program, which, when executed by a processor, implements the video recognition method as in the first aspect embodiment of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0033] Figure 1 It is an exemplary schematic diagram of a video recognition method according to an embodiment of the present application.

[0034] Figure 2 It is a schematic diagram of a classification recognition model according to an embodiment of the present application.

[0035] Figure 3 It is a structural diagram of a three-dimensional multi-layer perceptron network according to an embodiment of the present application.

[0036] Figure 4 It is a schematic diagram of element mixing in the time dimension of the time feature elements in the feature graph i according to an embodiment of the present application.

[0037] FIG5(a) is a schematic diagram of uniformly grouping time feature elements to obtain a first group of time feature elements.

[0038] FIG5( b ) is a schematic diagram of discretely sampling the time feature elements to obtain a second group of the time feature elements.

[0039] FIG5(c) is a schematic diagram of performing window shifting on the time feature elements to obtain a third grouping of the time feature elements on the basis of uniform grouping of the time feature elements.

[0040] FIG5(d) is a schematic diagram of obtaining a fourth grouping of time feature elements by grouping each time feature element with a preset number of time feature elements that are consecutive after the first time feature element, starting from the first time feature element.

[0041] Figure 6 This is a schematic diagram of a video recognition device according to an embodiment of the present application.

[0042] Figure 7 It is a schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0043] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0044] Figure 1 is an exemplary implementation of a video recognition method proposed in this application, such as Figure 1 As shown, the video recognition method comprises the following steps:

[0045] S101, obtaining an initial feature map of a video frame in a video to be identified.

[0046] Determine the video to be identified, and project the video frames in the video to be identified to the feature space to obtain an initial feature map. The video to be identified may be a three-channel input video with a height of H, a width of W, and a duration of T. For example, the local window used when projecting the video frames in the video to be identified to the feature space may be 7×7×4 and the sampling interval may be 4×4×4. Each sampled local pixel is projected by a shared linear layer to a feature extraction channel number C. 1 feature space to obtain the initial feature map.

[0047] S102, starting from the initial feature map, perform feature fusion processing N times in sequence, wherein the feature map i processed by the i-th feature fusion processing is the target feature map i-1 output by the i-1-th feature fusion processing, i and N are both positive integers, 1<i≤N; the feature fusion processing includes: mixing the feature elements of the feature map i in the feature dimension on each feature extraction channel to obtain mixed feature elements, and fusing the mixed feature elements in all feature dimensions to obtain a fused feature map i, and performing cross-channel perception on the fused feature map i to obtain a target feature map i.

[0048] After obtaining the initial feature map, N feature fusion processes are performed in sequence starting from the initial feature map, wherein the feature map i processed by the i-th feature fusion process is the target feature map i-1 output by the i-1-th feature fusion process, and i and N are both positive integers, 1<i≤N.

[0049] Among them, the feature fusion processing includes: mixing the feature elements of the feature map i in the feature dimension on each feature extraction channel to obtain the mixed feature elements, fusing the mixed feature elements in all feature dimensions to obtain the fused feature map i, and performing cross-channel perception on the fused feature map i to obtain the target feature map i.

[0050] Among them, the feature dimensions include height dimension, width dimension and time dimension. When the feature elements of feature graph i in the feature dimension are mixed to obtain mixed feature elements, the height feature elements in feature graph i are mixed in the height dimension to obtain mixed height feature elements; the width feature elements in feature graph i are mixed in the width dimension to obtain mixed width feature elements; the time feature elements in feature graph i are mixed in the time dimension to obtain mixed time feature elements.

[0051] Optionally, in order to reduce computational complexity and the number of parameters, the present application may group the time feature elements on the time feature dimension and then mix the elements.

[0052] S103, based on the target feature map N output by the Nth feature fusion processing, performing category recognition on the video to be recognized to obtain the target category of the video to be recognized.

[0053] According to the target feature map N output by the Nth feature fusion process, the video to be identified is identified to obtain the target category of the video to be identified. Exemplarily, the target feature map N may be averaged and pooled, and the mean feature map obtained after the average pooling operation is processed to obtain the target category of the video to be identified.

[0054] The target category of the video to be identified refers to the video category corresponding to the video to be identified, such as automobiles, pets, and food.

[0055] The video recognition method proposed in the embodiment of the present application decomposes the mixture between the feature elements in the video to be identified into the interaction of three independent feature dimensions, and then uses linear projection to combine the interaction results of different feature dimensions, thereby reducing the computational complexity and the number of parameters, and achieving a better balance between accuracy and computational efficiency in the video recognition problem.

[0056] The video recognition method proposed in this application can be implemented based on a classification recognition model. The overall structure of the classification recognition model includes a three-dimensional projection layer, N three-dimensional multi-layer perceptron networks, a transition layer, an average pooling layer and a fully connected layer.

[0057] Figure 2 It is a schematic diagram of the classification recognition model, such as Figure 2As shown, taking the number of three-dimensional multi-layer perceptron networks in the classification recognition model as 4 as an example, the specific steps of the video recognition method proposed in this application include:

[0058] The video frame in the video to be identified is input into the classification recognition model, and the three-dimensional projection layer in the classification recognition model projects the video frame to obtain the initial feature map. The video to be identified can be a three-channel input video with a height of H, a width of W, and a duration of T. The local window used by the three-dimensional projection layer is 7×7×4 and the sampling interval is 4×4×4. Each sampled local pixel is projected by a shared linear layer to a feature extraction channel with the number of C. 1 Therefore, the feature dimension of the initial feature map output by the three-dimensional projection layer is Since the number of feature extraction channels is C 1 , the initial feature map output by the 3D projection layer can be expressed as

[0059] After obtaining the initial feature map output by the 3D projection layer, the 4 3D multilayer perceptron networks in the classification recognition model perform 4 feature fusion processes in sequence starting from the initial feature map, so that the final target feature map is output by the 4th 3D multilayer perceptron network. Figure 4 , where four three-dimensional multilayer perceptron networks are connected in series, the input of the first three-dimensional multilayer perceptron network is the initial feature map, the feature map i input to the i-th three-dimensional multilayer perceptron network is the target feature map i-1 output by the i-1-th three-dimensional multilayer perceptron network, i and N are both positive integers, 1<i≤4.

[0060] Specifically, the input of the first 3D multilayer perceptron network is the initial feature map The input of the first 3D multilayer perceptron network is the target feature Figure 1

[0061] The input of the second 3D multilayer perceptron network is the target feature Figure 1 The output of the second 3D multilayer perceptron network is the target feature Figure 2

[0062] The input of the third 3D multilayer perceptron network is the target feature Figure 2 The output of the third three-dimensional multilayer perceptron network is the target feature Figure 3

[0063] The input of the fourth 3D multilayer perceptron network is the target feature Figure 3 The output of the fourth 3D multilayer perceptron network is the target feature Figure 4

[0064] The beginning of the second three-dimensional multilayer perceptron network, the third three-dimensional multilayer perceptron network and the fourth three-dimensional multilayer perceptron network all include a transition layer, and the target feature map i output by the previous three-dimensional multilayer perceptron network is input into the transition layer, and the transition layer increases the number of feature extraction channels for the target feature map i, that is, C 4 >C 3 >C 2 >C 1 ; and reduce the resolution of the target feature map i. After the transition layer, in the three-dimensional multi-layer perceptron network, the number of feature extraction channels of the target feature map i and the resolution of the target feature map i no longer change.

[0065] The target features output by the fourth three-dimensional multi-layer perceptron network Figure 4 The average pooling layer in the input classification recognition model is used to classify the target features. Figure 4 An average pooling operation is performed, and the mean feature map obtained after the average pooling operation is input into the fully connected layer in the classification recognition model to obtain the target category of the video to be recognized output by the fully connected layer.

[0066] The structure of the three-dimensional multi-layer perceptron network is described in detail below. Figure 3 is a schematic diagram of the structure of a three-dimensional multi-layer perceptron network shown in this application, such as Figure 3 As shown in the figure, the three-dimensional multi-layer perceptron network includes a token-mixing MLP and a channel-based perceptron, and the token-mixing MLP is subdivided into a height-based, width-based, time-based, and linear-projection subunits. Figure 3 As shown, both the feature element mixing unit and the cross-channel perception unit contain layer normalization (LayerNorm) and residual structure to help optimize the perceptron neural network.

[0067] Among them: the feature element mixing unit in the three-dimensional multilayer perceptron network is used to mix the feature elements of the feature map i in the feature dimension to obtain the mixed feature element, and the linear projection subunit fuses the mixed feature element in all feature dimensions to obtain the fused feature map i. Among them, the feature dimension includes the height dimension, the width dimension and the time dimension. When the feature elements of the feature map i are mixed in the feature dimension to obtain the mixed feature element, the height feature element mixing subunit mixes the height feature elements in the feature map i in the height dimension to obtain the mixed height feature element; the width feature element mixing subunit mixes the width feature elements in the feature map i in the width dimension to obtain the mixed width feature element; the time feature element mixing subunit mixes the time feature elements in the feature map i in the time dimension to obtain the mixed time feature element.

[0068] The present application proposes to mix feature elements in the feature dimensions of height, width and time only along one feature dimension at a time, so that the number of feature elements input for each feature element mixing can be significantly reduced. It can be expressed as a linear mapping of the mixture of elements in three feature dimensions:

[0069]

[0070] Where X H represents the mixed height feature element output by the height feature element mixing subunit, X W represents the mixed width feature element output by the width feature element mixing subunit, X T It represents the mixed time feature element output by the time feature element mixing subunit, and FC is the result of mixing different dimensions by the fully connected layer in the linear projection subunit. For the height feature element mixing subunit and the width feature element mixing subunit, the recurrent fully connected layer can be used to model the height feature element and the width feature element.

[0071] The cross-channel perception unit in the three-dimensional multi-layer perceptron network consists of two linear layers. A nonlinear activation layer is added between the two linear layers to improve the fitting ability of the network. The cross-channel perception unit is used to perform cross-channel perception on the fused feature map i and output the target feature map i.

[0072] Specifically, the target feature map i-1 of the input of any three-dimensional multilayer perceptron network is represented as feature X, and the calculation formula of the three-dimensional multilayer perceptron network is:

[0073] Y=Token-mixing-MLP(LN(X))+X

[0074] Z=Channel-MLP(LN(Y))+Y

[0075] Where LN is layer normalization (Layer Norm), the output Z of the three-dimensional multi-layer perceptron network, that is, the target feature map i, will be used as the input of the next three-dimensional multi-layer perceptron network.

[0076] The video recognition method proposed in the embodiment of the present application decomposes the mixture between the characteristic elements in the video to be recognized into the interaction of three independent characteristic dimensions, and then uses linear projection to combine the interaction results of different characteristic dimensions. In the element mixing of the time dimension, a grouped temporal mixing operation is proposed, which divides the characteristic elements in the video to be recognized into different groups in chronological order, and independently mixes information within each group, thereby reducing the computational complexity and number of parameters of the temporal mixing operation, and achieving a better balance between accuracy and computational efficiency in the video recognition problem.

[0077] Figure 4 is an exemplary implementation of a video recognition method proposed in this application, such as Figure 2 As shown, based on the above embodiment, the time feature elements in the feature graph i are mixed in the time dimension to obtain mixed time feature elements, including the following steps:

[0078] S401, obtaining time feature elements of feature graph i in the time dimension, and grouping the time feature elements.

[0079] Get the time feature elements of feature graph i in the time dimension. If all the time feature elements in the video to be identified are regarded as a group, all the time feature elements are directly mixed using linear mapping. When the input feature is When , the output of this timing mixture is:

[0080]

[0081] where W∈R TC×TC is a linear projection matrix. Although this method of putting all time feature elements in one group can obtain longer time element correlation, its computational complexity reaches O(HWT 2 C 2 ), the number of parameters is O(T 2 C 2 ), both indicators will increase quadratically with the length of the video to be identified.

[0082] Therefore, in order to reduce the computational complexity and the number of parameters, the present application groups the time feature elements.

[0083] When grouping time feature elements, the following four grouping methods are introduced:

[0084] As an achievable method, it is called the short-range GTM method: when grouping the time feature elements, the total number T of the time feature elements corresponding to the video to be identified and the total number S of the time feature elements that each group needs to contain after grouping can be obtained, that is, the time feature elements need to be evenly divided into T / S groups in order to obtain the first group of time feature elements. For example, FIG5(a) is a schematic diagram of uniformly grouping the time feature elements to obtain the first group of time feature elements. As shown in FIG5(a), it is assumed that there are 6 time feature elements, namely, X, S, and S. 1 , X 2 , X 3 , X 4 , X 5 and X 6 , every two time feature elements are grouped into one group, then X 1 and X 2 As a group, X 3 and X 4 As a group, X 5 and X 6 As a group, the grouping of the time feature elements obtained in this way is recorded as the first grouping. In another exemplary example, assuming that the total number T of time feature elements corresponding to the video to be identified is 100, and the total number S of time feature elements that each group needs to contain after grouping is 10, then the time feature elements need to be divided into 100 / 10=10 groups in order, that is, the 1st time feature element to the 10th time feature element is a group, the 11th time feature element to the 20th time feature element is a group, and so on, the 91st time feature element to the 100th time feature element is a group, and a total of 10 groups are divided.

[0085] As a feasible method, it is called the long-range GTM method: compared with the short-range GTM method, in order to obtain a longer-term correlation, the total number T of time feature elements corresponding to the video to be identified and the total number S of time feature elements that each group needs to contain after grouping can be obtained. When grouping the time feature elements, the time feature elements can be discretely sampled in sequence, and the time feature elements in each group are discontinuous to obtain the second grouping of the time feature elements. Figure 5(b) is a schematic diagram of discretely sampling the time feature elements to obtain the second grouping of the time feature elements. As shown in Figure 5(b), it is assumed that there are 6 time feature elements, namely X 1 , X 2 , X 3 , X 4 , X 5 and X 6, every two time feature elements are grouped together, and after discrete sampling of the time feature elements in sequence, X 1 and X 4 As a group, X 2 and X 5 As a group, X 3 and X 6 As a group, the grouping of the time feature elements obtained in this way is recorded as the second grouping. In another exemplary embodiment, assuming that the total number T of time feature elements corresponding to the video to be identified is 100, and the total number S of time feature elements that each group needs to contain after grouping is 10, then the time feature elements need to be divided into 100 / 10=10 groups in order, then the 1st, 11th, 21st, 31st...91st time feature elements are grouped into one group, the 2nd, 12th, 22nd, 32nd...92nd time feature elements are grouped into one group, and so on, the 10th, 20th, 30th,...100th time feature elements are grouped into one group, and a total of 10 groups are divided.

[0086] As an achievable method, it is called a shift-window GTM method: the total number T of time feature elements corresponding to the video to be identified and the total number S of time feature elements that each group needs to contain after grouping can be obtained. When the time feature elements are grouped, on the basis of uniformly grouping the time feature elements in the short-range GTM method as described above, the time feature elements are window-shifted to obtain a third group of time feature elements. Among them, window-shifting the time feature elements means that after the group is fixed, the time feature elements in each group are shifted. For example, assuming that the total number T of time feature elements corresponding to the video to be identified is 100, and the total number S of time feature elements that each group needs to contain after grouping is 10, the time feature elements need to be divided into 100 / 10=10 groups in order. After the above uniform grouping, the 1st to 10th time feature elements are obtained as a group, the 11th to 20th time feature elements are obtained as a group, and so on, the 91st to 100th time feature elements are obtained as a group. On the basis of grouping the time feature elements into a group, the time feature elements are window-shifted. If the number of window-shifts for the time feature elements is 5, the 6th to 15th time feature elements are grouped into a group, the 16th to 25th time feature elements are grouped into a group, and so on, the 86th to 95th time feature elements are grouped into a group. In addition, the first 1st to 5th time feature elements and the last 96th to 100th time feature elements are grouped into a group, and a total of 10 groups are divided. In another exemplary embodiment, FIG5(c) is a schematic diagram of window-shifting the time feature elements to obtain the third grouping of the time feature elements on the basis of uniform grouping of the time feature elements. As shown in FIG5(c), it is assumed that there are 6 time feature elements, namely, X 1 , X 2 , X 3 , X 4 , X 5 and X 6 , every two time feature elements are grouped into one group. On the basis of uniform grouping of the time feature elements, the time feature elements are window-shifted, and then X 1 and X 6 As a group, X 2 and X 3 As a group, X 4 and X 5 As a group, the grouping of the time feature elements obtained in this way is recorded as the third grouping.

[0087] As another feasible method, it is called the shift-token GTM method: the total number T of time feature elements corresponding to the video to be identified and the total number S of time feature elements that each group needs to contain after grouping can be obtained. When grouping the time feature elements, each time feature element and a preset number of time feature elements that are consecutive after the time feature element can be determined as a group starting from the first time feature element to obtain the fourth group of time feature elements. FIG5(d) is a schematic diagram of obtaining the fourth group of time feature elements by determining each time feature element and a preset number of time feature elements that are consecutive after the time feature element starting from the first time feature element. As shown in FIG5(d), it is assumed that there are 6 time feature elements in total, namely, X, , and , respectively. 1 , X 2 , X 3 , X 4 , X 5 and X 6 , every two time feature elements are grouped together, and starting from the first time feature element, each time feature element and a preset number of time feature elements after the time feature element are determined as a group, then X 1 and X 6 As a group, X 2 and X 1 As a group, X 3 and X 2 As a group, X 4 and X 3 As a group, X 5 and X 4 As a group, X 6 and X 5 As a group, the grouping of the time feature elements obtained in this way is recorded as the fourth grouping. In another exemplary embodiment, assuming that the total number T of time feature elements corresponding to the video to be identified is 100, and the total number S of time feature elements that each group needs to contain after grouping is 10, then starting from the first time feature element, each time feature element and the 9 consecutive time feature elements after the time feature element are determined as a group, that is, the first time feature element to the tenth time feature element are grouped, the second time feature element to the eleventh time feature element are grouped, the third time feature element to the twelfth time feature element are grouped, and so on.

[0088] It should be noted that, in the time feature element mixed subunits corresponding to different three-dimensional multi-layer perceptron networks in the above classification and recognition model, the above four grouping methods can be used in combination, or only one grouping method can be used to group the time feature elements in the time feature element mixed subunits corresponding to different three-dimensional multi-layer perceptron networks.

[0089] S402, mixing the time feature elements within each group to obtain mixed time feature elements.

[0090] The time feature elements in each group of the time feature element mixing subunit obtained above are mixed to obtain a mixed time feature element. Optionally, a linear projection method can be used to mix the time feature elements in each group. Y in FIG. 5(a), FIG. 5(b), FIG. 5(c) and FIG. 5(d) 1 , Y 2 , Y 3 , Y 4 , Y 5 and Y 6 They respectively represent the mixed time feature elements after mixing by linear projection.

[0091] The video recognition method proposed in the embodiment of the present application proposes a grouped temporal mixing operation in the element mixing in the time dimension, which divides the characteristic elements in the video to be identified into different groups in chronological order, and performs information mixing independently within each group, thereby reducing the computational complexity and number of parameters of the temporal mixing operation.

[0092] Figure 6 is a schematic diagram of a video recognition device shown in this application, such as Figure 6 As shown, the video recognition device 600 includes an acquisition module 601, a processing module 602 and a recognition module 603, wherein:

[0093] An acquisition module 601 is used to acquire an initial feature map of a video frame in a video to be identified;

[0094] The processing module 602 is used to perform N feature fusion processes in sequence starting from the initial feature map, wherein the feature map i processed by the i-th feature fusion process is the target feature map i-1 output by the i-1-th feature fusion process, i and N are both positive integers, 1<i≤N; the feature fusion process includes: mixing the feature elements of the feature map i in the feature dimension on each feature extraction channel to obtain a mixed feature element, and fusing the mixed feature elements in all feature dimensions to obtain a fused feature map i, and performing cross-channel perception on the fused feature map i to obtain a target feature map i;

[0095] The recognition module 603 is used to perform category recognition on the video to be recognized based on the target feature map N output by the Nth feature fusion processing to obtain the target category of the video to be recognized.

[0096] The video recognition device proposed in the embodiment of the present application decomposes the mixture between the feature elements in the video to be recognized into the interaction of three independent feature dimensions, and then uses linear projection to combine the interaction results of different feature dimensions, thereby reducing the computational complexity and the number of parameters, and achieving a better balance between accuracy and computational efficiency in the video recognition problem.

[0097] Furthermore, the processing module 602 is also used to: mix the feature elements of the feature graph i in the feature dimension to obtain a mixed feature element, including: mixing the elements of the height feature elements in the feature graph i in the height dimension to obtain a mixed height feature element; mixing the elements of the width feature elements in the feature graph i in the width dimension to obtain a mixed width feature element; mixing the elements of the time feature elements in the feature graph i in the time dimension to obtain a mixed time feature element.

[0098] Furthermore, the processing module 602 is also used to: mix the time feature elements in the feature graph i in the time dimension to obtain a mixed time feature element, including: obtaining the time feature elements of the feature graph i in the time dimension and grouping the time feature elements; mixing the time feature elements within each group to obtain a mixed time feature element.

[0099] Furthermore, the processing module 602 is further configured to: group the time feature elements, including: uniformly grouping the time feature elements to obtain a first group of the time feature elements.

[0100] Furthermore, the processing module 602 is further configured to: group the time feature elements, including: discretely sampling the time feature elements to obtain a second group of the time feature elements.

[0101] Furthermore, the processing module 602 is further configured to: group the time feature elements, including: performing window translation on the time feature elements on the basis of uniformly grouping the time feature elements, to obtain a third grouping of the time feature elements.

[0102] Furthermore, the processing module 602 is also used to: group the time feature elements, including: starting from the first time feature element, determining each time feature element and a preset number of time feature elements consecutively following the time feature element as a group to obtain a fourth group of time feature elements.

[0103] Furthermore, the acquisition module 601 is further used to: acquire an initial feature map of a video frame in the video to be identified, including: projecting the video frame in the video to be identified to a feature space to acquire the initial feature map.

[0104] Furthermore, in the video recognition device 600: an acquisition module 601 is used to input the video frame in the video to be recognized into the classification recognition model, and the three-dimensional projection layer in the classification recognition model projects the video frame to obtain an initial feature map; a processing module 602 is used to perform N feature fusion processes in sequence starting from the initial feature map by N three-dimensional multilayer perceptron networks in the classification recognition model to output a target feature map N; wherein the N three-dimensional multilayer perceptron networks are connected in series, the input of the first three-dimensional multilayer perceptron network is the initial feature map, the feature map i input by the i-th three-dimensional multilayer perceptron network is the target feature map i-1 output by the i-1-th three-dimensional multilayer perceptron network, i and N are both positive integers, 1<i≤N; an identification module 603 is used to input the target feature map N into the average pooling layer in the classification recognition model to perform an average pooling operation on the target feature map N, and input the mean feature map obtained after the average pooling operation into the fully connected layer in the classification recognition model to obtain the target category of the video to be recognized output by the fully connected layer.

[0105] Furthermore, in the processing module 602, the three-dimensional multilayer perceptron network includes a feature element mixing unit and a cross-channel perception unit, wherein: the feature element mixing unit mixes the feature elements of the feature map i in the feature dimension to obtain a mixed feature element, and fuses the mixed feature elements in all feature dimensions to obtain a fused feature map i; the cross-channel perception unit performs cross-channel perception on the fused feature map i and outputs a target feature map i.

[0106] Furthermore, in the processing module 602, the feature element mixing unit includes: a height feature element mixing subunit, a width feature element mixing subunit and a time feature element mixing subunit; the method also includes: the height feature element mixing subunit performs element mixing on the height feature elements in the feature graph i in the height dimension to obtain a mixed height feature element; the width feature element mixing subunit performs element mixing on the width feature elements in the feature graph i in the width dimension to obtain a mixed width feature element; the time feature element mixing subunit performs element mixing on the time feature elements in the feature graph i in the time dimension to obtain a mixed time feature element.

[0107] Furthermore, in the processing module 602, a transition layer is included between each two adjacent three-dimensional multi-layer perceptron networks, and the target feature map i is input into the transition layer, and the transition layer increases the number of feature extraction channels for the target feature map i and reduces the resolution of the target feature map i.

[0108] In order to implement the above embodiment, the present application embodiment also proposes an electronic device 700, such as Figure 7As shown, the electronic device 700 includes: a processor 701 and a memory 702 communicatively connected to the processor, the memory 702 stores instructions executable by at least one processor, and the instructions are executed by at least one processor 701 to implement the video recognition method shown in the above embodiment.

[0109] In order to implement the above embodiment, the embodiment of the present application also proposes a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to implement the video recognition method shown in the above embodiment.

[0110] In order to implement the above embodiment, the embodiment of the present application also proposes a computer program product, including a computer program, which implements the video recognition method shown in the above embodiment when executed by a processor.

[0111] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0112] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0113] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. A video recognition method, It is characterized in that include: Obtaining an initial feature map of a video frame in a video to be identified; Starting from the initial feature map, N feature fusion processes are performed in sequence, wherein the feature map i processed by the i-th feature fusion process is the target feature map i-1 output by the i-1-th feature fusion process, and i and N are both positive integers, 1<i≤N; The feature fusion processing includes: mixing the feature elements of the feature map i in the feature dimension on each feature extraction channel to obtain a mixed feature element, fusing the mixed feature element in the full feature dimension to obtain a fused feature map i, and performing cross-channel perception on the fused feature map i to obtain a target feature map i; Based on the target feature graph N output by the Nth feature fusion processing, performing category recognition on the video to be recognized to obtain the target category of the video to be recognized; The mixing of feature elements of the feature graph i in the feature dimension to obtain mixed feature elements includes: Obtaining the time feature elements of the feature graph i in the time dimension, and uniformly grouping the time feature elements to obtain a first group of the time feature elements, or discretely sampling the time feature elements to obtain a second group of the time feature elements, or, on the basis of uniformly grouping the time feature elements, performing window translation on the time feature elements to obtain a third group of the time feature elements, or, starting from the first time feature element, determining each of the time feature elements and a preset number of time feature elements consecutively following the time feature element as a group to obtain a fourth group of the time feature elements; The time feature elements within each of the groups are mixed to obtain a mixed time feature element.

2. The method according to claim 1, It is characterized in that The step of mixing the feature elements of the feature graph i in the feature dimension to obtain a mixed feature element further includes: Performing element mixing on the height characteristic elements in the feature graph i in the height dimension to obtain mixed height characteristic elements; The width feature elements in the feature map i are mixed in the width dimension to obtain mixed width feature elements.

3. The method according to any one of claims 1 to 2, It is characterized in that The step of obtaining an initial feature map of a video frame in a video to be identified includes: Project the video frames in the video to be identified onto the feature space to obtain the initial feature map.

4. The method according to claim 1, It is characterized in that The method comprises: Inputting a video frame in the video to be identified into a classification and recognition model, and projecting the video frame by a three-dimensional projection layer in the classification and recognition model to obtain the initial feature map; The N three-dimensional multilayer perceptron networks in the classification and recognition model sequentially perform the feature fusion process N times starting from the initial feature map to output the target feature map N; The N three-dimensional multilayer perceptron networks are connected in series, the input of the first three-dimensional multilayer perceptron network is an initial feature map, the feature map i input by the i-th three-dimensional multilayer perceptron network is a target feature map i-1 output by the i-1-th three-dimensional multilayer perceptron network, and i and N are both positive integers, 1<i≤N; The target feature map N is input into the average pooling layer in the classification recognition model to perform an average pooling operation on the target feature map N, and the mean feature map obtained after the average pooling operation is input into the fully connected layer in the classification recognition model to obtain the target category of the video to be identified output by the fully connected layer.

5. The method according to claim 4, It is characterized in that The three-dimensional multi-layer perceptron network includes a feature element mixing unit and a cross-channel perception unit, wherein: The feature element mixing unit mixes the feature elements of the feature map i in a feature dimension to obtain a mixed feature element, and fuses the mixed feature element in all feature dimensions to obtain a fused feature map i; The cross-channel perception unit performs cross-channel perception on the fused feature map i and outputs a target feature map i.

6. The method according to claim 5, It is characterized in that The feature element mixing unit includes: a height feature element mixing subunit, a width feature element mixing subunit and a time feature element mixing subunit; the method further includes: The highly characteristic element mixing subunit performs element mixing on the highly characteristic elements in the characteristic graph i in the height dimension to obtain a mixed highly characteristic element; The width feature element mixing subunit performs element mixing on the width feature elements in the feature graph i in the width dimension to obtain a mixed width feature element; The time feature element mixing subunit mixes the time feature elements in the feature graph i in the time dimension to obtain a mixed time feature element.

7. The method according to any one of claims 4 to 6, It is characterized in that A transition layer is included between each two adjacent three-dimensional multilayer perceptron networks, and the target feature map i is input into the transition layer. The transition layer increases the number of feature extraction channels for the target feature map i and reduces the resolution of the target feature map i.

8. A video recognition device, It is characterized in that include: An acquisition module, used to acquire an initial feature map of a video frame in a video to be identified; A processing module, used to perform N feature fusion processes in sequence starting from the initial feature map, wherein the feature map i processed by the i-th feature fusion process is the target feature map i-1 output by the i-1-th feature fusion process, and i and N are both positive integers, 1<i≤N; The feature fusion processing includes: mixing the feature elements of the feature map i in the feature dimension on each feature extraction channel to obtain a mixed feature element, fusing the mixed feature element in the full feature dimension to obtain a fused feature map i, and performing cross-channel perception on the fused feature map i to obtain a target feature map i; A recognition module, configured to perform category recognition on the video to be recognized based on a target feature map N output by the Nth feature fusion process, so as to obtain a target category of the video to be recognized; The mixing of feature elements of the feature graph i in the feature dimension to obtain mixed feature elements includes: Obtaining the time feature elements of the feature graph i in the time dimension, and uniformly grouping the time feature elements to obtain a first group of the time feature elements, or discretely sampling the time feature elements to obtain a second group of the time feature elements, or, on the basis of uniformly grouping the time feature elements, performing window translation on the time feature elements to obtain a third group of the time feature elements, or, starting from the first time feature element, determining each of the time feature elements and a preset number of time feature elements consecutively following the time feature element as a group to obtain a fourth group of the time feature elements; The time feature elements within each of the groups are mixed to obtain a mixed time feature element.

9. An electronic device, include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, in, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.

11. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.