PSC-TNet video action recognition method based on fusion of spatial features and frame difference information

By adopting a three-branch structure and feature enhancement module in image action recognition, combining time and space features, the shortcomings of deep learning methods in capturing spatiotemporal information and feature extraction are solved, and more efficient and accurate action recognition is achieved.

CN119964244AActive Publication Date: 2025-05-09ZHEJIANG SCI-TECH UNIV

Patent Information

Application Number
CN202510054148.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-09
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

The image action recognition method based on deep learning has shortcomings in capturing spatio-temporal information and feature extraction, and is easily disturbed by redundant information.

Method used

A PSC-TNet video action recognition method based on fusion spatial features and frame difference information is proposed. A three-branch structure (appearance branch, motion branch and action detail branch) is adopted and the temporal feature extraction module TtS and the spatial feature enhancement module PSC are designed to extract and enhance spatial and temporal features more effectively.

Benefits of technology

By combining temporal and spatial characteristics, the PSC-TNet method can more accurately identify actions in the video, improve the accuracy and efficiency of action recognition, and reduce the dependence on redundant information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964244A_ABST
    Figure CN119964244A_ABST
Patent Text Reader

Abstract

The invention discloses a PSC-TNet video action recognition method based on fusion of spatial features and frame difference information. According to the method, a video is subjected to image frame preprocessing and then is input into a PSC-TNet video action recognition model based on fusion spatial features and frame difference information, and the PSC-TNet video action recognition model comprises three branches based on different frame intervals and an action classification module. The three branches are respectively a motion branch, an appearance branch and an action detail branch which comprise spatial information, motion information and motion detail information, and the original video is respectively input into the motion branch, the appearance branch and the action detail branch to output respective action characteristics; and all the identification features are spliced and then input to an action classification module to obtain an action classification result. According to the method, information can be extracted from the time features of the video image to adjust the spatial features of the video image, so that the spatial features with the weights of the key space parts enhanced are obtained, and more accurate video image action recognition is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a method for action recognition, and in particular to an image action recognition method that associates time information and space information. Background Art

[0002] Video image action recognition is an important research direction in the field of computer vision, which aims to enable computers to automatically understand human actions in video content. This technology involves extracting action features from video sequences and analyzing these features using machine learning or deep learning models to identify ongoing actions or activities.

[0003] Action recognition is very important in many practical applications, such as video surveillance, human-computer interaction, sports analysis, health care, and entertainment industries. In the field of security monitoring, action recognition can help detect abnormal behavior or potential security threats. In sports events, it can be used to analyze athletes' performance and techniques. In healthcare, action recognition can monitor patients' recovery progress and daily activities to provide better care services, etc.

[0004] With the development of deep learning technology, deep learning has achieved great success in image classification, object detection, speech recognition and other fields. Deep learning can automatically learn high-level feature representations from data, avoiding the complexity and limitations of manually designed features. The accuracy and efficiency of deep learning in human action recognition have been significantly improved. It is usually based on models such as convolutional neural network CNN, recurrent neural network RNN ​​and graph neural network GNN, which can learn complex spatiotemporal features from video data. The addition of attention mechanism also further improves the performance of action recognition. There are two main methods of CNN: one is based on two-stream network, which divides the video into RGB frames and optical flow frames, and extracts spatial and temporal features respectively with convolutional neural network (CNN), and then combines the two features for classification; the other is based on three-dimensional convolutional neural network (3D CNN), which can directly perform three-dimensional convolution on video blocks, learn spatial and temporal features at the same time, and then use fully connected layers or pooling layers for classification. Both methods have their own advantages and disadvantages. For example, two-stream network can better capture the details of the action, but it requires additional calculation of optical flow; 3D CNN can better model the overall action, but it requires more parameters and computing resources. Summary of the invention

[0005] In order to solve the problems that image action recognition methods based on deep learning cannot capture spatiotemporal information sufficiently and feature extraction is easily interfered by redundant information, the present invention proposes a PSC-TNet video action recognition method based on the fusion of spatial features and frame difference information for video image processing.

[0006] The network framework of the present invention adopts a three-branch structure of appearance branch, motion branch and motion detail branch to extract spatial information, motion information and action detail information. The present invention designs two new plug-and-play modules, namely, the temporal feature extraction module TtS and the spatial feature enhancement module PSC. The temporal feature extraction module TtS aims to obtain weight scores based on the differential features of a certain frame interval, and more effectively locate the motion difference information of two frames to strengthen the positional features of the space where the character's action is located, so as to achieve secondary enhancement of the spatial features based on the temporal features of the video. The spatial feature enhancement module PSC, different from the traditional global pooling of the channel dimension, which leads to feature loss, adopts a filtering processing method, which compresses the channel dimension while keeping the orthogonal dimension, that is, the spatial dimension, at high resolution, and then performs matrix multiplication on the two to achieve the spatial feature weight without global pooling of the channel. The PSC spatial feature enhancement module adopts a cross-multiplication strategy, and the features are processed by convolution kernels of different scales, that is, the spatial feature weights obtained by the 1×1 convolution kernel are used to adjust the features of the 3×3 convolution kernel, and the spatial feature weights obtained by the 3×3 convolution kernel are used to adjust the features of the 1×1 convolution kernel to achieve cross-spatial information aggregation. The PSC-T module proposed in the present invention is the fusion of the PSC spatial feature enhancement module and the TtS temporal feature extraction module. The spatial features extracted by the TtS feature extraction module are added to the new weights obtained by adding the weights after cross-multiplication in the PSC module, and the new spatial feature weights are output, which has achieved good results in the action recognition classification task.

[0007] The technical solution of the present invention is as follows:

[0008] The method of the present invention is to preprocess the image frames of the input video, and the preprocessed video is input as the original video into a PSC-TNet video action recognition model based on fusion spatial features and frame difference information to extract and obtain the action recognition result; the PSC-TNet video action recognition model includes three branches based on different frame intervals and an action classification module, the three branches are respectively a motion branch, an appearance branch and an action detail branch, the original video is respectively input into the motion branch, the appearance branch and the action detail branch to output respective action features, and then all the recognition features are spliced ​​and input into the action classification module to obtain the action classification result.

[0009] The input video is an image video to be action recognized.

[0010] The input video is sampled at a higher frequency, a lower frequency, and a medium frequency to obtain a higher frequency video frame, a lower frequency video frame, and a medium frequency video frame to form an original video.

[0011] The middle frequency is a frequency between the higher frequency and the lower frequency.

[0012] In the PSC-TNet video action recognition model, higher-frequency video frames in the original video are used as feature tensors containing motion information and are passed into the motion branch, lower-frequency video frames in the original video are used as feature tensors containing spatial information and are passed into the appearance branch, and medium-frequency video frames in the original video are used as feature tensors containing action detail information and are passed into the action detail branch; the three branches all contain residual layers, each residual layer includes a residual module, the features of the intermediate outputs of the motion branch and the action detail branch are passed to the appearance branch for channel splicing and feature fusion, and then input into the residual layer of the appearance branch for convolution processing, and the three branches respectively output recognition features containing spatial information, motion information and motion detail information, and finally the recognition features output by the three branches are spliced ​​and fused on each channel of the appearance branch and input into the action classification module to obtain the final action recognition result.

[0013] The classifier adopts a fully connected layer.

[0014] In the PSC-TNet video action recognition model, the motion branch and the appearance branch both include a convolution block connected in sequence and four consecutive residual layers containing a spatial feature enhancement module PSC; the action detail branch includes a temporal feature extraction module TtS, a convolution block and four consecutive spatiotemporal feature fusion modules connected in sequence, each spatiotemporal feature fusion module includes a temporal feature extraction module TtS and a residual layer containing a spatial feature enhancement module PSC in parallel, the input of the spatiotemporal feature fusion module is respectively input into the temporal feature extraction module TtS and the residual layer to obtain respective results, and the respective results are added and fused as the output of the spatiotemporal feature fusion module; in the spatiotemporal feature fusion module, the spatial feature weight generated in the last residual layer is added to the output of the temporal feature extraction module to generate a new spatial feature weight to adjust the spatial feature.

[0015] The four residual layers of the motion branch, appearance branch and action detail branch respectively contain three, four, six and three residual modules in the transmission order; the convolution blocks of the motion branch and the action detail branch and the convolution blocks of the appearance branch are spliced ​​on the channel of the appearance branch and input into the first residual layer of the appearance branch, and the nth residual layer of the motion branch and the action detail branch and the nth residual layer of the appearance branch are spliced ​​on the channel of the appearance branch and input into the n+1th residual layer of the appearance branch.

[0016] The residual layers in the motion branch, the appearance branch and the action detail branch each include a plurality of residual modules connected in sequence.

[0017] The residual module includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a PSC-T attention mechanism module and a fourth convolutional layer. The first convolutional layer, the second convolutional layer, the third convolutional layer and the PSC-T attention mechanism module are connected in sequence. The input of the residual module is respectively input into the first convolutional layer and the fourth convolutional layer. The output of the PSC-T attention mechanism module and the output of the fourth convolutional layer are added as the output of the residual module; the first convolutional layer and the second convolutional layer are mainly composed of a convolution operation, a batch normalization operation, and an activation operation connected in sequence. The third convolutional layer is mainly composed of a convolution operation and a batch normalization operation connected in sequence, and the fourth convolutional layer is composed of one convolution operation.

[0018] Therefore, the input features of the residual module are respectively input into multiple convolutional layers and the PSC-T attention mechanism module to obtain the output features of the residual module.

[0019] The PSC-T attention mechanism module of the residual module in the residual layer of the motion branch, appearance branch and action detail branch adopts the spatial feature enhancement module PSC.

[0020] More specifically, the motion branch includes a first convolution block and four residual layers of a first residual layer to a fourth residual layer connected in sequence, the appearance branch includes a second convolution block and four residual layers of a fifth residual layer to an eighth residual layer connected in sequence, and each residual layer of the motion branch and the appearance branch includes a spatial feature enhancement module.

[0021] The action detail branch includes four residual layers, namely the third convolution block and the ninth residual layer to the twelfth residual layer, which are connected in sequence. A temporal feature extraction module is included between each residual layer of the action detail branch and before the third convolution block. A spatiotemporal feature fusion module is included in each residual layer of the action detail branch. The four residual layers of the motion branch, appearance branch and action detail branch have 3, 4, 6 and 3 residual modules respectively. The input of the action detail branch passes through the output of the first, second and third convolution layers in the residual module, and then is combined with the output of the temporal feature extraction module as the spatiotemporal feature fusion. The output of the first convolution block and the output of the third convolution block are spliced ​​and fused on the channel dimension of the second convolution block, the output of the first residual layer and the output of the ninth residual layer are spliced ​​and fused on the channel dimension of the fifth residual layer, the output of the second residual layer and the output of the sixth residual layer are spliced ​​and fused on the channel dimension of the tenth residual layer, the output of the third residual layer and the output of the seventh residual layer are spliced ​​and fused on the channel dimension of the eleventh residual layer, and the output of the fourth residual layer and the output of the eighth residual layer are spliced ​​and fused on the channel dimension of the twelfth residual layer.

[0022] The first convolution block, the second convolution block and the third convolution block have the same structure, and all include a convolution layer, a batch normalization layer, an activation layer, and a maximum pooling layer connected in sequence.

[0023] The temporal feature extraction module includes a video frame temporal feature extraction operation and a plurality of fifth convolution layers performed in sequence. The video frame temporal feature extraction operation is to perform a difference process between each two adjacent frames of images, and then input the difference result into a respective fifth convolution layer for feature extraction to obtain differential motion features, and all differential motion features are spliced ​​and fused as the output of the temporal feature extraction module;

[0024] The fifth convolutional layer is mainly composed of convolution operations.

[0025] The number of the fifth convolutional layer is equal to the number of medium video frames. The first video frame of the medium video frame is used as the input of the first convolutional layer. Each subsequent frame is subtracted from the previous video frame and then input into the corresponding convolutional layer for feature extraction to obtain differential motion features. Finally, the differential motion features corresponding to all video frames are concatenated and fused as the output of the temporal feature extraction module.

[0026] The spatial feature enhancement module PSC includes two spatial feature branches, an addition operation, multiple sigmoid activation functions and a point-by-point multiplication operation. The spatial feature enhancement module PSC is respectively input into two spatial feature branches, and the two spatial feature branches are respectively a high receptive field branch and a low receptive field branch. The convolution size of the convolution layer in the high receptive field branch is larger than that of the low receptive field branch. The convolution layer in the high receptive field branch uses 3×3 convolution for feature processing, and the convolution layer in the low receptive field branch uses 1×1 convolution for feature processing.

[0027] The topological structure of each spatial feature branch is the same, including the sixth convolutional layer, the seventh convolutional layer, the global pooling layer, the first reconstruction layer, the second reconstruction layer and the softmax activation function. The input of the spatial feature enhancement module PSC is used as the input of the spatial feature branch and is respectively input into the sixth convolutional layer and the seventh convolutional layer. The output of the sixth convolutional layer is successively passed through the global pooling layer, the first reconstruction layer, and the softmax activation function to obtain AA features. The output of the seventh convolutional layer is passed through the second reconstruction layer to obtain BB features. AA features and BB features are used as the outputs of the spatial feature branches.

[0028] The AA features output by the high receptive field branch and the BB features output by the low receptive field branch are processed by matrix multiplication and input into the first sigmoid activation function. The AA features output by the low receptive field branch and the BB features output by the high receptive field branch are processed by matrix multiplication and input into the second sigmoid activation function. The outputs of the first sigmoid activation function and the second sigmoid activation function are respectively input into the third sigmoid activation function after adding their respective pixels. The output of the third sigmoid activation function and the input of the spatial feature enhancement module PSC are output as the output of the spatial feature enhancement module PSC after point-by-point multiplication.

[0029] The spatial feature enhancement module PSC obtains the spatial feature weights through global maximum pooling and matrix multiplication, and then adjusts the features of the high-scale branches with the weights generated by the low-scale receptive field branches, and adjusts the features of the low-scale branches with the weights generated by the high-scale receptive field branches. Finally, the weights are added as the final spatial feature weights and multiplied with the input features.

[0030] The action classification module specifically adopts a linear classification structure, including a global average pooling layer, a Dropout layer and a classifier connected in sequence. The recognition features containing spatial information, motion information and motion detail information are used as inputs of the global average pooling layer, and the classifier outputs the action recognition result.

[0031] The computer program is an instruction corresponding to the method described in any one of claims 1 to 9.

[0032] The method of the present invention mainly comprises the following steps: firstly, the video frame of the input video image is preprocessed, and three branches based on different frame intervals are extracted, wherein the high-frequency video frame is passed into the motion branch as a feature tensor containing motion information, the low-frequency video frame is passed into the appearance branch as a feature tensor containing background information, and the medium-frequency video frame is subjected to frame feature extraction to obtain action detail signs and pass them into the action detail branch; after each residual module of the three branches, the feature tensors are spliced ​​on the channel on the appearance branch to perform feature fusion and are processed by convolution; finally, the feature tensor results on the three branches that have been processed differently are input into the linear classification layer to obtain the final action recognition result.

[0033] The present invention can extract information from the temporal features of video images to adjust the spatial features of video images, obtain spatial features that strengthen the weights of key spatial parts, and thereby obtain more accurate video image action recognition.

[0034] Beneficial effects of the present invention:

[0035] The present invention proposes a PSC-TNet video action recognition method based on the fusion of spatial features and frame difference information. The present invention designs two new plug-and-play modules, namely a temporal feature extraction module TtS and a spatial feature enhancement module PSC.

[0036] The temporal feature extraction module TtS aims to obtain weight scores based on the differential features of a certain frame interval, and more effectively locate the motion difference information between two frames to enhance the positional features of the space where the character's actions are located, so as to achieve secondary enhancement of the spatial features based on the temporal features of the video.

[0037] The spatial feature enhancement module PSC is different from the traditional global pooling of the channel dimension, which leads to feature loss. PSC adopts a filtering processing method. By compressing the channel dimension, the dimension in the orthogonal direction, that is, the spatial dimension, is kept at a high resolution, and then the two are matrix-multiplied to obtain the spatial feature weight without global pooling of the channel. The PSC spatial feature enhancement module adopts a cross-multiplication strategy. The features are processed by convolution kernels of different scales, that is, the spatial feature weights obtained by the 1×1 convolution kernel are used to adjust the features of the 3×3 convolution kernel, and the spatial feature weights obtained by the 3×3 convolution kernel are used to adjust the features of the 1×1 convolution kernel to achieve cross-spatial information aggregation.

[0038] The PSC-T module proposed in the present invention is the fusion of the PSC spatial feature enhancement module and the TtS temporal feature extraction module. The spatial features extracted by the TtS feature extraction module are added to the new weights obtained by adding the weights after cross-multiplication in the PSC module to output new spatial feature weights. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 The network framework diagram of the method of the present invention.

[0040] Figure 2 This is a schematic diagram of the structure of the coordination of the temporal feature extraction module TtS and the spatial feature enhancement module PSC used in the present invention.

[0041] Figure 3 This is a schematic diagram of the structure of the spatial feature enhancement module PSC used in the present invention.

[0042] Figure 4 Schematic diagram of the structure of the residual module in the network of the present invention. DETAILED DESCRIPTION

[0043] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0044] The embodiments of the present invention and their implementation processes and situations are as follows:

[0045] The dataset of the present invention adopts the HMDB51 dataset. HMDB51 contains 51 types of actions, a total of 6849 videos, each action contains at least 51 videos, the resolution is 320*240, from YouTube, Google videos, etc., a total of 2G.

[0046] 1) After preprocessing the input video image, obtain low-frequency video frames, high-frequency video frames and medium-frequency video; specifically, normalize the size of the original input video image to obtain video image frames. In the video image frame sequence, take video frames at a larger certain step length to obtain a low-frequency video frame sequence as the input of the appearance branch; take video frames at a smaller certain step length to obtain a high-frequency video frame sequence as the input of the motion branch; take video frames at a medium step length to obtain a medium-frequency video frame sequence. The size of the original input video image is normalized to a 64×112×112 3-channel RGB image, and 64×112×112 is used as the input size of the neural network. Then the 3-channel RGB image is standardized and mapped from integers between 0 and 255 to floating-point numbers between 0 and 1.

[0047] 2) Input the low-frequency video frames into the appearance branch, the high-frequency video frames into the motion branch, and the medium-frequency video frames into the motion detail branch. The feature tensor containing spatial information output by each layer in the appearance branch and the feature tensor containing motion information output by the corresponding layer in the motion branch are concatenated, fused and convolved in the channel dimension to obtain the final feature tensor.

[0048] In this embodiment, the input sizes of the appearance branch and the motion branch are 4×112×112 and 32×112×112, respectively, so as to extract spatial information and motion information features.

[0049] Figure 1 The figure shows the overall framework of PSC-TNet. The motion branch (the right branch) includes the first convolution block and the first residual layer to the fourth residual layer connected in sequence. A spatial feature enhancement module (PSC) is added to the residual module of each residual layer. The appearance branch (the center branch) includes the second convolution block and the fifth residual layer to the eighth residual layer connected in sequence. A spatial feature enhancement module (PSC) is added to the residual module of each residual layer; the motion detail branch (the left branch) includes the temporal feature extraction module (TtS), the third convolution block and the ninth residual layer to the twelfth residual layer connected in sequence. A temporal feature extraction module (TtS) is added in front of each residual layer, and a spatiotemporal feature fusion module (PSC-T) is added to the residual module of the residual layer.

[0050] The temporal feature extraction module (TtS) is as follows Figure 2As shown, it includes video frame time feature extraction operations and several fifth convolution layers performed in sequence. The video frame time feature extraction operation is to perform difference processing between each two adjacent frame images, and then input the difference result into each fifth convolution layer for feature extraction to obtain differential motion features, and all differential motion features are spliced ​​and fused as the output of the time feature extraction module; the fifth convolution layers are all composed of convolution operations.

[0051] The number of the fifth convolutional layer is equal to the number of medium video frames. The first video frame of the medium video frame is used as the input of the first convolutional layer. Each subsequent frame is subtracted from the previous video frame and then input into the corresponding convolutional layer for feature extraction to obtain differential motion features. Finally, the differential motion features corresponding to all video frames are concatenated and fused as the output of the temporal feature extraction module.

[0052] The video frame temporal feature extraction operation consists of a convolution layer with equal input channels and output channels to obtain the weighted scores. The specific formula is as follows:

[0053]

[0054] Score t =sigmoid(f t )

[0055] v=Concat(Score0,Score1,……,Score T )

[0056] Among them, f t Score is the motion information feature tensor obtained by subtracting frame T from frame T-1. t To obtain the weight score in the calculated space, Conv() is the spatial convolution function, Concat() is the stacking function, and x t is the tth input feature tensor, and v is the output feature tensor, i.e., the motion detail feature.

[0057] The spatial feature enhancement module (PSC) is as follows Figure 3 As shown in the figure, it is divided into two branches of different scales. The left branch is a 3×3 convolution, and the right branch is a 1×1 convolution. They are processed by global average maximum pooling and Softmax function respectively. The specific formulas are as follows:

[0058] x l1 =Softmax reshape (GlobalPooling(Conv 3×3 (x)))

[0059] x l2 = reshape(Conv 3×3(x))

[0060] x r1 =Softmax reshape (GlobalPooling(Conv 1×1 (x)))

[0061] x r2 = reshape(Conv 1×1 (x))

[0062] Among them, x is the initial input feature, x l1 and x l2 is the feature branch with high receptive field, x r1 and x r2 is a feature branch with low receptive field. In order to make up for the deficiency of single scale and perform multi-scale spatial modeling, the next step is to aggregate cross-space information, that is, to use x r2 x l1 To adjust the weight, use x l2 x r1 To adjust the weight, the specific formula is as follows:

[0063] x l =x l1 × r2

[0064] x r =x r1 × l2

[0065] x out =x l +x r

[0066] x=x×x out

[0067] Among them, x l is the spatial weight of the low receptive field branch, x r is the spatial weight of the high receptive field branch, which is added together to obtain the final x out The spatial weight is multiplied by the initial x feature to obtain the final output of the spatial feature enhancement module.

[0068] A spatiotemporal feature fusion module (PSC-T) is added to each residual module of the action detail branch, which is a combination of the spatial feature enhancement module (PSC) and the temporal feature extraction module (TtS). The spatiotemporal feature weights extracted by the temporal feature extraction module (TtS) are added to the final weight addition stage of the spatial feature module (PSC). The specific formula is as follows:

[0069] x weight =xl +x r +v

[0070] x=x×x weight

[0071] Among them, x weight This is the final feature weight map of the spatiotemporal feature fusion module, which is multiplied by the initial feature x to obtain the output of the spatiotemporal feature fusion module.

[0072] In the motion detail branch, a spatiotemporal feature fusion module (PSC-T) is added to each residual module to enhance the performance of the staggered block and its sensitivity to space. The specific formula is as follows:

[0073] x1 = Relu(BN(Conv(x)))

[0074] x2=Relu(BN(Conv(x1)))

[0075] x3=BN(Conv(x2))

[0076]

[0077] Among them, the initial feature x passes through the first convolution layer, the first normalization layer, the first activation layer, the second convolution layer, the second normalization layer, the second activation layer, the third convolution layer, the third normalization layer, and is input into the spatiotemporal feature fusion module (PSC-T) or the spatial feature enhancement module (PSC), that is, middle is the motion detail branch, and the spatiotemporal feature fusion module (PSC-T) is inserted into the residual module of the motion detail branch. Fast and slow are the motion branch and appearance branch, respectively, in which the spatial feature enhancement module (PSC) is inserted.

[0078] Each residual layer includes multiple residual modules connected in sequence, and the number of residual modules contained in the four residual layers is 3, 4, 6, and 3 respectively. Among them, the first residual module in each residual layer is downsampled, the purpose of which is to reduce the dimension of the feature. Each residual module includes a residual map and an identity map, wherein the input of each residual module is recorded as an input feature tensor, and the input feature tensor is extracted by the residual map to obtain the first space-motion feature tensor, and the input feature tensor is obtained by the identity map to obtain the second space-motion feature tensor. The output feature tensor of the residual module is obtained by adding the first space-motion feature tensor and the second space-motion feature tensor. The specific formula is as follows:

[0079] H(x)=F(x)+G(x)

[0080] Among them, H() is the output function of the residual module; F() is the residual mapping function; G() is the identity mapping function; x is the input feature tensor of the residual module.

[0081] The network structure diagram of the first residual module in each residual layer is as follows Figure 4 As shown, in the identity mapping, it is determined whether the number of channels of the output feature tensor (i.e., the first feature tensor) of the residual mapping is the same as the number of channels of the input feature tensor; if they are the same, the input feature tensor is directly used as the output feature tensor of the identity mapping; if they are not the same, the feature tensor after point-by-point convolution of the input feature tensor is used as the output feature tensor of the identity mapping, i.e., the second feature tensor. It can be set by the following formula:

[0082]

[0083] Where G() is the identity mapping function; is a convolution function with a kernel size of 1×1×1 and an output channel of C.

[0084] The feature fusion strategy is to concatenate and fuse the output of the first convolution block and the output of the third convolution block on the channel dimension of the output of the second convolution block to obtain the first space-motion feature tensor. The specific formula is as follows:

[0085] T(x)=Concat(DownSample fast (x fast ),DownSample middle (x middle ),x slow )

[0086] Among them, T(x) is the output function of the stacking result; Concat is the connection function, which performs splicing and fusion in the channel dimension; DownSample fast is the upsampling function on the motion branch, DownSample middle For the upsampling function on the motion detail branch, trilinear interpolation is used to maintain the stackable dimension size.

[0087] By analogy, the above is the first space-motion feature tensor. The output of the first residual layer and the output of the ninth residual layer are spliced ​​and fused on the channel dimension of the fifth residual layer output to obtain the second space-motion feature tensor. The second space-motion feature tensor, the output of the second residual layer and the output of the sixth residual layer are spliced, fused and convolved on the channel dimension to obtain the third space-motion feature tensor. The third space-motion feature tensor, the output of the third residual layer and the output of the seventh residual layer are spliced, fused and convolved on the channel dimension to obtain the fourth space-motion feature tensor. The fourth space-motion feature tensor, the output of the fourth residual layer and the output of the eighth residual layer are spliced, fused and convolved on the channel dimension to obtain the final feature tensor.

[0088] After the feature extraction of the three branches, the output results of the three branches are added and fused to obtain the output feature tensor containing texture information, time information and spatial information, which is processed by the following formula:

[0089] B(X)=B fast (x)+B middle (x)+B slow (x)

[0090] Among them, B() is the fusion output function; B fast (x) is the final output of the motion branch; B slow (x) is the final output of the appearance branch; B middle (x) is the final output of the motion detail branch; x is the input feature tensor.

[0091] 3) The final feature tensor is input into the action classification module to obtain the action recognition result. Among them, the action classification module includes a global average pooling layer, a Dropout layer and a classifier connected in sequence, and the feature tensor containing spatial information and motion information is used as the input of the global average pooling layer, and the classifier outputs the action recognition result. In this embodiment, the probability value mapped to 51 categories in the classifier is taken as the action category of the video image, and processed by the following formula:

[0092] O(x)=Max(Linear(Dropout(GlobalAvgPool(x))))

[0093] Among them, O() is the network output function; Linear() is the fully connected function; Dropout() is the regularization function, which randomly drops neurons with a given probability; Max() is the maximum value function, GlobalAvgPool() is the global average pooling function, and x is the input feature tensor.

[0094] The present invention also provides a storage medium storing a computer program, which, when executed by a processor, implements a PSC-TNet video action recognition method based on fusion of spatial features and frame differential information. The computer program is an instruction corresponding to the implementation of a PSC-TNet video action recognition method based on fusion of spatial features and frame differential information.

[0095] Finally, it should be noted that the above embodiments and explanations are only used to illustrate the technical solution of the present invention rather than to limit it. Those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope disclosed in the technical solution of the present invention, which should be included in the scope of protection of the claims of the present invention.

[0096] In summary, a new spatial feature enhancement module PSC and temporal feature extraction module TtS were proposed, and a three-branch structural innovation was made on the skeleton network slowfast. A new action recognition network structure was designed, which achieved certain improvements on HMDB51 compared with the skeleton network slowfast.

Claims

1. A PSC-TNet video action recognition method based on fusion of spatial features and frame difference information, characterized in that: The method comprises the following steps: preprocessing image frames of an input video, inputting the preprocessed video as the original video into a PSC-TNet video action recognition model based on fusion spatial features and frame difference information, and extracting and obtaining action recognition results; the PSC-TNet video action recognition model comprises three branches based on different frame intervals and an action classification module, the three branches are respectively a motion branch, an appearance branch and an action detail branch, the original video is respectively input into the motion branch, the appearance branch and the action detail branch to output respective action features, and then all the recognition features are spliced ​​and input into the action classification module to obtain the action classification result.

2. The PSC-TNet video action recognition method based on fusion of spatial features and frame difference information according to claim 1 is characterized in that: The input video is sampled at a higher frequency, a lower frequency, and a medium frequency to obtain a higher frequency video frame, a lower frequency video frame, and a medium frequency video frame to form an original video.

3. The PSC-TNet video action recognition method based on fusion of spatial features and frame difference information according to claim 1 is characterized in that: In the PSC-TNet video action recognition model, higher-frequency video frames in the original video are used as feature tensors containing motion information and are passed into the motion branch, lower-frequency video frames in the original video are used as feature tensors containing spatial information and are passed into the appearance branch, and medium-frequency video frames in the original video are used as feature tensors containing action detail information and are passed into the action detail branch; the three branches all contain residual layers, and the features of the intermediate outputs of the motion branch and the action detail branch are passed to the appearance branch for channel splicing and feature fusion, and then input into the residual layer of the appearance branch, and finally the recognition features output by the three branches are spliced ​​and fused on each channel and input into the action classification module to obtain the final action recognition result.

4. The PSC-TNet video action recognition method based on fusion of spatial features and frame difference information according to claim 1 is characterized in that: In the PSC-TNet video action recognition model, The motion branch and the appearance branch each include a sequentially connected convolution block and four consecutive residual layers including a spatial feature enhancement module PSC; the action detail branch includes a sequentially connected temporal feature extraction module TtS, a convolution block and four consecutive spatiotemporal feature fusion modules, each spatiotemporal feature fusion module includes a parallel temporal feature extraction module TtS and a residual layer including a spatial feature enhancement module PSC, the input of the spatiotemporal feature fusion module is respectively input into the temporal feature extraction module TtS and the residual layer to obtain respective results, and the respective results are added and fused as the output of the spatiotemporal feature fusion module; The four residual layers of the motion branch, appearance branch and action detail branch respectively contain three, four, six and three residual modules in the transmission order; the convolution blocks of the motion branch and the action detail branch and the convolution blocks of the appearance branch are spliced ​​on the channel of the appearance branch and input into the first residual layer of the appearance branch, and the nth residual layer of the motion branch and the action detail branch and the nth residual layer of the appearance branch are spliced ​​on the channel of the appearance branch and input into the n+1th residual layer of the appearance branch.

5. The PSC-TNet video action recognition method based on fusion of spatial features and frame difference information according to claim 4 is characterized in that: The residual module includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a PSC-T attention mechanism module and a fourth convolutional layer. The first convolutional layer, the second convolutional layer, the third convolutional layer and the PSC-T attention mechanism module are connected in sequence. The input of the residual module is respectively input into the first convolutional layer and the fourth convolutional layer. The output of the PSC-T attention mechanism module and the output of the fourth convolutional layer are added as the output of the residual module; the first convolutional layer and the second convolutional layer are mainly composed of a convolution operation, a batch normalization operation, and an activation operation connected in sequence. The third convolutional layer is mainly composed of a convolution operation and a batch normalization operation connected in sequence, and the fourth convolutional layer is composed of one convolution operation.

6. The PSC-TNet video action recognition method based on fusion of spatial features and frame difference information according to claim 5 is characterized in that: The PSC-T attention mechanism module of the residual module in the residual layer of the motion branch, appearance branch and action detail branch adopts the spatial feature enhancement module PSC.

7. The PSC-TNet video action recognition method based on fusion of spatial features and frame difference information according to claim 4 is characterized in that: The temporal feature extraction module includes a video frame temporal feature extraction operation and several fifth convolutional layers performed in sequence. The video frame temporal feature extraction operation is to perform a difference process between each two adjacent frames of images, and then input the difference result into each fifth convolutional layer for feature extraction to obtain differential motion features, and all differential motion features are spliced ​​and fused as the output of the temporal feature extraction module; the fifth convolutional layers are mainly composed of convolution operations.

8. The PSC-TNet video action recognition method based on fusion of spatial features and frame difference information according to claim 4 is characterized in that: The spatial feature enhancement module PSC includes two spatial feature branches, an addition operation, multiple sigmoid activation functions and a point-by-point multiplication operation. The spatial feature enhancement module PSC is respectively input into the two spatial feature branches, which are respectively a high receptive field branch and a low receptive field branch. The convolution size of the convolution layer in the high receptive field branch is larger than that of the low receptive field branch. Each spatial feature branch has the same topological structure, including a sixth convolution layer, a seventh convolution layer, a global pooling layer, a first reconstruction layer, a second reconstruction layer and a softmax activation function. The input of the spatial feature enhancement module PSC is used as the input of the spatial feature branch and is respectively input into the sixth convolution layer and the seventh convolution layer. The output of the sixth convolution layer is sequentially subjected to the global pooling layer, the first reconstruction layer and the softmax activation function to obtain AA features. The output of the seventh convolution layer is subjected to the second reconstruction layer to obtain BB features. The AA features and the BB features are used as the outputs of the spatial feature branches. The AA features output by the high receptive field branch and the BB features output by the low receptive field branch are processed by matrix multiplication and input into the first sigmoid activation function. The AA features output by the low receptive field branch and the BB features output by the high receptive field branch are processed by matrix multiplication and input into the second sigmoid activation function. The outputs of the first sigmoid activation function and the second sigmoid activation function are respectively input into the third sigmoid activation function after adding their respective pixels. The output of the third sigmoid activation function and the input of the spatial feature enhancement module PSC are output as the output of the spatial feature enhancement module PSC after point-by-point multiplication.

9. The PSC-TNet video action recognition method based on fusion of spatial features and frame difference information according to claim 1, characterized in that: The action classification module includes a global average pooling layer, a Dropout layer and a classifier connected in sequence. The recognition features containing spatial information, motion information and motion detail information are used as inputs of the global average pooling layer, and the classifier outputs the action recognition result.

10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Human body action recognition method based on fine space network

    CN110378194A

  • Human body behavior recognition method based on time-space and operation information fusion

    CN114220170A

  • Lightweight behavior recognition method and system based on feature compression

    CN116110121A

  • Methods, apparatus, servers, and systems for vital signs detection and monitoring

    US20190007256A1

Cited By

  • Video generation method and device, electronic equipment, storage medium and program product

    CN120455806A