A Human Behavior Recognition Method for Multi-Path Network Based on Key Frame Selection
Through the multi-path network method based on keyframe selection, the problems of spatiotemporal information redundancy and keyframe screening difficulties in video behavior recognition are solved, multimodal feature fusion is realized, and the recognition accuracy of unclipped videos is improved.
Patent Information
- Application Number
- CN202410818306.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-06-24
AI Technical Summary
The prior art has problems such as spatiotemporal information redundancy, local texture information missing and keyframe screening in video behavior recognition, especially in the recognition effect of unedited long videos, and most methods only use single modal data.
Using a multi-path network method based on keyframe selection, video segments related to human body movements are screened out through a lightweight keyframe sampling module, and combined with a multi-path video-text encoder, the time and space encoder are used to learn spatiotemporal features, and feature similarity calculation is performed in combination with a text encoder to realize multi-modal feature fusion.
Effectively filter out keyframes, reduce irrelevant behavior interference, and improve recognition effect, especially to achieve high-precision recognition on unedited long video datasets, improving overall recognition accuracy.
Smart Images

Figure CN118968609B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human behavior recognition, and in particular to a human behavior recognition method based on a multi-path network with key frame selection. Background Art
[0002] Human behavior recognition refers to the process of identifying human behaviors by a computer through processing and analyzing videos. This task is an important research topic in the field of computer vision and has received extensive attention in both academia and industry. Its research results have shown broad application prospects in fields such as autonomous driving, motion analysis, and smart homes. At the same time, as a fundamental task in the field of video understanding, the research results based on video behavior recognition have a direct promoting effect on the development of a series of downstream tasks such as video detection, video segmentation, and video localization.
[0003] The rise of deep learning has promoted the development of the field of video behavior recognition, but also brought new challenges. Compared with image data, video data provides richer background appearance information and context temporal information. There is spatio-temporal correlation between consecutive frames in a video, enabling the model to understand the persistence and coherence of behaviors. However, there is a large amount of spatio-temporal information redundancy in videos, especially in unclipped long videos, and only some video segments are related to human behaviors, which poses a huge challenge to the training and actual deployment of the model. The method based on convolutional neural network has achieved remarkable success in the field of behavior recognition, but there are problems such as insufficient spatio-temporal information modeling and difficulty in capturing global information. In order to perform long-distance spatio-temporal modeling, many researchers have introduced Transformer into the field of behavior recognition and performed human action recognition by fine-tuning a pre-trained image model. This end-to-end fine-tuning can effectively perform global spatio-temporal modeling, but also brings a large computational burden. In addition, existing visual Transformer models usually manually segment image frames into image sequences first, which will destroy the inherent structure of the image frames and cause the model to lack the ability to fully model local information.
[0004] (1) The problem of spatio-temporal information redundancy in videos. Consecutive frames in a video are highly similar and contain a large amount of redundant information. The model only needs to identify a small number of frames and partial positions to obtain the correct classification result. The additional spatio-temporal information redundancy will increase the computational cost and pose a huge challenge to the practical application of behavior recognition.
[0005] (2) The problem of missing local texture information caused by manual segmentation of image frames. The method based on the Transformer model usually performs manual segmentation and position encoding on the input image frames and divides the complete picture into a Tokens sequence. This will destroy the complete structure of the image and make it difficult for the model to effectively model the local edge texture information of the image frames.
[0006] (3) Problem that the existing models have difficulty in screening key frames in videos. Only a small number of key frames in the video are related to human behaviors. The existing sampling strategies cannot screen out key frames well, which is particularly obvious in the unclipped long video dataset. The lengths of unclipped videos vary, and the irrelevant image frames contain interfering targets and scenes, affecting the overall recognition effect.
[0007] Currently, the behavior recognition methods based on RGB data in the existing technologies can be summarized into three categories: methods based on 2D convolutional neural networks, methods based on 3D convolutional neural networks, and methods based on Transformer models.
[0008] An EVL (Efficient Video Learners) method based on CLIP (Contrastive language-image pretraining) in the existing technologies includes: extracting initial features in the video through the CLIP pre-trained model, embedding a local temporal module in the lightweight decoder, and jointly learning spatio-temporal features in the video. The processing steps of its main algorithm flow include:
[0009] The pre-trained model extracts video features. First, use the pre-trained model to extract video features, and input the features extracted by the backbone network into the decoder layer to further learn spatio-temporal information.
[0010] Temporal convolution captures temporal information. The pre-trained model extracts powerful spatial features but lacks temporal information. By performing one-dimensional convolution operations in the temporal dimension, the temporal information in the video is supplemented.
[0011] The cross-frame attention mechanism enhances temporal information. By performing the attention mechanism between adjacent frames, the motion change information in the video can be effectively captured, enhancing the representation of temporal information.
[0012] Capturing global spatio-temporal information through the multi-head attention mechanism. Finally, input the temporally enhanced features into the multi-head attention module, and the global spatio-temporal information can be effectively captured through the attention mechanism.
[0013] The disadvantages of the above-mentioned CLIP-based EVL method in the existing technologies include:
[0014] The methods based on the Transformer model usually perform manual segmentation and position encoding on the input image frames, dividing the complete picture into a Tokens sequence. This will destroy the complete structure of the image, making it difficult for the model to effectively model the local edge texture information of the image frames.
[0015] Only a few key segments in the video contain human behavior information. A large amount of irrelevant redundant information will interfere with action recognition and affect the recognition accuracy. This network performs well on the clipped video segment dataset, but has a poor recognition effect on unclipped long videos.
[0016] This network can fully capture the spatio-temporal information in the video. However, only RGB single-modal data is used, and no other modal data is used for supplementation (such as sound modality and text modality), so there is still room for further improvement in recognition performance.
[0017] A video behavior recognition method based on adaptive token sampling in the prior art includes: This method designs a parameter-free adaptive Token sampler module, which scores and filters important tokens, thereby endowing the model with the ability of adaptive sampling.
[0018] The processing steps of this method include:
[0019] Tokens scoring. This network integrates an adaptive token selection module, divides the input video into tokens at each stage of the network, calculates the correlation between these tokens through the attention mechanism, and generates an attention matrix, thereby obtaining the importance of each token.
[0020] Tokens sampling. This step clips the input sequence according to the score of each token, thereby retaining important information and clipping off irrelevant tokens. Specifically, after obtaining the score of each token, the cumulative probability density function of the input tokens can be calculated, and the sampling function is generated by taking the inverse of the probability density function. Sampling the input token sequence through the sampling function can reduce redundant information and retain key tokens.
[0021] Information extraction by the attention mechanism. Inputting the sampled key information into the attention mechanism layer can effectively extract the information in the video and obtain the final classification result.
[0022] The disadvantages of a video behavior recognition method based on adaptive token sampling in the above prior art include:
[0023] This network is directly migrated from the image field to the video field. Although it can effectively reduce information redundancy and screen out key information, it does not consider the motion change information in the video and does not fully model the spatio-temporal information of the video. Therefore, the recognition effect on the video is poor.
[0024] Only single-modal data is used, and no other modal data is used for supplementation (such as sound modality and text modality), so there is still room for further improvement in recognition performance.
[0025] The key frames in the video are not pre-screened, which is more effective for segmented clips of the video but less effective for long unedited videos. Summary of the Invention
[0026] Embodiments of the present invention provide a human behavior recognition method based on a multi-path network with key frame selection to effectively recognize human behaviors in video data.
[0027] To achieve the above object, the present invention adopts the following technical solutions.
[0028] A human behavior recognition method based on a multi-path network with key frame selection includes:
[0029] Sampling the video data to be recognized to obtain multiple video segments;
[0030] Collecting features for each video segment, using a multi-layer perceptron and a normalization function to generate a probability distribution, and screening out the video segments where human actions are located according to the probability distribution;
[0031] Inputting the video segments where the human actions are located into a multi-path video-text encoder classification network, learning spatio-temporal features from the video segments through a time encoder and a space encoder, learning text features in the video segments through a text encoder, and obtaining the recognition result of the human behavior of the video to be recognized by calculating the similarity between the spatio-temporal features and the text features.
[0032] Preferably, the sampling of the video data to be recognized to obtain multiple video segments includes:
[0033] Using a lightweight key frame sampling module to decode the video data to be recognized into an image frame sequence using FFmpeg, and performing sparse sampling on the image frame sequence using the random sampling strategy of the Time-Sensitive Network (TSN). The number of sampled frames is T frames, and the T image frames are segmented to obtain N video segments: V1, V2, V3... V N , and each video segment contains multiple image frames.
[0034] Preferably, the collecting features for each video segment, using a multi-layer perceptron and a normalization function to generate a probability distribution, and screening out the video segments where human actions are located according to the probability distribution includes:
[0035] The key frame sampling module uses a pre-trained ResNet-18 to extract features from N video segments: V1, V2, V3... V N The features of each video segment are concatenated, and the probability distribution of each frame in each video segment is generated through a multi-layer perceptron and a normalization function. The probability distribution corresponding to each video segment is The sum of the probability distributions corresponding to all video segments is 1, that is,
[0036] For the probability distribution of the j-th frame in the i-th video segment V i The calculation method is as follows: The calculation method is:
[0037]
[0038] R represents the pre-trained ResNet-18 network, MLP represents a multi-layer perceptron, and softmax represents a normalization function; assume that each video segment V i Contains D frame images, and the probability distributions corresponding to each frame image are
[0039]
[0040] According to the probability distribution corresponding to the video segment, through the key video segment selection strategy, the video segments F1, F2, F3... F where the human action is located are screened out Z , the key video segment selection strategy includes:
[0041] (a) Select the two video segments where the maximum and the second maximum values of the probability distribution of the frame image are located. The two video segments where the maximum and the second maximum values of the probability distribution of the frame image are located.
[0042] (b) Select the two video segments with the largest sum of the probability distributions of all frame images The two video segments with the largest sum of the probability distributions of all frame images.
[0043] (c) Select the two video segments with the largest number of frame images whose probability distributions Exceed the probability mean The two video segments with the largest number of frame images whose probability distributions exceed the probability mean.
[0044] Preferably, inputting the video segments where the human action is located into the multi-path video-text encoder classification network, learning spatio-temporal features from the video segments through the time encoder and the space encoder, learning text features in the video segments through the text encoder, and obtaining the recognition result of the human behavior of the video to be recognized by calculating the similarity between the spatio-temporal features and the text features, including:
[0045] The multi-path video-text encoder network includes a video encoder branch and a text encoder branch. The video encoder branch includes a spatial stream structure and a temporal stream structure. The spatial stream structure collects image information with different resolutions, and the temporal stream structure collects image information with different frame rates.
[0046] The spatial stream structure samples key video segments F1, F2, F3... F output from the key frame sampling module ZRandomly sample T frame images from each video segment and perform random cropping to obtain high-resolution images with the input video dimension of T×C×H×W;
[0047] S in = Crop(Random(F1, F2, F3…F Z )) (2―3)
[0048] where Random represents randomly selecting T frames, Crop represents performing random cropping operations on the image frames, and S in represents the high-resolution video input to the spatial encoder;
[0049] The temporal stream structure selects all key video segments F1, F2, F3…F Z , performs random cropping and downsampling on each video segment, and obtains the input video dimension of Input the high-frame-rate frame sequence into the network;
[0050] T in = DownSample(Crop(F1, F2, F3…F Z )) (2―4)
[0051] where Crop represents performing random cropping operations on the image frames, DownSample represents performing downsampling operations on the video, and T in represents the high-frame-rate video finally input to the temporal encoder;
[0052] The spatial encoder uses the backbone network to extract features from the input high-resolution video to obtain spatial information. The temporal encoder uses the backbone network to extract features from the input high-frame-rate video to obtain temporal features, align and fuse the temporal features with the spatial information to obtain the spatio-temporal features of the entire video;
[0053] F v = S out + Reshape(T out ) (2―5)
[0054] where Reshape represents performing dimension reorganization operations on the features, adjusting the feature map size to T×N×D, and F v represents the spatio-temporal features of the entire video finally;
[0055] The text encoder branch uses a Transformer with a network depth of 12 layers and eight attention heads to extract features and obtain the text features of the entire video;
[0056] Calculate the similarity between the spatio-temporal features and the text features of each video, and determine the classification result y of the human behavior of the video according to the similarity* ;
[0057] y * = argmax(P(f(x, y)|θ) (2 - 6)
[0058] where x represents the spatio-temporal features of the input video, y represents the text information of the corresponding label of the video, f represents the function for calculating similarity, and the label with the highest similarity to the video is the classification result y of the human behavior in the video * 。
[0059] As can be seen from the technical solutions provided by the embodiments of the present invention above, the method proposed by the present invention can effectively screen out the image frames related to human actions, reduce the interference of irrelevant behaviors, and improve the overall recognition effect.
[0060] Additional aspects and advantages of the present invention will be given in part in the following description, and these will become apparent from the following description, or can be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0062] Figure 1 It is a schematic diagram of the implementation of a human behavior recognition method based on key frame selection and multi-path network provided by an embodiment of the present invention;
[0063] Figure 2 It is a processing flow chart of a human behavior recognition method based on key frame selection and multi-path network provided by an embodiment of the present invention;
[0064] Figure 3 It is a structural diagram of a multi-path video-text encoder network provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] The following will describe in detail the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention, and cannot be construed as a limitation of the present invention.
[0066] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the stated features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The phrase "and / or" used herein includes any unit and all combinations of one or more of the associated listed items.
[0067] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as here.
[0068] For the convenience of understanding the embodiments of the present invention, the following will further explain with several specific embodiments in conjunction with the accompanying drawings, and each embodiment does not constitute a limitation on the embodiments of the present invention.
[0069] The embodiment of the present invention provides a MAR-Net (Multi-path Action Recognition Network based on Key Frame Selection) based on key frame selection. This network mainly consists of two parts: a key frame sampling module (Key Frame Sampling Module, KFS Module) and a multi-path video-text encoder. The key frame sampling module is constructed based on a lightweight CNN, which can pre-identify the video segments where human behaviors are located at a relatively low computational cost and perform frame sampling on the key video segments. Finally, the key frames are input into the multi-path video-text encoder network to extract image and text information and obtain the classification result of the video.
[0070] The implementation principle diagram of a human behavior recognition method based on a multi-path network with key frame selection provided by the embodiment of the present invention is as Figure 1 shown, including a key frame sampling module and a multi-path video-text encoder network for classification. The processing flow chart of this method is as Figure 2 shown, including the following processing steps;
[0071] Step S21: The key-frame sampling module samples the video data to be recognized, obtaining multiple video segments.
[0072] Obtain the video data to be recognized. After the lightweight key-frame sampling module decodes the video data to be recognized, it uses the random sampling strategy of TSN (Time-Sensitive Network) to sample the input video. The T frames obtained by sampling are pre-divided into N video segment sequences: V1, V2, V3…V N , and each segment represents a video segment clip.
[0073] Step S22: The key-frame sampling module collects features for each video segment, uses a multi-layer perceptron and a normalization function (softmax) to generate a probability distribution, and filters out the video segment where the human action is located according to the probability distribution.
[0074] Step S23: Input the video segment where the human action is located into the multi-path video-text encoder classification network. Learn spatio-temporal features from the video segment through the time encoder and the space encoder, learn text features in the video segment through the text encoder, and obtain the recognition result of the human behavior of the video to be recognized by calculating the similarity between the spatio-temporal features and the text features.
[0075] Specifically, the above steps S21 and S22 include that the main function of the key-frame sampling module (KFS Module) is to filter out the key video segment where the human action is located from the input video segment sequence. Its main body is a lightweight convolutional neural network ResNet-18, and the additional computational cost introduced is small. The key-frame sampling module mainly consists of the following parts:
[0076] 1. FFmpeg decoding, initially sampling image frames.
[0077] 2. Video segmentation.
[0078] 3. Extract features and generate a probability distribution.
[0079] 4. Secondarily sample image frames, key-frame screening.
[0080] The detailed process of the key-frame sampling module is as follows:
[0081] 1. Initially sample image frames
[0082] This part first uses FFmpeg to decode the video to be recognized to obtain a sequence of image frames, and then uses the sampling strategy of TSN to sparsely sample the input image sequence, with the number of sampled frames being T frames. Taking T = 64 frames as an example, the dimension of the video initially input to the sampling module is T×C×H×W = 64×3×224×224, where C represents the number of channels, and H and W represent the height and width of the image frame.
[0083] 2. Video segmentation
[0084] This part mainly manually segments the 64 image frames sampled initially, with the number of segments being N. Taking N = 4 as an example, each video segment after segmentation contains 16 image frames, and the video dimensions corresponding to V1, V2, V3, and V4 are
[0085] T′×C×H×W = 16×3×224×224.
[0086] 3. Extract features and generate probability distribution
[0087] First, use the pre-trained ResNet-18 to extract features for each frame in the four video segments V1, V2, V3, and V4, and then splice the features and generate the probability distribution of each frame in each video segment through a multi-layer perceptron and a normalization function. The probability distribution corresponding to each video segment is The sum of the probability distributions corresponding to all video segments is 1, that is
[0088] The probability distribution of the jth frame in the ith video segment V i is calculated as follows: The calculation method is:
[0089]
[0090] R represents the pre-trained ResNet-18 network, MLP represents a multi-layer perceptron, and softmax represents the normalization function; assume that each video segment V i contains D image frames, and the probability distributions corresponding to each frame image are
[0091]
[0092] 4. Secondary sampling of image frames
[0093] After obtaining the probability distributions of each video segment V1, V2, V3, and V4, determine the video segment where the human behavior is located by designing a segment selection strategy. In order to select two key video segments from the four video segments, the present invention designs three evaluation indicators for screening key video segments, and the comparison of the three selection strategies is as follows:
[0094] (a) Probability distribution of selected frame images Two video segments where the maximum value and the second maximum value are located.
[0095] (b) Sum of probability distributions of all selected frame images Two largest video segments.
[0096] (c) Probability distribution of selected frame images Exceeding the probability mean Two video segments with the largest number.
[0097] For a video to be recognized, two key video segments are selected, and each key video segment contains 16 frame images. Then, eight frame images are screened from each key video segment, and different sampling strategies are used for the training set and the test set. During training, to increase the diversity of data, eight frames are randomly screened as key frames in each segment; during testing, the probability distribution P j i The eight largest video frames. A total of 16 frames F1, F2, F3…F are screened from the two segments 16 .
[0098] It is proved by experiments that the screening effect of strategy b is the most effective.
[0099] Specifically, the above step S23 includes: The structure of a multi-path video-text encoder network provided by an embodiment of the present invention is as Figure 3 shown. Specifically, the multi-path video-text encoder network is composed of a video encoder branch and a text encoder branch, which are respectively used to learn the image information of image frames and the text information of classification labels, match the correct action label for the input video through similarity measurement, and output the final human action recognition result.
[0100] In order to perform sufficient spatio-temporal modeling on the Z key video segments F1, F2, F3…F output by the key frame sampling module Z Here, Z = 16 is taken as an example for illustration.
[0101] The present invention designs the video encoder branch as a spatio-temporal two-stream structure, and respectively collects image information with different frame rates and different resolutions.
[0102] The spatial stream sacrifices the number of input image frames and uses high-resolution images to enhance the extraction of background texture information. The specific principle is shown in formula 2-3. The spatial stream starts from the key video segments F1, F2, F3…F 16Randomly sample four frames from each video segment and perform random cropping. The input video dimension obtained is T×C×H×W = 4×3×224×224. Then, input the high-resolution images into the backbone network, which can effectively capture the spatial information in the video.
[0103] S in = Crop(Random(F1, F2, F3…F 16 )) (2―3)
[0104] where F1, F2, F3…F 16 represent the sequence of key video segments selected by the key frame module, Random represents randomly selecting four frames, Crop represents performing a random cropping operation on the image frames, and S in represents the high-resolution video input to the spatial encoder.
[0105] The temporal stream sacrifices the resolution of the input video and uses a larger number of input frames to enhance the model's ability to capture information on continuous motion changes. The specific principle is shown in Equation 2-4. The temporal stream selects all the key frames F1, F2, F3…F 16 , performs random cropping and downsampling on the image frames. The input video dimension obtained is T×C×H×W = 16×3×112×112. Then, input the high-frame-rate frame sequence into the network, which can effectively capture the motion change information in the video.
[0106] T in = DownSample(Crop(F1, F2, F3…F 16 )) (2―4)
[0107] where Crop represents performing a random cropping operation on the image frames, DownSample represents performing a downsampling operation on the video to reduce the video resolution. T in represents the high-frame-rate video finally input to the temporal encoder.
[0108] The temporal encoder and the spatial encoder use the same backbone network for feature extraction. S out The dimension of the output feature map is T×N×D. T out The size of the output feature map is where T represents the number of input frames and N represents the number of Tokens. By performing a Reshape operation on the high-frame-rate feature T out align and fuse the temporal features with the spatial information to obtain the spatio-temporal features of the entire video, so as to enhance the representation of spatio-temporal information in the video.
[0109] F v = S out + Reshape(Tout ) (2 - 5)
[0110] Among them, Reshape represents performing a dimensional reorganization operation on the features, adjusting the feature map size to T×N×D. F v represents the spatio-temporal features of the entire video finally.
[0111] The text encoder branch uses a Transformer with a network depth of 12 layers and eight attention heads to extract features and obtain the video text features F in the label. t .
[0112] Then, the action recognition task is modeled as a video-text multi-modal learning problem. By calculating the similarity to match the text label for the input video, the final classification result is obtained. The principle is shown in Equation 2 - 6:
[0113] y * = argmax(P(f(F v , F t )│θ) (2 - 6)
[0114] Among them, F v represents the spatio-temporal features of the input video, F t represents the text information of the video corresponding label, f represents the function for calculating the similarity, and the label with the highest similarity to the video is the classification result y of the human action in the video. * .
[0115] The method of the present invention can be applied to many real-world scenarios such as autonomous driving, intelligent monitoring, and motion analysis, and involves fields such as industry, culture, and commerce. For example, predicting the actions of pedestrians during autonomous driving of a car; performing intelligent monitoring on parking lots and factory assembly lines; and performing real-time analysis on a sports competition during a sports event.
[0116] In summary, the embodiments of the present invention alleviate the following three problems of the existing Transformer-based action recognition algorithm:
[0117] The problem of insufficient spatio-temporal information modeling. The present invention designs a video encoder with a two-stream structure to collect image information with different frame rates and different resolutions respectively. The high frame rate path can effectively capture the motion information in the video, and the high resolution path can effectively capture the background information in the video. By fusing the spatio-temporal information, the spatio-temporal information representation can be effectively enhanced.
[0118] The problem of being unable to screen key frames in videos. Only a few key segments in the video contain human behavior information, and a large amount of irrelevant redundant information will interfere with action recognition and affect the overall recognition effect. Through designing a key frame sampling module, the present invention can pre-identify the video segments where human behaviors are located at a relatively low computational cost, and perform frame sampling on the key video segments, so as to screen out the image frames containing human behavior information.
[0119] The problem of single data modality used. Existing methods usually only use the RGB data modality and do not use additional data for supervision. The present invention inputs the text information associated with the video into an encoder for feature extraction, and obtains the recognition result of human actions by calculating the similarity between the video features and the text features.
[0120] The present invention focuses on solving the problems that existing models are difficult to screen out key frames in videos and have poor recognition effects on long videos, improves the defects and deficiencies of existing solutions, and promotes its wider application in practice. This method not only performs well on the clipped video segment dataset UCF-101; it can achieve an average recognition accuracy of 92.5% on the unclipped long video dataset ActivityNet.
[0121] Those of ordinary skill in the art can understand that the drawings are only schematic diagrams of an embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.
[0122] From the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0123] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the apparatus or system embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description of the method embodiments. The apparatus and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0124] As mentioned above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A human behavior recognition method for a multi-path network based on key frame selection, characterized in that, Including: Sampling the video data to be recognized to obtain multiple video segments; Performing feature acquisition on each video segment, using a multi-layer perceptron and a normalization function to generate a probability distribution, and screening out the video segment where the human action is located according to the probability distribution; Inputting the video segment where the human action is located into a multi-path video-text encoder classification network, learning spatio-temporal features from the video segment through a temporal encoder and a spatial encoder, learning text features in the video segment through a text encoder, and obtaining the recognition result of the human behavior of the video to be recognized by calculating the similarity between the spatio-temporal features and the text features The performing feature acquisition on each video segment, using a multi-layer perceptron and a normalization function to generate a probability distribution, includes: The key-frame sampling module uses the pre-trained ResNet-18 to segment N videos: V1, V2, V3…V N for feature extraction. The features of each video segment are concatenated, and the probability distribution of each frame in each video segment is generated through a multi-layer perceptron and a normalization function. For each video segment the corresponding probability distribution is The sum of the probability distributions corresponding to all video segments is 1, that is The probability distribution of the j-th frame in the i-th video segment V i is calculated as follows: R represents a pre-trained ResNet-18 network, MLP represents a multi-layer perceptron, and softmax represents a normalization function; Let each video segment V i contain D frame images, and the probability distribution corresponding to each frame image is 2. The method according to claim 1, wherein The sampling the video data to be recognized to obtain multiple video segments, includes: Use a lightweight key-frame sampling module to decode the video data to be recognized into a sequence of image frames using FFmpeg. Use the random sampling strategy of the Temporal Segment Network (TSN) to sparsely sample the sequence of image frames. The number of sampled frames is T frames. Segment the T image frames to obtain N video segments: V1, V2, V3…V N , and each video segment contains multiple image frames.
3. The method according to claim 1, wherein The screening out the video segment where the human action is located according to the probability distribution, includes: Select video segments F1, F2, F3... F where human actions are located through a key video segment selection strategy based on the probability distribution corresponding to video segmentation. Z , and the key video segment selection strategy includes: (a) Probability distribution of selected frame images Two video segments where the maximum value and the second maximum value are located; (b) Sum of probability distributions of all frame images The two largest video segments; (c) Probability distribution of selected frame images Beyond probability mean The two video segments with the largest number, T = 64.
4. The method according to claim 3, characterized in that, The inputting the video segment where the human action is located into a multi-path video-text encoder classification network, learning spatio-temporal features from the video segment through a temporal encoder and a spatial encoder, learning text features in the video segment through a text encoder, and obtaining the recognition result of the human behavior of the video to be recognized by calculating the similarity between the spatio-temporal features and the text features, includes: The multi-path video-text encoder network includes a video encoder branch and a text encoder branch. The video encoder branch includes a spatial flow structure and a temporal flow structure. The spatial flow structure acquires image information of different resolutions, and the temporal flow structure acquires image information of different frame rates; The spatial stream structure randomly samples T frames of images from each of the key video segments F1, F2, F3... F output by the key frame sampling module, and performs random cropping to obtain high-resolution images with an input video dimension of T×C×H×W; Z S in = Crop(Random(F1, F2, F3…F Z )) (2―3) Among them, Random represents randomly selecting T frames, Crop represents performing a random cropping operation on the image frames, and S in represents the high-resolution video input to the spatial encoder; The time flow structure selects all the key video segments F1, F2, F3…F Z , randomly crops and downsamples each video segment to obtain an input video dimension of Input the high-frame-rate frame sequence into the network; T in = DownSample(Crop(F1, F2, F3…F Z )) (2―4) Among them, Crop represents performing a random cropping operation on the image frame, DownSample represents performing a downsampling operation on the video, and T in represents the high-frame-rate video finally input into the time-domain encoder; The spatial encoder uses a backbone network to extract features from the input high-resolution video, obtaining spatial information S out , and the temporal encoder uses a backbone network to extract features from the input high-frame-rate video, obtaining temporal features T out . The temporal features are aligned and fused with the spatial information to obtain spatio-temporal features F of the entire video v ; F v = S out + Reshape(T out ) (2 - 5) Among them, Reshape represents performing a dimensional reorganization operation on the features, adjusting the size of its feature map to T×N×D, and F v represents the spatio-temporal features of the entire video in the end; The text encoder branch uses a Transformer with a network depth of 12 layers and eight attention heads for feature extraction to obtain the text features of the entire video; Calculate the similarity between the spatio-temporal features and text features of each video, and determine the classification result y of the human behavior of the video according to the similarity * ; y * = argmax(P(f(x, y)|θ) (2-6) Among them, x represents the spatio-temporal features of the input video, y represents the text information of the corresponding label of the video, f represents the function for calculating similarity, and the label with the highest similarity to the video is the classification result y of the human behavior in the video * .
Citation Information
Patent Citations
Human body action recognition method and system based on multi-modal semantic embedding
CN115953714A
Product production quality monitoring method and device
CN116797091A