Human Behavior Recognition Method, Device, Equipment and Medium for Low-Quality Videos

By obtaining and aggregating frame difference maps in low-quality videos, combining CLIP model and video-specific prompt generator, the problem of human behavior recognition accuracy in low-quality videos is solved, and efficient and accurate behavior recognition is achieved.

CN120126221BActive Publication Date: 2025-07-11PEKING UNIV SHENZHEN GRADUATE SCHOOL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510610669.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-07-11
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

When the prior art recognizes human behavior on low-quality videos, poor video quality leads to a decrease in the accuracy of human detection and recognition.

Method used

By obtaining the pre-order frame difference graph, post-order frame difference graph and average frame difference graph of the video frame, cross-frame semantic aggregation is performed, and the behavior label is determined by combining the CLIP model and the video-specific prompt generator.

Benefits of technology

Effectively suppress background noise, improve the accuracy and efficiency of human behavior recognition in low-quality videos, and maintain efficient and accurate identification of human behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126221B_ABST
    Figure CN120126221B_ABST
Patent Text Reader

Abstract

The present application discloses a human behavior recognition method, device, equipment and medium for low-quality videos. The method includes a pre-frame difference map, a post-frame difference map and an average frame difference map corresponding to video frames; performing cross-frame semantic aggregation on the pre-frame difference map, the post-frame difference map and the average frame difference map corresponding to each video frame to obtain a feature representation corresponding to each video frame; and determining a behavior label of the video sequence to be recognized based on the feature representation corresponding to each video frame. The present application first obtains the pre-frame difference map, the post-frame difference map and the average frame difference map to perform inter-frame noise suppression, and then performs cross-frame semantic aggregation based on the pre-frame difference map, the post-frame difference map and the average frame difference map to aggregate rich spatio-temporal information. In this way, not only can background noise and interference be reduced while maintaining key contour information, but also rich spatio-temporal information can be obtained, effectively improving the accuracy of behavior recognition in encrypted videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and particularly relates to a method, device, equipment and medium for human behavior recognition for low-quality videos. Background Art

[0002] With the wide application of intelligent devices and monitoring systems, human behavior recognition technology plays an important role in fields such as public security, medical supervision, and smart homes. However, due to equipment aging and other uncontrollable factors, the captured videos may have situations such as out-of-focus, overexposure, and low resolution. How to maintain the accuracy of human behavior recognition on such low-quality videos has become a key challenge. The existing behavior recognition methods generally first use publicly available object detection algorithms (such as YOLO) to detect humans, and then recognize the human behavior of the detected humans. However, when performing human behavior recognition on low-quality videos, the accuracy of human detection will be affected by the poor video quality, and further affect the accuracy of the recognized human behavior.

[0003] Therefore, the existing technology still needs to be improved. Summary of the Invention

[0004] The technical problem to be solved by this application is to provide a method, device, equipment and medium for human behavior recognition for low-quality videos in view of the deficiencies of the existing technology.

[0005] To solve the above technical problem, the first aspect of this application provides a method for human behavior recognition for low-quality videos. Specifically, the method for human behavior recognition for low-quality videos includes:

[0006] Obtain the pre-frame difference map between each video frame in the video sequence to be recognized and its corresponding previous video frame, the post-frame difference map between the video frame and its corresponding subsequent video frame, and the average frame difference map between the video frame and the average frame of the video sequence to be recognized;

[0007] Perform cross-frame semantic aggregation on the pre-frame difference map, post-frame difference map, and average frame difference map corresponding to each video frame to obtain the feature representation corresponding to each video frame;

[0008] Based on the feature representation corresponding to each video frame in the video sequence to be recognized, determine the behavior label of the video sequence to be recognized.

[0009] The method for human behavior recognition for low-quality videos, wherein the step of obtaining the pre-frame difference map between each video frame in the video sequence to be recognized and its corresponding previous video frame, the post-frame difference map between the video frame and its corresponding subsequent video frame, and the average frame difference map between the video frame and the average frame of the video sequence to be recognized specifically includes:

[0010] Obtain the average frame of the video sequence to be recognized, as well as the previous video frame and the subsequent video frame of each video frame.

[0011] Perform frame difference operations on each video frame respectively with its previous video frame, the video frame itself, and the average frame to obtain the corresponding previous frame difference map, subsequent frame difference map, and average frame difference map for each video frame.

[0012] The method for human behavior recognition for low-quality videos, wherein the process of obtaining the previous video frame and the subsequent video frame of the video frame specifically includes:

[0013] Read the body ratio and motion amplitude of the video frame, and determine the frame interval corresponding to the video frame based on the body ratio and the motion amplitude.

[0014] Select the previous video frame and the subsequent video frame for the video frame in the video sequence to be recognized according to the frame interval.

[0015] The method for human behavior recognition for low-quality videos, wherein the cross-frame semantic aggregation of the previous frame difference map, subsequent frame difference map, and average frame difference map corresponding to each video frame to obtain the feature representation corresponding to each video frame specifically includes:

[0016] Obtain the previous weight of the previous frame difference map corresponding to each video frame, the subsequent weight of the subsequent frame difference map, and the average weight of the average frame difference map.

[0017] Based on the previous weight, subsequent weight, and average weight, perform weighted combination of the previous frame difference map, subsequent frame difference map, and average frame difference map corresponding to each video frame to obtain the fused frame difference map for each video frame.

[0018] Extract features from the fused frame difference map of each video frame to obtain the feature representation of each video frame.

[0019] The method for human behavior recognition for low-quality videos, wherein the determination of the behavior label of the video sequence to be recognized based on the feature representation corresponding to each video frame in the video sequence to be recognized specifically includes:

[0020] Obtain the text information corresponding to the video sequence to be recognized, and determine the text representation corresponding to the text information through the text encoder in the CLIP model.

[0021] Determine the high-dimensional feature representation through the video encoder in the CLIP model based on the feature representation corresponding to each video frame in the video sequence to be recognized, and determine the global video representation based on the high-dimensional feature representation of each video frame.

[0022] Based on the text representation and the global video representation, determine the action label of the video sequence to be recognized.

[0023] The above-mentioned human action recognition method for low-quality videos, wherein the determination of the global video representation based on the high-dimensional feature representation of each video frame specifically includes:

[0024] Input the high-dimensional feature representation corresponding to each video frame into the cross-frame interaction Transformer, and output the spatio-temporal feature representation corresponding to each video frame through the cross-frame interaction module;

[0025] Input the spatio-temporal feature representation of each video frame into the spatio-temporal fusion module, and output the global video representation through the spatio-temporal fusion module.

[0026] The above-mentioned human action recognition method for low-quality videos, wherein the determination of the action label of the video sequence to be recognized based on the text representation and the global video representation specifically includes:

[0027] Input the text representation and the global video representation into a video-specific prompt generator;

[0028] Capture the dependency association between the text representation and the global video representation through the self-attention mechanism in the video-specific prompt generator to form an intermediate text representation;

[0029] Determine the video-specific prompt based on the intermediate text representation and the global video representation through the feed-forward network in the video-specific prompt generator;

[0030] Fuse the video-specific prompt with the text representation to obtain an enhanced text representation;

[0031] Calculate the similarity between the enhanced text representation and the global video representation, and determine the action label of the video sequence to be recognized based on the similarity.

[0032] The second aspect of the present application provides a human action recognition device for low-quality videos, wherein the human action recognition device for low-quality videos specifically includes:

[0033] An inter-frame noise suppression module, configured to obtain the pre-frame difference map between each video frame in the video sequence to be recognized and its corresponding previous video frame, the post-frame difference map between the video frame and its corresponding subsequent video frame, and the average frame difference map between the video frame and the average frame of the video sequence to be recognized;

[0034] A cross-frame semantic aggregation module, configured to perform cross-frame semantic aggregation on the pre-frame difference map, the post-frame difference map, and the average frame difference map corresponding to each video frame to obtain the feature representation corresponding to each video frame;

[0035] A behavior recognition module, configured to determine a behavior label of the video sequence to be recognized based on the feature representation corresponding to each video frame in the video sequence to be recognized.

[0036] A third aspect of the present application provides a computer-readable storage medium storing one or more programs, which can be executed by one or more processors to implement the steps in any one of the above-mentioned human behavior recognition methods for low-quality videos.

[0037] A fourth aspect of the present application provides a terminal device, which includes: a processor and a memory;

[0038] The memory stores a computer-readable program executable by the processor;

[0039] When the processor executes the computer-readable program, the steps in any one of the above-mentioned human behavior recognition methods for low-quality videos are implemented.

[0040] Beneficial effects: Compared with the prior art, the present application provides a human behavior recognition method, device, equipment and medium for low-quality videos. The method includes obtaining a pre-frame difference map between each video frame in the video sequence to be recognized and its corresponding pre-order video frame, a post-frame difference map between the video frame and its corresponding post-order video frame, and an average frame difference map between the video frame and the average frame of the video sequence to be recognized; performing cross-frame semantic aggregation on the pre-frame difference map, the post-frame difference map, and the average frame difference map corresponding to each video frame to obtain a feature representation corresponding to each video frame; and determining a behavior label of the video sequence to be recognized based on the feature representation corresponding to each video frame in the video sequence to be recognized. The present application first obtains the pre-frame difference map, the post-frame difference map, and the average frame difference map to suppress inter-frame noise, and then performs cross-frame semantic aggregation based on the pre-frame difference map, the post-frame difference map, and the average frame difference map to aggregate rich spatio-temporal information. In this way, not only can background noise and interference be reduced while maintaining key contour information, but also rich spatio-temporal information can be obtained, effectively improving the accuracy of behavior recognition in encrypted videos. Especially for low-quality videos, it is also possible to eliminate interference, highlight the main body, and at the same time maintain the efficient and accurate recognition of human behavior. Description of the Drawings

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained without creative efforts based on these drawings.

[0042] Figure 1 Flowchart of the human behavior recognition method for low-quality videos provided by the embodiments of the present application.

[0043] Figure 2 Principle flowchart of an example of the human behavior recognition method for low-quality videos provided by the embodiments of the present application.

[0044] Figure 3 Principle block diagram of the human behavior recognition device for low-quality videos provided by the embodiments of the present application.

[0045] Figure 4 Principle block diagram of the terminal device provided by the embodiments of the present application. Detailed implementation manners

[0046] The embodiments of the present application provide a human behavior recognition method, device, equipment and medium for low-quality videos. To make the objectives, technical solutions and effects of the present application clearer and more definite, the following further describes the present application in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0047] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0048] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0049] It should be understood that the sequence numbers and magnitudes of the steps in this embodiment do not mean the order of execution. The execution order of each process is determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0050] The following further describes the application content by describing the embodiments in conjunction with the accompanying drawings.

[0051] This embodiment provides a human behavior recognition method for low-quality videos, as Figure 1 shown, the method includes:

[0052] S10. Obtain the pre-frame difference map between each video frame in the video sequence to be recognized and its corresponding previous video frame, the post-frame difference map between the video frame and its corresponding subsequent video frame, and the average frame difference map between the video frame and the average frame of the video sequence to be recognized.

[0053] Specifically, the video sequence to be recognized may include all video frames in the video collected by an image acquisition device, or may include some video frames in the video collected by the image acquisition device. In the embodiments of the present application, the video sequence to be recognized is a low-quality video sequence. For example, the video sequence to be recognized is a video frame sequence collected by a camera installed in a public place, etc.

[0054] Furthermore, since the background noise in the low-quality video is large and the human body area occupies a small part of the video, when performing human body recognition on the low-quality video, the background noise and interference in the low-quality video can be reduced first. For this reason, in the embodiments of the present application, after obtaining the video sequence to be recognized, the frame difference method can be used to effectively reduce the background noise and interference while maintaining the key contour information, as Figure 2 shown, the frame difference includes the pre-frame difference map, the post-frame difference map, and the average frame difference map. The pre-frame difference map is used to reflect the residual information between the video frame and its previous video frame, the post-frame difference map is used to reflect the residual information between the video frame and its subsequent video frame, and the average frame difference map is used to reflect the residual information between the video frame and the average video frame of the video sequence to be recognized.

[0055] In one implementation, the obtaining of the pre-frame difference map between each video frame in the video sequence to be recognized and its corresponding previous video frame, the post-frame difference map between the video frame and its corresponding subsequent video frame, and the average frame difference map between the video frame and the average frame of the video sequence to be recognized specifically includes:

[0056] Obtain the average frame of the video sequence to be recognized and the previous video frame and the subsequent video frame of each video frame;

[0057] Perform frame difference operations on each video frame and its previous video frame, each video frame, and the average frame respectively to obtain the pre-frame difference map, the post-frame difference map, and the average frame difference map corresponding to each video frame.

[0058] Specifically, the average frame is the average frame of the video sequence to be recognized. The background noise can be suppressed through the average frame. Herein, the average frame is determined based on all video frames of the video sequence to be recognized; it can also be determined based on some video frames of the video sequence to be recognized. For example, several key video frames in the video sequence to be recognized are selected, and then the average frame is determined based on these several key video frames. In the embodiments of the present application, the average frame is determined based on all video frames in the video sequence to be recognized. Then, for the video sequence to be recognized, it is expressed as , represents the video frame at time , represents the number of video frames in the video sequence to be recognized. The pixel value of each pixel point in each video frame in the video sequence to be recognized is , represents the pixel position of the pixel point. The average frame can be expressed as:

[0059] ,

[0060] wherein, represents the average frame.

[0061] After obtaining the average frame, the frame - to - frame residual between the video frame and the average frame can be calculated to obtain the average - frame difference map, the frame - to - frame residual between the video frame and the previous video frame can be calculated to obtain the previous - frame difference map, and the frame - to - frame residual between the video frame and the subsequent video frame can be calculated to obtain the subsequent - frame difference map. Herein, the average - frame difference map, the previous - frame difference map, and the subsequent - frame difference map are all single - channel grayscale images, and can be respectively expressed as:

[0062] ,

[0063] ,

[0064] ,

[0065] wherein, represents the previous - frame difference map, represents the subsequent - frame difference map, represents the average - frame difference map, represents the video frame, represents the previous video frame, represents the subsequent video frame.

[0066] Further, the preceding video frame and the subsequent video frame corresponding to the video frame are both included in the video frame sequence to be recognized. The preceding video frame is the video frame located before the video frame in chronological order, and the subsequent video frame is the video frame located after the video frame in chronological order. Among them, the preceding video frame and the subsequent video frame can be randomly selected from the video frame sequence to be recognized, or can be selected according to a preset rule. Moreover, the preceding video frame and the subsequent video frame can be consecutive video frames of the video frame, or can be non-consecutive video frames of the video frame, etc.

[0067] In a specific implementation manner, the process of obtaining the preceding video frame and the subsequent video frame of the video frame specifically includes:

[0068] Read the body ratio and motion amplitude of the video frame, and determine the frame interval corresponding to the video frame based on the body ratio and the motion amplitude;

[0069] Select the preceding video frame and the subsequent video frame for the video frame in the video sequence to be recognized according to the frame interval.

[0070] Specifically, the body ratio refers to the proportion of the human body in the video frame, and the motion amplitude refers to the motion amplitude of the human body in the video frame. That is to say, when obtaining the preceding video frame and the subsequent video frame, the body ratio and the motion amplitude in the video frame will be recognized first, and then the frame interval corresponding to the video frame will be determined based on the body ratio and the motion amplitude. Among them, the frame interval refers to the number of video frames between the video frame and the preceding video frame or the subsequent video frame, and the frame interval between the video frame and the preceding video frame is the same as the frame interval between the video frame and the subsequent video frame. That is to say, after obtaining the frame interval, the preceding video frame and the subsequent video frame can be determined according to the frame interval. That is, the preceding video frame is selected from the video frames before the video frame in chronological order, and the subsequent video frame is selected from the video frames after the video frame in chronological order. The preceding video frame is separated from the video frame by the number of video frames of the frame interval, and the subsequent video frame is also separated from the video frame by the number of video frames of the frame interval. Of course, in practical applications, the frame interval between the preceding video frame and the video frame and the frame interval between the subsequent video frame and the video frame can also be different. For example, the first frame interval between the preceding video frame and the video frame can be determined first based on the body ratio and the motion amplitude, and then the second frame interval between the subsequent video frame and the video frame can be determined based on the first frame interval between the preceding video frame and the video frame (such as presetting the difference between the first frame interval and the second frame interval, and after obtaining the first frame interval, calculating the sum of the first frame interval and the difference to obtain the second frame interval, etc.).

[0071] In the embodiments of the present application, the frame interval of each video frame can be determined according to the proportion of the main body and the movement amplitude of each video frame. In this way, the appropriate frame interval can be flexibly selected according to the time characteristics or movement patterns in the video, so as to improve the ability to recognize behaviors at different time scales, and further improve the accuracy of human form recognition. At the same time, in the embodiments of the present application, by obtaining the pre-frame difference map, the post-frame difference map, and the average frame difference map to determine the subsequent feature representation, the background noise and interference in the low-quality video can be effectively suppressed, and at the same time, the key contour information can be maintained, enhancing the recognition ability of the moving subject.

[0072] S20. Perform cross-frame semantic aggregation on the pre-frame difference map, the post-frame difference map, and the average frame difference map corresponding to each video frame to obtain the feature representation corresponding to each video frame.

[0073] Specifically, after obtaining the pre-frame difference map, the post-frame difference map, and the average frame difference map through inter-frame noise suppression, perform cross-frame semantic aggregation on the pre-frame difference map, the post-frame difference map, and the average frame difference map to further process the cross-frame semantic information. Among them, cross-frame semantic aggregation can combine the residuals between different frames through dynamic weighting, so that the feature representation corresponding to each video frame aggregates rich spatio-temporal information.

[0074] Exemplarily, the performing cross-frame semantic aggregation on the pre-frame difference map, the post-frame difference map, and the average frame difference map corresponding to each video frame to obtain the feature representation corresponding to each video frame specifically includes:

[0075] Obtain the pre-weight corresponding to the pre-frame difference map of each video frame, the post-weight corresponding to the post-frame difference map, and the average weight corresponding to the average frame difference map;

[0076] Based on the pre-weight, the post-weight, and the average weight, perform weighted combination on the pre-frame difference map, the post-frame difference map, and the average frame difference map corresponding to each video frame to obtain the fused frame difference map of each video frame;

[0077] Perform feature extraction on the fused frame difference map of each video frame to obtain the feature representation of each video frame.

[0078] Specifically, the pre-weight, the post-weight, and the average weight are the weight coefficients for performing weighted combination on the pre-frame difference map, the post-frame difference map, and the average frame difference map. Among them, the pre-weight, the post-weight, and the average weight of each video frame are the same, or the pre-weight, the post-weight, and the average weight of some video frames are the same, and the pre-weight, the post-weight, and the average weight of some video frames are different, or the pre-weight, the post-weight, and the average weight of each video frame are all different.

[0079] In a specific implementation, the average weight of the video frame adopts the default weight, and the preceding weight and subsequent weight of the video frame can be calculated based on the number of video frames spaced between the preceding video frame, the video frame and the subsequent video frame. Specifically, the process of obtaining the preceding weight and the subsequent weight can be to first obtain the first number of video frames spaced between the preceding video frame and the video frame (i.e., used to determine the frame interval of the preceding video frame) and the second number of video frames spaced between the subsequent video frame and the video frame (i.e., used to determine the frame interval of the subsequent video frame), and then obtain the third number of video frames spaced between the subsequent video frame and the preceding video frame (i.e., used to determine the frame interval of the preceding video frame + the frame interval for determining the subsequent video frame + 1), and then determine the preceding weight and the subsequent weight based on the first number of video frames, the second number of video frames and the third number of video frames. Accordingly, the preceding weight and the subsequent weight can be expressed as:

[0080] ,

[0081] ,

[0082] in, represents the post-order weight, represents the preceding weight, Indicates the frame number of the subsequent video frame. Indicates the frame number of the video frame, Indicates the frame number of the previous video frame.

[0083] Further, after obtaining the preceding weight, the succeeding weight, and the average weight, the preceding frame difference map, the succeeding frame difference map, and the average frame difference map are weightedly combined based on the preceding weight, the succeeding weight, and the average weight to obtain a fused frame difference map, wherein the fused frame difference map can be expressed as:

[0084] ,

[0085] in, represents the fused frame difference map, Represents a weighted join operation.

[0086] After obtaining the fused frame difference map, feature extraction can be performed on the fused frame difference map to obtain a feature representation of the video frame, wherein the fused frame difference map can be passed through a convolution module (such as Convolution module, etc.), through which the feature representation of the video frame is output.

[0087] The embodiments of the present application can more comprehensively capture the motion details in the video and effectively retain the dynamic change information by fusing the semantic information contained in the previous video frame, the video frame, the subsequent video frame and the average frame, thereby further improving the recognition ability of low-quality videos.

[0088] S30. Determine the behavior label of the video sequence to be recognized based on the feature representation corresponding to each video frame in the video sequence to be recognized.

[0089] Specifically, the behavior label is used to represent the behavior category of human behavior, which is determined based on the feature representation corresponding to each video frame in the video sequence to be recognized. That is, the feature representations corresponding to each video frame in the video sequence to be recognized can be fused into a global video representation, and then the behavior label is determined based on the global video representation.

[0090] Exemplarily, the determining the behavior label of the video sequence to be recognized based on the feature representation corresponding to each video frame in the video sequence to be recognized specifically includes:

[0091] Obtain the text information corresponding to the video sequence to be recognized, and determine the text representation corresponding to the text information through the text encoder in the CLIP model;

[0092] Determine the high-dimensional feature representation based on the feature representation corresponding to each video frame in the video sequence to be recognized through the video encoder in the CLIP model, and determine the global video representation based on the high-dimensional feature representations of each video frame;

[0093] Determine the behavior label of the video sequence to be recognized based on the text representation and the global video representation.

[0094] Specifically, the CLIP model includes a text encoder and a video encoder. The input item of the video encoder is the feature representation corresponding to each video frame. The feature representations corresponding to each video frame are encoded through the video encoder to determine the global video representation corresponding to the video sequence to be recognized. The input of the text encoder is the text information corresponding to the video sequence to be recognized, and the text information is encoded through the text encoder to obtain the text representation. The text information can be text information of an inherent category, etc. Among them, the video encoder can encode the feature representation into a high-dimensional feature representation through Patch Embedding (Patch encoder), and then the high-dimensional feature representations corresponding to all video frames perform interactive learning through cross-frame interaction to capture spatio-temporal dependence relationships, and generate a global video representation based on the captured spatio-temporal dependence relationships.

[0095] Exemplarily, the determining the global video representation based on the high-dimensional feature representations of each video frame specifically includes:

[0096] Input the high-dimensional feature representations corresponding to each video frame into the cross-frame interaction module, and output the spatio-temporal feature representations corresponding to each video frame through the cross-frame interaction module;

[0097] Input the spatio-temporal feature representation of each video frame into the spatio-temporal fusion module, and output the global video representation through the spatio-temporal fusion module.

[0098] Specifically, as Figure 2 shown, the cross-frame interaction module is used to perform interactive learning on the high-dimensional features corresponding to the video frames to capture spatio-temporal dependencies to obtain the spatio-temporal feature representation of each frame. Among them, the cross-frame interaction module can adopt the Transformer model, and perform interactive learning on the high-dimensional feature representations of each video frame through the Transformer model to obtain the spatio-temporal feature representation corresponding to each video frame. The spatio-temporal fusion module is used to fuse the spatio-temporal representations of each video frame to generate the global video representation. Among them, the spatio-temporal fusion module can also adopt the Transformer model, and fuse the spatio-temporal representations of each video frame through the Transformer model to generate the global video representation.

[0099] Furthermore, after obtaining the global video representation, visual cues can be generated based on the global video representation and the text representation, and then the action label of the video sequence to be recognized can be determined based on the visual cues and the global video representation. Specifically, determining the action label of the video sequence to be recognized based on the text representation and the global video representation specifically includes:

[0100] Input the text representation and the global video representation into the video-specific cue generator;

[0101] Capture the dependency relationship between the text representation and the global video representation through the self-attention mechanism in the video-specific cue generator to form an intermediate text representation;

[0102] Determine the video-specific cue based on the intermediate text representation and the global video representation through the feed-forward network in the video-specific cue generator;

[0103] Fuse the video-specific cue with the text representation to obtain the enhanced text representation;

[0104] Calculate the similarity between the enhanced text representation and the global video representation, and determine the action label of the video sequence to be recognized based on the similarity.

[0105] Specifically, the video-specific prompt can be determined by a video-specific prompt generator, which is used to combine the text representation with relevant visual information in the video (such as the visual information "bookstore" in the video and the text "reading books", etc.) to provide a more precise semantic context during classification. Among them, the input items of the video-specific prompt generator include the text prompt and the global video prompt. The self-attention mechanism in the video-specific prompt generator captures the dependencies between the text representation and the global video representation to form the video-specific prompt, and then enhances the text representation based on the video-specific prompt to obtain the enhanced text representation. The following uses to represent the text representation, and

[0106]

[0106]

[0107]

[0108] Furthermore, after obtaining the enhanced text representation, the similarity between the global video representation and the enhanced text representation can be calculated, and then the human body label corresponding to the video sequence to be recognized can be determined according to the similarity. Among them, the similarity between the global video representation and the enhanced text representation can use the cosine similarity.

[0108] Of course, in practical applications, other methods can also be used to determine the behavior labels. For example, after obtaining the global video features, the behavior labels can be determined through a classification head, or the behavior labels can be determined through regression operations, etc. There is no specific limitation here. The above implementation method is only an example of a specific implementation method.

[0109] It should be noted that the embodiments of this application illustrate the specific process of the human behavior recognition method for low-quality videos in the form of a method flow. In practical applications, the recognition process of human behavior for low-quality videos can be carried out through a deep learning module (denoted as a human behavior recognition model). Specifically, as Figure 2 shown, the human behavior recognition model may include an inter-frame noise suppression module, a cross-frame semantic aggregation module, and a behavior recognition backbone network. The inter-frame noise suppression module is connected to the cross-frame semantic aggregation module, and the cross-frame semantic aggregation module is connected to the behavior recognition backbone network. The inter-frame noise suppression module is used to obtain the pre-frame difference map between each video frame in the video sequence to be recognized and its corresponding previous video frame, the post-frame difference map between the video frame and its corresponding subsequent video frame, and the average frame difference map between the video frame and the average frame of the video sequence to be recognized. The cross-frame semantic aggregation module is used to perform cross-frame semantic aggregation on the pre-frame difference map, post-frame difference map, and average frame difference map corresponding to each video frame to obtain the feature representation corresponding to each video frame. The behavior recognition backbone network is used to determine the behavior label of the video sequence to be recognized based on the feature representation corresponding to each video frame in the video sequence to be recognized. Among them, the specific implementation processes of the inter-frame noise suppression module, the cross-frame semantic aggregation module, and the behavior recognition backbone network are the same as those in the method flow embodiment above, and will not be specifically described here. Only the loss function used in the training process of the human behavior recognition model will be described here. The loss function used in the training process can be:

[0110] ,

[0111] where, represents the loss function, represents the cosine similarity, represents the number of intrinsic lists in the text information, represents the enhanced text representation, represents the th text representation of the intrinsic category.

[0112] In summary, this embodiment provides a human behavior recognition method for low-quality videos. The method includes obtaining a pre-frame difference map between each video frame in the video sequence to be recognized and its corresponding previous video frame, a post-frame difference map between the video frame and its corresponding subsequent video frame, and an average frame difference map between the video frame and the average frame of the video sequence to be recognized; performing cross-frame semantic aggregation on the pre-frame difference map, post-frame difference map, and average frame difference map corresponding to each video frame to obtain a feature representation corresponding to each video frame; and determining the behavior label of the video sequence to be recognized based on the feature representation corresponding to each video frame in the video sequence to be recognized. In this application, the pre-frame difference map, post-frame difference map, and average frame difference map are first obtained for inter-frame noise suppression, and then cross-frame semantic aggregation is performed based on the pre-frame difference map, post-frame difference map, and average frame difference map to aggregate rich spatio-temporal information. This can not only reduce background noise and interference while maintaining key contour information, but also obtain rich spatio-temporal information, effectively improving the accuracy of behavior recognition in encrypted videos. Especially for low-quality videos, it can eliminate interference, highlight the main body, and at the same time maintain efficient and accurate recognition of human behavior.

[0113] Based on the above human behavior recognition method for low-quality videos, this embodiment provides a human behavior recognition device for low-quality videos, as Figure 3 shown. The human behavior recognition device for low-quality videos specifically includes:

[0114] An inter-frame noise suppression module 100, configured to obtain a pre-frame difference map between each video frame in the video sequence to be recognized and its corresponding previous video frame, a post-frame difference map between the video frame and its corresponding subsequent video frame, and an average frame difference map between the video frame and the average frame of the video sequence to be recognized;

[0115] A cross-frame semantic aggregation module 200, configured to perform cross-frame semantic aggregation on the pre-frame difference map, post-frame difference map, and average frame difference map corresponding to each video frame to obtain a feature representation corresponding to each video frame;

[0116] A behavior recognition module 300, configured to determine the behavior label of the video sequence to be recognized based on the feature representation corresponding to each video frame in the video sequence to be recognized.

[0117] Based on the above human behavior recognition method for low-quality videos, this embodiment provides a computer-readable storage medium. The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the human behavior recognition method for low-quality videos as described in the above embodiment.

[0118] Based on the above-mentioned human behavior recognition method for low-quality videos, the present application also provides a terminal device, such as Figure 4 shown, which includes at least one processor 20; a display screen 21; and a memory 22, and may also include a communication interface 23 and a bus 24. Among them, the processor 20, the display screen 21, the memory 22, and the communication interface 23 can complete mutual communication through the bus 24. The display screen 21 is set to display a user guidance interface preset in the initial setting mode. The communication interface 23 can transmit information. The processor 20 can call the logical instructions in the memory 22 to execute the method in the above-mentioned embodiment.

[0119] In addition, when the logical instructions in the above-mentioned memory 22 can be implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium.

[0120] The memory 22, as a computer-readable storage medium, can be set to store software programs and computer-executable programs, such as program instructions or modules corresponding to the methods in the embodiments of the present disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions, or modules stored in the memory 22, that is, implements the methods in the above-mentioned embodiments.

[0121] The memory 22 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal device, etc. In addition, the memory 22 may include a high-speed random access memory and may also include a non-volatile memory. For example, various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs can also be transient storage media.

[0122] In addition, the specific processes of loading and executing multiple instructions in the above-mentioned storage medium and the terminal device have been described in detail in the above method and will not be repeated here one by one.

[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A human behavior recognition method for low-quality videos, characterized in that, The described human behavior recognition method for low-quality videos specifically includes: Obtain the pre-frame difference map between each video frame in the video sequence to be recognized and its corresponding previous video frame, the post-frame difference map between the video frame and its corresponding subsequent video frame, and the average frame difference map between the video frame and the average frame of the video sequence to be recognized; Perform cross-frame semantic aggregation on the pre-frame difference map, post-frame difference map, and average frame difference map corresponding to each video frame to obtain the feature representation corresponding to each video frame; Based on the feature representation corresponding to each video frame in the video sequence to be recognized, determine the behavior label of the video sequence to be recognized; Among them, the step of determining the behavior label of the video sequence to be recognized based on the feature representation corresponding to each video frame in the video sequence to be recognized specifically includes: Obtain the text information corresponding to the video sequence to be recognized, and determine the text representation corresponding to the text information through the text encoder in the CLIP model; Based on the feature representation corresponding to each video frame in the video sequence to be recognized, determine the high-dimensional feature representation through the video encoder in the CLIP model, and determine the global video representation based on the high-dimensional feature representation of each video frame; Based on the text representation and the global video representation, determine the behavior label of the video sequence to be recognized; The step of determining the behavior label of the video sequence to be recognized based on the text representation and the global video representation specifically includes: Input the text representation and the global video representation into the video-specific prompt generator; Capture the dependency association between the text representation and the global video representation through the self-attention mechanism in the video-specific prompt generator to form an intermediate text representation; Based on the intermediate text representation and the global video representation, determine the video-specific prompt through the feed-forward network in the video-specific prompt generator; Fuse the video-specific prompt with the text representation to obtain an enhanced text representation; Calculate the similarity between the enhanced text representation and the global video representation, and determine the behavior label of the video sequence to be recognized based on the similarity.

2. The human behavior recognition method for low-quality videos according to claim 1, wherein The step of obtaining the pre-frame difference map between each video frame in the video sequence to be recognized and its corresponding previous video frame, the post-frame difference map between the video frame and its corresponding subsequent video frame, and the average frame difference map between the video frame and the average frame of the video sequence to be recognized specifically includes: Obtain the average frame of the video sequence to be recognized, as well as the previous video frame and the subsequent video frame of each video frame; Perform frame difference operations on each video frame and its previous video frame, each video frame and the average frame respectively to obtain the pre-frame difference map, post-frame difference map, and average frame difference map corresponding to each video frame.

3. The human behavior recognition method for low-quality videos according to claim 2, wherein The process of obtaining the previous video frame and the subsequent video frame of the video frame specifically includes: Read the body occupancy ratio and motion amplitude of the video frame, and determine the frame interval corresponding to the video frame based on the body occupancy ratio and the motion amplitude; Select the previous video frame and the subsequent video frame for the video frame in the video sequence to be recognized according to the frame interval.

4. The human behavior recognition method for low-quality videos according to claim 1, characterized in that Performing cross-frame semantic aggregation on the pre-frame difference map, post-frame difference map, and average-frame difference map corresponding to each video frame to obtain the feature representation corresponding to each video frame specifically includes: Obtaining the pre-order weight of the pre-frame difference map corresponding to each video frame, the post-order weight of the post-frame difference map, and the average weight of the average-frame difference map; Based on the pre-order weight, post-order weight, and average weight, performing weighted combination of the pre-frame difference map, post-frame difference map, and average-frame difference map corresponding to each video frame to obtain the fused frame difference map of each video frame; Performing feature extraction on the fused frame difference map of each video frame to obtain the feature representation of each video frame.

5. The human behavior recognition method for low-quality videos according to claim 1, characterized in that The determining the global video representation based on the high-dimensional feature representation of each video frame specifically includes: Inputting the high-dimensional feature representation corresponding to each video frame into a cross-frame interaction Transformer, and outputting the spatio-temporal feature representation corresponding to each video frame through the cross-frame interaction module; Inputting the spatio-temporal feature representation of each video frame into a spatio-temporal fusion module, and outputting the global video representation through the spatio-temporal fusion module.

6. An apparatus for human behavior recognition for low-quality videos, characterized in that, The human behavior recognition device for low-quality videos specifically includes: An inter-frame noise suppression module, configured to obtain the pre-frame difference map between each video frame in the video sequence to be recognized and its corresponding pre-order video frame, the post-frame difference map between it and its corresponding post-order video frame, and the average-frame difference map between it and the average frame of the video sequence to be recognized; A cross-frame semantic aggregation module, configured to perform cross-frame semantic aggregation on the pre-frame difference map, post-frame difference map, and average-frame difference map corresponding to each video frame to obtain the feature representation corresponding to each video frame; A behavior recognition module, configured to determine the behavior label of the video sequence to be recognized based on the feature representation corresponding to each video frame in the video sequence to be recognized; Among them, the determining the behavior label of the video sequence to be recognized based on the feature representation corresponding to each video frame in the video sequence to be recognized specifically includes: Obtaining the text information corresponding to the video sequence to be recognized, and determining the text representation corresponding to the text information through the text encoder in the CLIP model; Determining the high-dimensional feature representation based on the feature representation corresponding to each video frame in the video sequence to be recognized through the video encoder in the CLIP model, and determining the global video representation based on the high-dimensional feature representation of each video frame; Based on the text representation and the global video representation, determining the behavior label of the video sequence to be recognized; The determining the behavior label of the video sequence to be recognized based on the text representation and the global video representation specifically includes: Inputting the text representation and the global video representation into a video-specific prompt generator; Capturing the dependency relationship between the text representation and the global video representation through the self-attention mechanism in the video-specific prompt generator to form an intermediate text representation; Determining the video-specific prompt based on the intermediate text representation and the global video representation through the feed-forward network in the video-specific prompt generator; Fusing the video-specific prompt with the text representation to obtain an enhanced text representation; Calculate the similarity between the enhanced text representation and the global video representation, and determine the action label of the video sequence to be recognized based on the similarity.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the human action recognition method for low-quality videos according to any one of claims 1-5.

8. A terminal device, characterized in that, Including: A processor and a memory; A computer-readable program executable by the processor is stored on the memory; When the processor executes the computer-readable program, the steps in the human action recognition method for low-quality videos according to any one of claims 1-5 are implemented.

Citation Information

Patent Citations

  • Behavior recognition method, device and equipment

    CN116778568A