Human body behavior recognition method and device for low-quality video, equipment and medium

By obtaining frame difference maps in low-quality videos and performing cross-frame semantic aggregation, the problem of low-quality videos with low-quality videos is solved, and efficient and accurate recognition of human behavior is achieved.

CN120126221AActive Publication Date: 2025-06-10PEKING UNIV SHENZHEN GRADUATE SCHOOL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510610669.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-06-10
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

When the prior art performs human behavior recognition on low-quality videos, the poor video quality will lead to a decrease in the accuracy of human detection and behavior recognition.

Method used

By obtaining the frame difference map between each video frame and its predecessor, post-order and average frames, and performing cross-frame semantic aggregation, the feature representation of each video frame is obtained, and the behavior label is determined based on these feature representations.

Benefits of technology

Effectively suppress background noise and interference in low-quality videos, maintain key contour information, enhance the recognition ability of moving subjects, and improve the accuracy of behavior recognition in low-quality videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126221A_ABST
    Figure CN120126221A_ABST
Patent Text Reader

Abstract

The invention discloses a human body behavior recognition method and device for a low-quality video, equipment and a medium. The method comprises a preorder frame difference chart, a postorder frame difference chart and an average frame difference chart corresponding to video frames. Performing cross-frame semantic aggregation on the preorder frame difference chart, the postorder frame difference chart and the average frame difference chart corresponding to each video frame to obtain a feature representation corresponding to each video frame; and determining a behavior label of the to-be-identified video sequence based on the feature representation corresponding to each video frame. According to the method, the preorder frame difference chart, the post-order frame difference chart and the average frame difference chart are firstly acquired to carry out inter-frame noise suppression, and then cross-frame semantic aggregation is carried out based on the preorder frame difference chart, the post-order frame difference chart and the average frame difference chart to aggregate rich spatio-temporal information; in this way, background noise and interference can be reduced on the premise that key contour information is kept, rich spatio-temporal information can be obtained, and the behavior recognition accuracy in the encrypted video is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and particularly relates to a method, device, equipment and medium for human behavior recognition for low-quality videos. Background Art

[0002] With the wide application of intelligent devices and monitoring systems, human behavior recognition technology plays an important role in fields such as public safety, medical supervision, and smart home. However, due to equipment aging and other uncontrollable factors, the captured videos may have problems such as out-of-focus, overexposure, and low resolution. How to maintain the accuracy of human behavior recognition on such low-quality videos has become a key challenge. Existing behavior recognition methods generally first use publicly available object detection algorithms (such as YOLO) to detect humans, and then recognize the human behavior of the detected humans. However, when performing human behavior recognition on low-quality videos, the accuracy of human detection will be affected due to the poor video quality, and further affect the accuracy of the recognized human behavior.

[0003] Therefore, the existing technology still needs to be improved. Summary of the Invention

[0004] The technical problem to be solved by the present application is to provide a method, device, equipment and medium for human behavior recognition for low-quality videos in view of the deficiencies of the existing technology.

[0005] To solve the above technical problem, a first aspect of the present application provides a method for human behavior recognition for low-quality videos, wherein the method for human behavior recognition for low-quality videos specifically includes: Obtain a pre-frame difference map between each video frame in the video sequence to be recognized and its corresponding previous video frame, a post-frame difference map between the video frame and its corresponding subsequent video frame, and an average frame difference map between the video frame and the average frame of the video sequence to be recognized; Perform cross-frame semantic aggregation on the pre-frame difference map, post-frame difference map, and average frame difference map corresponding to each video frame to obtain a feature representation corresponding to each video frame; Based on the feature representations corresponding to each video frame in the video sequence to be recognized, determine the behavior label of the video sequence to be recognized.

[0006] In the method for human behavior recognition for low-quality videos, the step of obtaining a pre-frame difference map between each video frame in the video sequence to be recognized and its corresponding previous video frame, a post-frame difference map between the video frame and its corresponding subsequent video frame, and an average frame difference map between the video frame and the average frame of the video sequence to be recognized specifically includes: Obtain the average frame of the video sequence to be recognized, as well as the previous video frame and the subsequent video frame of each video frame; Perform frame difference operations on each video frame respectively with the previous video frame of each video frame, the video frame of each video frame, and the average frame to obtain the previous frame difference map, the subsequent frame difference map, and the average frame difference map corresponding to each video frame.

[0007] The human behavior recognition method for low-quality videos, wherein the process of obtaining the previous video frame and the subsequent video frame of the video frame specifically includes: Read the body ratio and motion amplitude of the video frame, and determine the frame interval corresponding to the video frame based on the body ratio and the motion amplitude; Select the previous video frame and the subsequent video frame for the video frame in the video sequence to be recognized according to the frame interval.

[0008] The human behavior recognition method for low-quality videos, wherein the cross-frame semantic aggregation of the previous frame difference map, the subsequent frame difference map, and the average frame difference map corresponding to each video frame to obtain the feature representation corresponding to each video frame specifically includes: Obtain the previous weight of the previous frame difference map corresponding to each video frame, the subsequent weight of the subsequent frame difference map, and the average weight of the average frame difference map; Based on the previous weight, the subsequent weight, and the average weight, perform weighted combination of the previous frame difference map, the subsequent frame difference map, and the average frame difference map corresponding to each video frame to obtain the fused frame difference map of each video frame; Extract features from the fused frame difference map of each video frame to obtain the feature representation of each video frame.

[0009] The human behavior recognition method for low-quality videos, wherein the determination of the behavior label of the video sequence to be recognized based on the feature representation corresponding to each video frame in the video sequence to be recognized specifically includes: Obtain the text information corresponding to the video sequence to be recognized, and determine the text representation corresponding to the text information through the text encoder in the CLIP model; Determine the high-dimensional feature representation through the video encoder in the CLIP model based on the feature representation corresponding to each video frame in the video sequence to be recognized, and determine the global video representation based on the high-dimensional feature representation of each video frame; Based on the text representation and the global video representation, determine the behavior label of the video sequence to be recognized.

[0010] The human behavior recognition method for low-quality videos, wherein the determination of the global video representation based on the high-dimensional feature representation of each video frame specifically includes: Input the high-dimensional feature representation corresponding to each video frame into the cross-frame interaction Transformer, and output the spatio-temporal feature representation corresponding to each video frame through the cross-frame interaction module; Input the spatio-temporal feature representation of each video frame into the spatio-temporal fusion module, and output the global video representation through the spatio-temporal fusion module.

[0011] The method for human behavior recognition for low-quality videos, wherein determining the behavior label of the video sequence to be recognized based on the text representation and the global video representation specifically includes: Input the text representation and the global video representation into the video-specific prompt generator; Capture the dependency relationship between the text representation and the global video representation through the self-attention mechanism in the video-specific prompt generator to form an intermediate text representation; Determine the video-specific prompt based on the intermediate text representation and the global video representation through the feed-forward network in the video-specific prompt generator; Fuse the video-specific prompt with the text representation to obtain an enhanced text representation; Calculate the similarity between the enhanced text representation and the global video representation, and determine the behavior label of the video sequence to be recognized based on the similarity.

[0012] The second aspect of the present application provides a device for human behavior recognition for low-quality videos, wherein the device for human behavior recognition for low-quality videos specifically includes: An inter-frame noise suppression module, configured to obtain the pre-frame difference map between each video frame in the video sequence to be recognized and its corresponding previous video frame, the post-frame difference map between the video frame and its corresponding subsequent video frame, and the average frame difference map between the video frame and the average frame of the video sequence to be recognized; A cross-frame semantic aggregation module, configured to perform cross-frame semantic aggregation on the pre-frame difference map, the post-frame difference map, and the average frame difference map corresponding to each video frame to obtain the feature representation corresponding to each video frame; A behavior recognition module, configured to determine the behavior label of the video sequence to be recognized based on the feature representation corresponding to each video frame in the video sequence to be recognized.

[0013] The third aspect of the present application provides a computer-readable storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in any one of the above-mentioned methods for human behavior recognition for low-quality videos.

[0014] The fourth aspect of the present application provides a terminal device, which includes: a processor and a memory; A computer-readable program executable by the processor is stored on the memory; When the processor executes the computer-readable program, the steps in any of the above-mentioned human behavior recognition methods for low-quality videos are implemented.

[0015] Beneficial effects: Compared with the prior art, the present application provides a human behavior recognition method, device, equipment and medium for low-quality videos. The method includes obtaining a pre-frame difference map between each video frame in the video sequence to be recognized and its corresponding pre-order video frame, a post-frame difference map between the video frame and its corresponding post-order video frame, and an average frame difference map between the video frame and the average frame of the video sequence to be recognized; performing cross-frame semantic aggregation on the pre-frame difference map, post-frame difference map and average frame difference map corresponding to each video frame to obtain a feature representation corresponding to each video frame; and determining a behavior label of the video sequence to be recognized based on the feature representation corresponding to each video frame in the video sequence to be recognized. The present application first obtains the pre-frame difference map, post-frame difference map and average frame difference map to suppress inter-frame noise, and then performs cross-frame semantic aggregation based on the pre-frame difference map, post-frame difference map and average frame difference map to aggregate rich spatio-temporal information. In this way, not only can background noise and interference be reduced while maintaining key contour information, but also rich spatio-temporal information can be obtained, effectively improving the accuracy of behavior recognition in encrypted videos. Especially for low-quality videos, it can also eliminate interference, highlight the main body, and maintain efficient and accurate recognition of human behavior. Description of the Drawings

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a flowchart of the human behavior recognition method for low-quality videos provided by the embodiment of the present application.

[0018] Figure 2 It is a principle flowchart of an example of the human behavior recognition method for low-quality videos provided by the embodiment of the present application.

[0019] Figure 3 It is a principle block diagram of the human behavior recognition device for low-quality videos provided by the embodiment of the present application.

[0020] Figure 4 It is a principle block diagram of the terminal device provided by the embodiment of the present application. Detailed Embodiments

[0021] An embodiment of the present application provides a method, device, equipment and medium for human behavior recognition in low-quality videos. To make the purpose, technical solution and effect of the present application clearer and more definite, the following further describes the present application in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0022] Those skilled in the art of the present technology can understand that unless specifically stated, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used here may include wireless connection or wireless coupling. The phrase "and / or" used here includes all or any unit and all combinations of one or more related listed items.

[0023] Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used here have the same meaning as the general understanding of those of ordinary skill in the art to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0024] It should be understood that the sequence numbers and magnitudes of the steps in this embodiment do not mean the order of execution. The order of execution of each process is determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0025] The following further describes the application content with reference to the accompanying drawings and by way of description of the embodiments.

[0026] This embodiment provides a method for human behavior recognition in low-quality videos, as Figure 1 shown, the method includes: S10. Obtain a pre-order frame difference map between each video frame in the video sequence to be recognized and its corresponding pre-order video frame, a post-order frame difference map between the video frame and its corresponding post-order video frame, and an average frame difference map between the video frame and the average frame of the video sequence to be recognized.

[0027] Specifically, the video sequence to be recognized may include all video frames in the video captured by the image acquisition device, or may include some video frames in the video captured by the image acquisition device. In the embodiments of the present application, the video sequence to be recognized is a low-quality video sequence. For example, the video sequence to be recognized is a video frame sequence captured by a camera installed in a public place, etc.

[0028] Furthermore, since the background noise in the low-quality video is large and the human body region occupies a small part of the video, when performing human body recognition on the low-quality video, the background noise and interference in the low-quality video can be reduced first. For this purpose, in the embodiments of the present application, after obtaining the video sequence to be recognized, the frame difference method can be used to effectively reduce the background noise and interference while maintaining the key contour information, as Figure 2 shown, the frame difference includes the pre-frame difference map, the post-frame difference map, and the average frame difference map. The pre-frame difference map is used to reflect the residual information between a video frame and its previous video frame, the post-frame difference map is used to reflect the residual information between a video frame and its subsequent video frame, and the average frame difference map is used to reflect the residual information between a video frame and the average-order video frame of the video sequence to be recognized.

[0029] In one implementation manner, obtaining the pre-frame difference map between each video frame in the video sequence to be recognized and its corresponding previous video frame, the post-frame difference map between each video frame and its corresponding subsequent video frame, and the average frame difference map between each video frame and the average frame of the video sequence to be recognized specifically includes: Obtaining the average frame of the video sequence to be recognized, as well as the previous video frame and the subsequent video frame of each video frame; Performing frame difference operations on each video frame respectively with its previous video frame, its own video frame, and the average frame to obtain the pre-frame difference map, the post-frame difference map, and the average frame difference map corresponding to each video frame.

[0030] Specifically, the average frame is the average frame of the video sequence to be recognized. The background noise can be suppressed through the average frame. Among them, the average frame is determined based on all video frames of the video sequence to be recognized; it can also be determined based on some video frames of the video sequence to be recognized. For example, several key video frames in the video sequence to be recognized are selected, and then the average frame is determined based on the several key video frames, etc. In the embodiments of the present application, the average frame is determined based on all video frames in the video sequence to be recognized. Then, for the video sequence to be recognized, it is expressed as , representing the video frame at time , representing the number of video frames in the video sequence to be recognized, and each video frame in the video sequence to be recognized has pixel values of the pixel points as , Indicates the pixel position of a pixel point, and the average frame can be expressed as: , wherein, represents the average frame.

[0031] After obtaining the average frame, the inter-frame residual between the video frame and the average frame can be calculated to obtain the average frame difference map, the inter-frame residual between the video frame and the previous video frame to obtain the previous frame difference map, and the inter-frame residual between the video frame and the subsequent video frame to obtain the subsequent frame difference map. Among them, the average frame difference map, the previous frame difference map, and the subsequent frame difference map are all single-channel grayscale images and can be respectively expressed as: , , , wherein, represents the previous frame difference map, represents the subsequent frame difference map, represents the average frame difference map, represents the video frame, represents the previous video frame, represents the subsequent video frame.

[0032] Furthermore, the previous video frame and the subsequent video frame corresponding to the video frame are both included in the video frame sequence to be recognized, and the previous video frame is the video frame located before the video frame in chronological order, and the subsequent video frame is the video frame located after the video frame in chronological order. Among them, the previous video frame and the subsequent video frame can be randomly selected from the video frame sequence to be recognized, or can be selected according to a preset rule, and the previous video frame and the subsequent video frame can be consecutive video frames of the video frame, or can be non-consecutive video frames of the video frame, etc.

[0033] In a specific implementation manner, the process of obtaining the previous video frame and the subsequent video frame of the video frame specifically includes: Read the body ratio and motion amplitude of the video frame, and determine the frame interval corresponding to the video frame based on the body ratio and the motion amplitude; Select the previous video frame and the subsequent video frame for the video frame in the video sequence to be recognized according to the frame interval.

[0034] Specifically, the body proportion refers to the proportion of the human body in the video frame, and the movement amplitude refers to the movement amplitude of the human body in the video frame. That is, when obtaining the previous video frame and the subsequent video frame, the body proportion and the movement amplitude in the video frame will be identified first, and then the frame interval corresponding to the video frame is determined based on the body proportion and the movement amplitude. Here, the frame interval refers to the number of video frames between the video frame and the previous video frame or the subsequent video frame, and the frame interval between the video frame and the previous video frame is the same as the frame interval between the video frame and the subsequent video frame. That is, after obtaining the frame interval, the previous video frame and the subsequent video frame can be determined according to the frame interval, that is, the previous video frame is selected from the video frames before the video frame in chronological order and the subsequent video frame is selected from the video frames after the video frame. The previous video frame is separated from the video frame by the number of video frames of the frame interval, and the subsequent video frame is also separated from the video frame by the number of video frames of the frame interval. Of course, in practical applications, the frame interval between the previous video frame and the video may also be different from the frame interval between the subsequent video frame and the video. For example, the first frame interval between the previous video frame and the video can be determined first based on the body proportion and the movement amplitude, and then the second frame interval between the subsequent video frame and the video is determined based on the first frame interval between the previous video frame and the video (such as presetting the difference between the first frame interval and the second frame interval, and after obtaining the first frame interval, calculating the sum of the first frame interval and the difference to obtain the second frame interval, etc.).

[0035] In the embodiment of the present application, the frame interval of each video frame can be determined according to the body proportion and the movement amplitude of each video frame, so that a suitable frame interval can be flexibly selected according to the time characteristics or movement patterns in the video, thereby improving the ability to identify behaviors at different time scales, and further improving the accuracy of human form recognition. At the same time, in the embodiment of the present application, by obtaining the previous frame difference map, the subsequent frame difference map, and the average frame difference map to determine the subsequent feature representation, the background noise and interference in the low-quality video can be effectively suppressed, and at the same time, the key contour information can be maintained, enhancing the recognition ability of the moving body.

[0036] S20. Perform cross-frame semantic aggregation on the previous frame difference map, the subsequent frame difference map, and the average frame difference map corresponding to each video frame to obtain the feature representation corresponding to each video frame.

[0037] Specifically, after obtaining the previous frame difference map, the subsequent frame difference map, and the average frame difference map through inter-frame noise suppression, perform cross-frame semantic aggregation on the previous frame difference map, the subsequent frame difference map, and the average frame difference map to further process the cross-frame semantic information. Here, the cross-frame semantic aggregation can combine the residuals between different frames through dynamic weighting, so that the feature representation corresponding to each video frame aggregates rich spatio-temporal information.

[0038] Exemplarily, performing cross-frame semantic aggregation on the preceding frame difference map, the succeeding frame difference map, and the average frame difference map corresponding to each video frame to obtain the feature representation corresponding to each video frame specifically includes: Obtain the preceding weight of the preceding frame difference map corresponding to each video frame, the succeeding weight corresponding to the succeeding frame difference map, and the average weight corresponding to the average frame difference map; Based on the preceding weight, the succeeding weight and the average weight, the preceding frame difference map, the succeeding frame difference map and the average frame difference map corresponding to each video frame are weightedly combined to obtain a fused frame difference map for each video frame; Feature extraction is performed on the fused frame difference map of each video frame to obtain a feature representation of each video frame.

[0039] Specifically, the preceding weight, the succeeding weight and the average weight are weight coefficients for weighted combination of the preceding frame difference map, the succeeding frame difference map and the average frame difference map, wherein the preceding weight, succeeding weight and the average weight of each video frame are the same, or the preceding weight, succeeding weight and the average weight of some video frames may be the same, and the preceding weight, succeeding weight and the average weight of some video frames may be different, or the preceding weight, succeeding weight and the average weight of each video frame may be different.

[0040] In a specific implementation, the average weight of the video frame adopts the default weight, and the preceding weight and subsequent weight of the video frame can be calculated based on the number of video frames spaced between the preceding video frame, the video frame and the subsequent video frame. Specifically, the process of obtaining the preceding weight and the subsequent weight can be to first obtain the first number of video frames spaced between the preceding video frame and the video frame (i.e., used to determine the frame interval of the preceding video frame) and the second number of video frames spaced between the subsequent video frame and the video frame (i.e., used to determine the frame interval of the subsequent video frame), and then obtain the third number of video frames spaced between the subsequent video frame and the preceding video frame (i.e., used to determine the frame interval of the preceding video frame + the frame interval for determining the subsequent video frame + 1), and then determine the preceding weight and the subsequent weight based on the first number of video frames, the second number of video frames and the third number of video frames. Accordingly, the preceding weight and the subsequent weight can be expressed as: , , in, represents the post-order weight, represents the preceding weight, Indicates the frame number of the subsequent video frame. Indicates the frame number of the video frame, Indicates the frame number of the previous video frame.

[0041] Further, after obtaining the pre-weight, post-weight, and average weight, the pre-frame difference map, post-frame difference map, and average frame difference map are weighted and combined based on the pre-weight, post-weight, and average weight to obtain a fused frame difference map, where the fused frame difference map can be expressed as: , where, represents the fused frame difference map, represents the weighted combination operation.

[0042] After obtaining the fused frame difference map, feature extraction can be performed on the fused frame difference map to obtain the feature representation of the video frame. Among them, the fused frame difference map can be passed through a convolutional module (such as convolutional module, etc.), and the feature representation of the video frame is output through this convolutional module.

[0043] By fusing the semantic information contained in the pre-video frame, video frame, post-video frame, and average frame, the embodiments of the present application can capture the motion details in the video more comprehensively, and can effectively retain the dynamic change information, further improving the recognition ability of low-quality videos.

[0044] S30. Determine the behavior label of the to-be-recognized video sequence based on the feature representation corresponding to each video frame in the to-be-recognized video sequence.

[0045] Specifically, the behavior label is used to represent the behavior category of the human behavior, which is determined based on the feature representation corresponding to each video frame in the to-be-recognized video sequence. That is, the feature representation corresponding to each video frame in the to-be-recognized video sequence can be fused into a global video representation, and then the behavior label is determined based on the global video representation.

[0046] Exemplarily, the determining the behavior label of the to-be-recognized video sequence based on the feature representation corresponding to each video frame in the to-be-recognized video sequence specifically includes: Obtain the text information corresponding to the to-be-recognized video sequence, and determine the text representation corresponding to the text information through the text encoder in the CLIP model; Determine the high-dimensional feature representation based on the feature representation corresponding to each video frame in the to-be-recognized video sequence through the video encoder in the CLIP model, and determine the global video representation based on the high-dimensional feature representation of each video frame; Determine the behavior label of the to-be-recognized video sequence based on the text representation and the global video representation.

[0047] Specifically, the CLIP model includes a text encoder and a video encoder. The input item of the video encoder is the feature representation corresponding to each video frame. The video encoder encodes the feature representation corresponding to each video frame to determine the global video representation corresponding to the video sequence to be recognized. The input of the text encoder is the text information corresponding to the video sequence to be recognized, and the text encoder encodes the text information to obtain a text representation. The text information can be text information of an inherent category, etc. Among them, the video encoder can encode the feature representation into a high-dimensional feature representation through Patch Embedding (Patch encoder), and then the high-dimensional feature representations corresponding to all video frames perform interactive learning through cross-frame interaction to capture spatio-temporal dependence relationships, and generate a global video representation based on the captured spatio-temporal dependence relationships.

[0048] Exemplarily, determining the global video representation based on the high-dimensional feature representation of each video frame specifically includes: Input the high-dimensional feature representation corresponding to each video frame into the cross-frame interaction module, and output the spatio-temporal feature representation corresponding to each video frame through the cross-frame interaction module; Input the spatio-temporal feature representation of each video frame into the spatio-temporal fusion module, and output the global video representation through the spatio-temporal fusion module.

[0049] Specifically, as Figure 2 shown, the cross-frame interaction module is used to perform interactive learning on the high-dimensional features corresponding to the video frames to capture spatio-temporal dependence relationships to obtain the spatio-temporal feature representation of each frame. Among them, the cross-frame interaction module can adopt a Transformer model to perform interactive learning on the high-dimensional feature representations of each video frame to obtain the spatio-temporal feature representation corresponding to each video frame. The spatio-temporal fusion module is used to fuse the spatio-temporal feature representations of each video frame to generate a global video representation. Among them, the spatio-temporal fusion module can also adopt a Transformer model to fuse the spatio-temporal feature representations of each video frame to generate a global video representation.

[0050] Furthermore, after obtaining the global video representation, a visual cue can be generated based on the global video representation and the text representation, and then the action label of the video sequence to be recognized can be determined based on the visual cue and the global video representation. Specifically, determining the action label of the video sequence to be recognized based on the text representation and the global video representation specifically includes: Input the text representation and the global video representation into the video-specific cue generator; Capture the dependence association between the text representation and the global video representation through the self-attention mechanism in the video-specific cue generator to form an intermediate text representation; Determine a video-specific prompt based on the intermediate text representation and the global video representation through a feed-forward network in the video-specific prompt generator; Fuse the video-specific prompt with the text representation to obtain an enhanced text representation; Calculate the similarity between the enhanced text representation and the global video representation, and determine the action label of the video sequence to be recognized based on the similarity.

[0051] Specifically, the video-specific prompt can be determined by a video-specific prompt generator, which is used to combine the text representation with relevant visual information in the video (such as the visual information "bookstore" in the video and the text "reading a book", etc.) to provide a more accurate semantic context during classification. Among them, the input items of the video-specific prompt generator include a text prompt and a global video prompt. The self-attention mechanism in the video-specific prompt generator captures the dependencies between the text representation and the global video representation to form a video-specific prompt, and then enhances the text representation based on the video-specific prompt to obtain an enhanced text representation. The following takes denoting the text representation and denoting the global video representation as an example to illustrate the working process of the video-specific prompt generator. The working process is specifically as follows: and First, input the text representation and the global video representation into the video-specific prompt generator. Use the text representation as the query vector, and the global video representation as the key vector and value vector. The self-attention mechanism in the video-specific prompt generator captures the dependencies between the text representation and and to generate an intermediate text representation Refine the global video representation through the feed-forward network in the video-specific prompt generator, and fuse the refined global video representation with the intermediate text representation to obtain a video-specific prompt and denoting the feed-forward network; finally, fuse the text representation with the video-specific prompt to obtain an enhanced text representation and and denoting the weighting coefficient.

[0052] Further, after obtaining the enhanced text representation, the similarity between the global video representation and the enhanced text representation can be calculated, and then the human body label corresponding to the video sequence to be recognized can be determined according to the similarity. Among them, the similarity between the global video representation and the enhanced text representation can adopt cosine similarity.

[0053] Of course, in practical applications, other methods can also be used to determine the behavior label. For example, after obtaining the global video features, the behavior label can be determined through a classification head, or the formation label can be determined through a regression operation, etc. There is no specific limitation here. The above implementation method is only an example of a specific implementation method.

[0054] It should be noted that the embodiments of the present application illustrate the specific process of the human behavior recognition method for low-quality videos in the form of a method flow. In practical applications, the recognition process of human behavior in low-quality videos can be carried out through a deep learning module (denoted as a human behavior recognition model). Specifically, as Figure 2 shown, the human behavior recognition model may include an inter-frame noise suppression module, a cross-frame semantic aggregation module, and a behavior recognition backbone network. The inter-frame noise suppression module is connected to the cross-frame semantic aggregation module, and the cross-frame semantic aggregation module is connected to the behavior recognition backbone network. The inter-frame noise suppression module is used to obtain the pre-frame difference map between each video frame in the video sequence to be recognized and its corresponding previous video frame, the post-frame difference map between it and its corresponding subsequent video frame, and the average frame difference map between it and the average frame of the video sequence to be recognized. The cross-frame semantic aggregation module is used to perform cross-frame semantic aggregation on the pre-frame difference map, post-frame difference map, and average frame difference map corresponding to each video frame to obtain the feature representation corresponding to each video frame. The behavior recognition backbone network is used to determine the behavior label of the video sequence to be recognized based on the feature representation corresponding to each video frame in the video sequence to be recognized. Among them, the specific implementation processes of the inter-frame noise suppression module, the cross-frame semantic aggregation module, and the behavior recognition backbone network are the same as those in the method flow embodiment above, and will not be specifically described here. Only the loss function used in the training process of the human behavior recognition model will be described here. The loss function used in the training process can be: , where, represents the loss function, represents the cosine similarity, represents the number of inherent lists in the text information, represents the enhanced text representation, represents the th text representation of the inherent category.

[0055] In summary, this embodiment provides a method for human behavior recognition in low-quality videos. The method includes obtaining a pre-frame difference map between each video frame in the video sequence to be recognized and its corresponding previous video frame, a post-frame difference map between the video frame and its corresponding subsequent video frame, and an average frame difference map between the video frame and the average frame of the video sequence to be recognized; performing cross-frame semantic aggregation on the pre-frame difference map, the post-frame difference map, and the average frame difference map corresponding to each video frame to obtain a feature representation corresponding to each video frame; and determining the behavior label of the video sequence to be recognized based on the feature representation corresponding to each video frame in the video sequence to be recognized. In this application, the pre-frame difference map, the post-frame difference map, and the average frame difference map are first obtained to suppress inter-frame noise, and then cross-frame semantic aggregation is performed based on the pre-frame difference map, the post-frame difference map, and the average frame difference map to aggregate rich spatio-temporal information. This can not only reduce background noise and interference while maintaining key contour information, but also obtain rich spatio-temporal information, effectively improving the accuracy of behavior recognition in encrypted videos. Especially for low-quality videos, it can eliminate interference, highlight the main body, and at the same time maintain efficient and accurate recognition of human behavior.

[0056] Based on the above method for human behavior recognition in low-quality videos, this embodiment provides a device for human behavior recognition in low-quality videos, as Figure 3 shown. The device for human behavior recognition in low-quality videos specifically includes: An inter-frame noise suppression module 100, configured to obtain a pre-frame difference map between each video frame in the video sequence to be recognized and its corresponding previous video frame, a post-frame difference map between the video frame and its corresponding subsequent video frame, and an average frame difference map between the video frame and the average frame of the video sequence to be recognized; A cross-frame semantic aggregation module 200, configured to perform cross-frame semantic aggregation on the pre-frame difference map, the post-frame difference map, and the average frame difference map corresponding to each video frame to obtain a feature representation corresponding to each video frame; A behavior recognition module 300, configured to determine the behavior label of the video sequence to be recognized based on the feature representation corresponding to each video frame in the video sequence to be recognized.

[0057] Based on the above method for human behavior recognition in low-quality videos, this embodiment provides a computer-readable storage medium. The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the method for human behavior recognition in low-quality videos as described in the above embodiment.

[0058] Based on the above method for human behavior recognition in low-quality videos, this application also provides a terminal device, as Figure 4As shown in the figure, it includes at least one processor 20; a display screen 21; and a memory 22. It may also include a communications interface 23 and a bus 24. Among them, the processor 20, the display screen 21, the memory 22, and the communications interface 23 can communicate with each other through the bus 24. The display screen 21 is set to display a user guidance interface preset in the initial setting mode. The communications interface 23 can transmit information. The processor 20 can call the logical instructions in the memory 22 to execute the method in the above-mentioned embodiments.

[0059] In addition, when the logical instructions in the above-mentioned memory 22 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.

[0060] The memory 22, as a computer-readable storage medium, can be set to store software programs and computer-executable programs, such as the program instructions or modules corresponding to the methods in the embodiments of the present disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions, or modules stored in the memory 22, that is, implements the methods in the above-mentioned embodiments.

[0061] The memory 22 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal device, etc. In addition, the memory 22 may include high-speed random access memory and may also include non-volatile memory. For example, various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, can also be transient storage media.

[0062] In addition, the specific processes of loading and executing multiple instructions by the above-mentioned storage medium and the instruction processor in the terminal device have been described in detail in the above method and will not be repeated here one by one.

[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A human behavior recognition method for low-quality videos, characterized in that: The human behavior recognition method for low-quality videos specifically includes: Obtain a preceding frame difference map between each video frame in the to-be-identified video sequence and its corresponding preceding frame, a succeeding frame difference map between its corresponding succeeding video frames, and an average frame difference map between the average frame of the to-be-identified video sequence; Perform cross-frame semantic aggregation on the previous frame difference map, the subsequent frame difference map and the average frame difference map corresponding to each video frame to obtain the feature representation corresponding to each video frame; Based on the feature representation corresponding to each video frame in the to-be-recognized video sequence, a behavior label of the to-be-recognized video sequence is determined.

2. The human behavior recognition method for low-quality videos according to claim 1, characterized in that: The step of obtaining a preceding frame difference map between each video frame in the to-be-identified video sequence and its corresponding preceding frame, a subsequent frame difference map between its corresponding subsequent video frames, and an average frame difference map between the average frames of the to-be-identified video sequence specifically comprises: Obtain an average frame of the video sequence to be identified and a preceding video frame and a succeeding video frame of each video frame; Frame difference operations are performed on each video frame and its preceding video frame, each video frame and the average frame to obtain a preceding frame difference map, a succeeding frame difference map and an average frame difference map corresponding to each video frame.

3. The human behavior recognition method for low-quality videos according to claim 2, characterized in that: The acquisition process of the preceding video frame and the succeeding video frame of the video frame specifically includes: Reading the subject ratio and the motion amplitude of the video frame, and determining the frame interval corresponding to the video frame based on the subject ratio and the motion amplitude; A preceding video frame and a succeeding video frame are selected for the video frame in the to-be-identified video sequence according to the frame interval.

4. The human behavior recognition method for low-quality videos according to claim 1, characterized in that: The cross-frame semantic aggregation of the preceding frame difference map, the succeeding frame difference map and the average frame difference map corresponding to each video frame to obtain the feature representation corresponding to each video frame specifically includes: Obtain the preceding weight of the preceding frame difference map corresponding to each video frame, the succeeding weight corresponding to the succeeding frame difference map, and the average weight corresponding to the average frame difference map; Based on the preceding weight, the succeeding weight and the average weight, the preceding frame difference map, the succeeding frame difference map and the average frame difference map corresponding to each video frame are weightedly combined to obtain a fused frame difference map for each video frame; Feature extraction is performed on the fused frame difference map of each video frame to obtain a feature representation of each video frame.

5. The human behavior recognition method for low-quality videos according to claim 1, characterized in that: The determining of the behavior label of the video sequence to be identified based on the corresponding feature representation of each video frame in the video sequence to be identified specifically includes: Acquire text information corresponding to the video sequence to be identified, and determine a text representation corresponding to the text information through a text encoder in a CLIP model; Determine a high-dimensional feature representation based on the feature representation corresponding to each video frame in the to-be-identified video sequence by a video encoder in the CLIP model, and determine a global video representation based on the high-dimensional feature representation of each video frame; Based on the text representation and the global video representation, a behavior label of the to-be-identified video sequence is determined.

6. The human behavior recognition method for low-quality videos according to claim 5, characterized in that: Determining the global video representation based on the high-dimensional feature representation of each video frame specifically includes: Input the high-dimensional feature representation corresponding to each video frame into the cross-frame interactive Transformer, and output the spatiotemporal feature representation corresponding to each video frame through the cross-frame interactive module; The spatiotemporal feature representation of each video frame is input into a spatiotemporal fusion module, and a global video representation is output through the spatiotemporal fusion module.

7. The human behavior recognition method for low-quality videos according to claim 5, characterized in that: The determining the behavior label of the to-be-identified video sequence based on the text representation and the global video representation specifically includes: inputting the textual representation and the global video representation into a video-specific cue generator; Capturing the dependency between the text representation and the global video representation through a self-attention mechanism in the video-specific cue generator to form an intermediate text representation; determining, by a feed-forward network in the video-specific cue generator, video-specific cues based on the intermediate text representation and the global video representation; fusing the video specific cue with the text representation to obtain an enhanced text representation; The similarity between the enhanced text representation and the global video representation is calculated, and the behavior label of the to-be-identified video sequence is determined based on the similarity.

8. A human behavior recognition device for low-quality videos, characterized in that: The human behavior recognition device for low-quality videos specifically includes: An inter-frame noise suppression module is used to obtain a preceding frame difference map between each video frame in the to-be-identified video sequence and its corresponding preceding frame, a subsequent frame difference map between its corresponding subsequent video frames, and an average frame difference map between the average frame of the to-be-identified video sequence; A cross-frame semantic aggregation module is used to perform cross-frame semantic aggregation on the preceding frame difference map, the succeeding frame difference map and the average frame difference map corresponding to each video frame to obtain a feature representation corresponding to each video frame; The behavior recognition module is used to determine the behavior label of the video sequence to be recognized based on the corresponding feature representation of each video frame in the video sequence to be recognized.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the human behavior recognition method for low-quality videos as described in any one of claims 1-7.

10. A terminal device, characterized in that: include: Processor and memory; The memory stores a computer-readable program executable by the processor; When the processor executes the computer-readable program, the steps in the method for human behavior recognition for low-quality videos as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Behavior recognition method, device and equipment

    CN116778568A