Video Spatiotemporal Action Detection Method Based on Peak Region Adaptive Diffusion
By adopting the adaptive diffusion method of peak area in space-time action detection, combined with the Gaussian nuclear heat map and the mean diffusion module of Gestalt principle, the problem of poor detection effect in the existing technology when dealing with fast deformation and bit movement is solved, and more efficient and accurate space-time action detection is achieved.
Patent Information
- Application Number
- CN202310138839.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-09
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2043-02-09
AI Technical Summary
When dealing with large movements, such as the rapid deformation and rapid displacement of moving targets, existing spatiotemporal action detection methods can easily lead to error accumulation during model optimization, resulting in poor spatial and temporal action detection results.
The video space-time action detection method based on adaptive diffusion of the peak area is adopted. By constructing a top-down Gaussian nuclear heat map peak area mining module, the target motion trajectory is depicted, and the mean diffusion positioning module based on the Gestalt principle is designed to adjust the target bounding box size to adapt to the change of the action amplitude.
This method can quickly adapt to action changes, reduce model calculation overhead, improve the accuracy and stability of space-time action detection, and is suitable for space-time action detection in complex scenarios.
Smart Images

Figure CN116109984B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, especially the action localization technology field in video processing, and relates to a spatio-temporal action detection method based on adaptive diffusion of peak regions. Background Art
[0002] Video content on the Internet is diverse, mixed, and growing exponentially in scale, and the security of video content has been increasingly emphasized. Among them, the targets and their associated actions in videos are the key to content review, while manual review is time-consuming and laborious and may cause misjudgment. Therefore, how to quickly and accurately detect actions and their associated targets in videos, that is, spatio-temporal action detection, has become an important research topic. This task is oriented to fine spatio-temporal marking, takes untrimmed videos containing multiple targets and multiple actions as input, and outputs the start and end times of all actions in the video, the spatial positions of the targets associated with the actions, and the corresponding action categories. Spatio-temporal action detection has broad application prospects in practical scenarios such as intelligent monitoring in parks, intelligent transportation guarantee, and warning of dangerous behaviors. For example, for an intelligent transportation system, the spatio-temporal action detection method can use sky cameras to monitor in real time illegal behaviors occurring on the road, such as cars driving in reverse and pedestrians running red lights, and give early warnings in time to reduce the traffic accident rate; in addition, it can also be applied to sports event scenarios to detect illegal video segments, such as malicious injury and illegal line crossing, to maintain and improve the fairness of the event.
[0003] Spatio-temporal action detection methods are mainly divided into two ways: single-frame input (Frame-level) and multi-frame input (Tubelet-level). Among them, Tubelet-level mainly solves the problem of difficult mining of temporal relationships in the Frame-level method. Tubelet-level spatio-temporal action detection mainly adopts a two-stage paradigm, that is, the description of the motion trajectory is divided into a coarse-grained stage and a fine-grained stage; in the coarse-grained stage, the target bounding box proposal of a given key frame is extended into a 3D temporal target bounding box, that is, having the same initial position and shape in the spatial dimension of different frames, and then it is input into an action category discriminator to obtain the action category; in the fine-grained stage, a frame-level detector is used to linearly correct the target bounding boxes whose intersection over union does not reach the preset threshold to describe the target motion trajectory. However, when dealing with large actions, such as rapid deformation (diving) and rapid displacement (running) of moving targets, due to the deviation of action features, it is easy to cause error accumulation during model optimization, resulting in poor spatio-temporal action detection effects.
[0004] The deficiencies of the above spatio-temporal action detection method are mainly manifested in two aspects: (1) Although the multi-frame input method can well mine the temporal relationship between targets, the direct expansion method cannot quickly cope with camera jitter and rapid target displacement; (2) For the action categories discriminated based on the 3D temporal object bounding box, due to the fixed same spatial position, and the target is prone to deformation during movement, resulting in a deviation between the target features and the true features, causing model detection errors. Therefore, aiming at the problems of target localization deviation and incorrect discrimination of action category in the model results caused by the direct expansion method, there is an urgent need to design a spatio-temporal action detection method that can depict the target motion trajectory and adjust the size of the target bounding box according to the target appearance. Summary of the Invention
[0005] In view of the deficiencies of the existing methods, the present invention proposes a video spatio-temporal action detection method based on adaptive diffusion in the peak region. The method of the present invention constructs a top-down Gaussian kernel heat map peak region mining module to mine the region of interest features to depict the target motion trajectory to cope with camera jitter and rapid displacement of moving targets; at the same time, a mean diffusion localization module based on the Gestalt principle is designed to adjust the size of the target bounding box associated with the action to cope with the problem of target deformation caused by drastic changes in the action amplitude.
[0006] The method of the present invention performs the following operations on the video data set with given action categories and spatio-temporal markings of actions in sequence:
[0007] Step (1) Preprocess the video to obtain a video frame sequence, and use two-dimensional and three-dimensional convolutional neural networks and faster region convolutional neural networks to extract the initial target bounding box tuple, video frame features, and video spatio-temporal feature maps;
[0008] Step (2) Construct a peak region mining module, with the input being the initial target bounding box tuple and video frame features, and the output being the peak region and its central position coordinates;
[0009] Step (3) Establish a Gestalt mean diffusion module, with the input being the original video frame sequence and the central position coordinates of the peak region, and the output being the target bounding box tuple of all targets at the current moment;
[0010] Step (4) Construct a channel pooling module, with the input being the video spatio-temporal feature map and the target bounding box tuple, and the output being the targets associated with the action and the action categories at the current moment;
[0011] Step (5) Use the stochastic gradient descent algorithm to optimize the spatio-temporal action detection model composed of the peak region mining module, the Gestalt mean diffusion module, and the channel pooling module, and sequentially execute steps (1) to (4) on the new video sequence to obtain the target bounding boxes and action categories of all targets associated with the action at different moments.
[0012] Furthermore, step (1) is specifically as follows:
[0013] (1-1) Sample the video at a sampling rate of N frames per second, where 5 ≤ N ≤ 10, to obtain a set of frame sequences containing T' frames represents the real number field, U s represents the frame sequence of the s-th frame, and H', W', and 3 represent the height, width, and RGB three channels of the video frame respectively;
[0014] (1-2) Divide the video frame sequence into T video segments The length of a single video segment is 2·N frames, V t represents the t-th video segment, and then input V t into a three-dimensional convolutional neural network to generate the spatio-temporal feature map of the t-th video segment H, W, and 2·N are the height, width, and number of channels of the feature map respectively, and thus obtain the spatio-temporal feature maps of all video segments;
[0015] (1-3) Use a faster region convolutional neural network to perform object detection on the middle frame of the video segment to obtain a set of initial object bounding box tuples The middle frame is the N-th frame of the video segment; N t,N represents the number of objects existing in the middle frame of the video segment V t ; represents the bounding box of the i-th object in the middle frame of the video segment V t ; respectively represent the abscissa and ordinate of the upper left corner of the bounding box of the i-th object in the middle frame of the video segment V t ; respectively represent the abscissa and ordinate of the lower right corner of the bounding box of the i-th object in the middle frame of the video segment V t ; Input the video frames of the video segment V t into a two-dimensional convolutional neural network to obtain video frame features C is the number of channels, 1 < n < 2·N.
[0016] Furthermore, step (2) is specifically as follows:
[0017] (2-1) Construct a peak region mining module to obtain the center position coordinates and sizes of the bounding boxes of all objects, and the center position coordinates of the bounding box of the i-th object The size of the bounding box of the i-th object
[0018] According to calculate the Gaussian kernel variance value to adjust the Gaussian kernel size, σ 0is the preset variance, 0 < σ 0 < 1, calculate the Gaussian value relative to the i-th target at the coordinate (x, y) Similarly, obtain the Gaussian heatmap distribution of target i and the Gaussian heatmap distributions of other targets, through obtain the Gaussian heatmap distribution matrix of the N-th frame of the t-th video segment;
[0019] (2-2) Obtain the peak region features ⊙ represents the element-wise multiplication operation, maxpool(·) represents the max pooling operation, the parameter max(·, ·) represents taking the maximum value; then calculate the cosine similarity score = cossim(F t,N,peak ·F t,N+1,can ) extracts the region features by means of a sliding window; select the top-k regions with the highest similarity and score > δ 0 the preset threshold 0 < δ 0 < 1, select the intersection of the top-k regions as the peak region tuple of the current frame respectively represent the abscissa and ordinate of the upper left corner of the i-th peak region of the (N + 1)-th frame of the video segment V t , respectively represent the abscissa and ordinate of the lower right corner of the i-th target peak region of the (N + 1)-th frame of the video segment V t , and calculate the center position coordinates of the peak region of the current frame accordingly Thus, obtain the peak regions and their center position coordinates of all frames of the current segment;
[0020] (2-3) Use the center position of the real result target bounding box and the center position of the peak region to calculate the localization offset loss where represents the real target bounding box of the i-th target of the n-th frame of the video segment V t , ||·|| 1 represents the l 1 norm, N t,n represents the number of targets in the n-th frame of the video segment V t .
[0021] Furthermore, step (3) is specifically:
[0022] (3-1) Construct a Gestalt mean diffusion module consisting of a target tracking sub-module and a spatial gradient sub-module. The target tracking sub-module uses color probability distribution to make a coarse-grained discrimination of target localization, and the spatial gradient sub-module is used to extract texture features to refine target localization. Map the original video frame sequence to the HSV color space and arrange it into a vector in the form of [Hue, Saturation, Value], where 0 < Hue < 180, 0 < Saturation < 255, 0 < Value < 255, Hue represents hue, Saturation represents saturation, and Value represents brightness; use histogram backprojection, that is, replace all pixel values of the video frame with the probability of their color appearance to generate a color probability distribution map, and obtain the probability of target pixels m j represents the number of pixel points of color j, and thus obtain the color distribution matrix Thus obtain the color histogram probability distribution matrix within the target bounding box;
[0023] (3-2) Convert the original video frame into a grayscale image, and the grayscale value of the point (x, y) R(·, ·) represents the R-channel component, G(·, ·) represents the G-channel component, B(·, ·) represents the B-channel component, 0 < x < H′, 0 < y < W′;
[0024] Use a preset spatial gradient operator Obtain the current frame texture matrix According to the obtained peak region tuple and the center position coordinates of the peak region Initialize the size of the target bounding box
[0025] (3-3) Divide the probability distribution matrix within the target bounding box into grids in the way of length and width Aggregate the grids within the closed loop according to the texture feature matrix At the same time, measure the distance between the grid distributions within different closed loops through , where the numerator represents the minimum Manhattan distance between two grids, and the denominator represents the Manhattan distance between the center coordinates of two grids. left, right, up, down, and cen respectively represent the leftmost, rightmost, topmost, bottommost, and center positions, x and y respectively represent the abscissa and ordinate, a and b respectively represent grid a and grid b. Aggregate adjacent grids with a distance less than the preset threshold δ 1 0 < δ 1 < 1; Measure the color histogram probability distribution matrix within the target bounding box at adjacent times, If it is less than the preset threshold δ 2Output the target bounding box tuples of all targets at the current moment 0 < δ 2 < 1, otherwise, preset threshold δ 1 Decay to half of the original value and re - execute step (3 - 3);
[0026] (3 - 4) Calculate the distance intersection - over - union loss function of the model Calculate the intersection - over - union of the predicted target bounding box and the true target bounding box is the true target bounding box, is the upper - left coordinate of the target bounding box, is the representation of the lower - right coordinate of the target bounding box, represents the upper - left coordinate of the smallest bounding box that can enclose both the true bounding box and the predicted bounding box, represents the lower - right coordinate of the smallest bounding box that can enclose both the true bounding box and the predicted bounding box, min(·,·) represents taking the minimum value.
[0027] Furthermore, step (4) is specifically:
[0028] (4 - 1) Construct a channel pooling module composed of spatial max - pooling and temporal max - pooling, and encode the target features using bilinear interpolation operation on the video frame feature maps at different moments based on the target bounding box tuples And perform channel concatenation, and obtain the target context features through a two - dimensional convolutional layer Conv2D 3 (·) represents a two - dimensional convolutional layer with input channels C′ = 2·N·C, output channels C, and convolutional kernel size 1×1×C', concat(·,·) represents channel concatenation;
[0029] (4 - 2) Concatenate the target context features and the spatio - temporal feature map of the video segment along the channel dimension, input the concatenated features into a two - dimensional convolutional layer and then perform a spatial global pooling operation to obtain the target classification score GAP(·) represents global average pooling in the spatial dimension;
[0030] (4 - 3) Use the Softmax function to process the target classification score to obtain the output probability that the i - th target in video segment V t belongs to the action category u as M t is the number of targets in video segment V t ; Calculate the cross - entropy loss function where is the true label, represents video segment V tThe i-th target contains an action with an action category of u;
[0031] (4-4) The intersection over union metric Stitch together video segments with the same action category to obtain a complete action instance.
[0032] Furthermore, step (5) is specifically:
[0033] (5-1) Construct a spatio-temporal action detection model composed of a peak region mining module, a Gestalt mean diffusion module, and a channel pooling module; use the stochastic gradient descent algorithm to optimize the above spatio-temporal action model, detect and iteratively train the model until convergence to obtain an optimized spatio-temporal action detection model;
[0034] (5-2) For a new video, obtain a video frame sequence and video segments where T″ is the number of video frames, and input the above optimized spatio-temporal action detection model, and sequentially execute according to steps (1) to (4) to output the current spatial positions of all action-related targets in the video segment at the current moment respectively represent the abscissa and ordinate of the upper left corner of the bounding box of the i'-th target in the n'-th frame of the t'-th video segment, respectively represent the abscissa and ordinate of the lower right corner of the bounding box of the i'-th target in the n'-th frame of the t'-th video segment, represents the probability value that the i'-th target in the t'-th video segment belongs to the action category u. When the action probability result exceeds the preset threshold δ 3 it is classified as the action category u, 0.5 ≤ δ 3 ≤ 1.
[0035] The present invention proposes a spatio-temporal action detection method for videos based on peak region adaptive diffusion. This method has the following characteristics: (1) The designed top-down Gaussian kernel adjustment algorithm can adaptively change the size of the local feature region for mining targets according to the target size to describe the motion trajectory of the target to adapt to the rapid movement of the target; (2) The designed Gestalt principle mean diffusion algorithm uses the most stable color attribute features to generate a pixel probability distribution map and integrates regions with obvious and certain similarities in appearance such as shape and direction to adapt to the rapid deformation of the target.
[0036] The method of the present invention is applicable to spatio-temporal action detection in complex scenarios, such as target occlusion, large movements, etc. The beneficial effects include: (1) By adjusting the Gaussian kernel size, the local feature region is adaptively adjusted, reducing the computational overhead of the model to quickly adapt to the situation where actions change rapidly; (2) The color histogram is not affected by changes in the shape, posture, etc. of the target. Therefore, matching based on the color distribution has the characteristics of good stability, resistance to partial occlusion, simple calculation method, and small computational amount.
[0037] The peak region mining module and the Gestalt mean diffusion module of the present invention can well ensure the applicability of the model and can be applied to fields such as traffic safety detection and illegal content identification. Description of the Drawings
[0038] Figure 1 is a flowchart of the method of the present invention. Detailed Embodiments
[0039] The present invention will be further described below with reference to the accompanying drawings.
[0040] Such as Figure 1 , a video spatio-temporal action detection method based on peak region adaptive diffusion. The method first uniformly samples the original video, and uses a detector and a convolutional neural network to extract the target bounding box tuple, video frame features, and video spatio-temporal feature map; uses the constructed peak region mining module to obtain the center position coordinates of the peak region; then inputs the original video frame, the peak region, and its center position coordinates into the Gestalt mean diffusion module to obtain the corrected target bounding box tuple of all targets; finally, uses the channel pooling module to obtain the spatial positions and action categories of all targets at different times. The method adopts a local strategy, adaptively adjusts the peak region according to the initial target bounding box size to quickly capture the target motion trajectory; also corrects the target spatio-temporal action pipeline through the Gestalt principle mean diffusion module to learn features that can better depict the target motion.
[0041] The method performs the following operations in sequence on the video data set with given action categories and action spatio-temporal tags:
[0042] Step (1) Preprocess the video to obtain a video frame sequence, and use two-dimensional and three-dimensional convolutional neural networks and a faster region convolutional neural network to extract the initial target bounding box tuple, video frame features, and video spatio-temporal feature map; specifically:
[0043] (1-1) Sample the video at a sampling rate of N frames per second, 5 ≤ N ≤ 10, to obtain a frame sequence set containing T′ frames represents the real number field, U s represents the frame sequence of the s-th frame, and H′, W′, and 3 respectively represent the height, width, and RGB three channels of the video frame;
[0044] (1 - 2) Divide the video frame sequence into T video segments The length of a single video segment is 2·N frames, V t denotes the t-th video segment, and then input V t into a three-dimensional convolutional neural network to generate the spatio-temporal feature map of the t-th video segment H, W, and 2·N are the height, width, and number of channels of the feature map respectively, thus obtaining the spatio-temporal feature maps of all video segments;
[0045] (1 - 3) Use a faster region convolutional neural network to perform object detection on the middle frame of the video segment to obtain a set of initial object bounding box tuples The middle frame is the N-th frame of the video segment; N t,N denotes the number of objects existing in the middle frame of the video segment V t denotes the bounding box of the i-th object in the middle frame of the video segment V t respectively denote the abscissa and ordinate of the upper left corner of the bounding box of the i-th object in the middle frame of the video segment V t respectively denote the abscissa and ordinate of the lower right corner of the bounding box of the i-th object in the middle frame of the video segment V; input the video frames of the video segment V t into a two-dimensional convolutional neural network to obtain video frame features t C is the number of channels, 1 < n < 2·N.
[0046] Step (2) Construct a peak region mining module, with the input being the initial object bounding box tuples and video frame features, and the output being the peak region and its central position coordinates; specifically:
[0047] (2 - 1) Construct a peak region mining module to obtain the central position coordinates and sizes of the bounding boxes of all objects, and the central position coordinates of the bounding box of the i-th object The size of the bounding box of the i-th object
[0048] According to calculate the Gaussian kernel variance value to adjust the Gaussian kernel size, with the preset variance 0 < σ 0 < 1, and in this embodiment, σ 0 = 0.2, calculate the Gaussian value relative to the i-th object at the coordinate (x, y) Similarly, obtain the Gaussian heat map distribution of object i and the Gaussian heat map distributions of other objects, through Obtain the Gaussian heatmap distribution matrix of the Nth frame of the tth video segment;
[0049] (2-2) Obtain the peak region features ⊙ represents the element-wise multiplication operation, maxpool(·) represents the max pooling operation, and the parameter max(·,·) represents taking the maximum value; then calculate the cosine similarity score = cossim(F t,N,peak ·F t,N+1,can ) extracts region features by means of a sliding window; select the top-k regions with the highest similarity and score > δ 0 The preset threshold is 0 < δ 0 < 1, and in this embodiment, δ 0 = 0.5; select the intersection of the top-k regions as the peak region tuple of the current frame respectively represent the abscissa and ordinate of the upper left corner of the ith peak region of the (N+1)th frame of the video segment V t , respectively represent the abscissa and ordinate of the lower right corner of the ith target peak region of the (N+1)th frame of the video segment V t , and calculate the center position coordinates of the peak region of the current frame accordingly Thus, obtain the peak regions and their center position coordinates of all frames of the current segment;
[0050] (2-3) Use the center position of the true result target bounding box and the center position of the peak region to calculate the localization offset loss where represents the true target bounding box of the ith target in the nth frame of the video segment V t , ||·|| 1 represents the l 1 norm, and N t,n represents the number of targets in the nth frame of the video segment V t .
[0051] Step (3) Establish a Gestalt mean diffusion module, with the input being the original video frame sequence and the center position coordinates of the peak regions, and the output being the target bounding box tuple of all targets at the current moment; specifically:
[0052] (3-1) Construct a Gestalt mean diffusion module composed of a target tracking sub-module and a spatial gradient sub-module. The target tracking sub-module uses color probability distribution to make a coarse-grained discrimination of target localization, and the spatial gradient sub-module is used to extract texture features to refine target localization. Map the original video frame sequence to the HSV color space and arrange it into a vector in the form of [Hue, Saturation, Value], where 0 < Hue < 180, 0 < Saturation < 255, 0 < Value < 255, Hue represents hue, Saturation represents saturation, and Value represents brightness; use histogram backprojection, that is, replace all pixel values of the video frame with the probability of their color appearance to generate a color probability distribution map and obtain the probability of target pixels m j represents the number of pixel points of color j, and thus obtain a color distribution matrix Thus, obtain the color histogram probability distribution matrix within the target bounding box;
[0053] (3-2) Convert the original video frame into a grayscale image, and the grayscale value of the point (x, y) R(·, ·) represents the R-channel component, G(·, ·) represents the G-channel component, B(·, ·) represents the B-channel component, 0 < x < H′, 0 < y < W′;
[0054] Use a preset spatial gradient operator Obtain the texture matrix of the current frame According to the obtained peak region tuple and the center position coordinates of the peak region Initialize the size of the target bounding box
[0055] (3-3) Divide the probability distribution matrix within the target bounding box into grids in the manner of length and width Aggregate the grids within the closed loop according to the texture feature matrix At the same time, measure the distance between the grid distributions within different closed loops through where the numerator represents the minimum Manhattan distance between two grids, and the denominator represents the Manhattan distance between the center coordinates of the two grids. left, right, up, down, and cen represent the leftmost, rightmost, topmost, bottommost, and center positions respectively, x and y represent the abscissa and ordinate respectively, a and b represent grids a and b, and aggregate adjacent grids smaller than the preset threshold δ 1 0 < δ 1 < 1, and in this embodiment, δ 1 = 0.5; then measure the color histogram probability distribution matrix within the target bounding box at adjacent times, If it is less than the preset threshold δ 2 then output the target bounding box tuples of all targets at the current moment 0 < δ 2 < 1, in this embodiment, δ 2 = 0.5, otherwise, decay the preset threshold δ 1 to half of its original value and re - execute step (3 - 3);
[0056] (3 - 4) Calculate the distance intersection - over - union loss function of the model The intersection - over - union of the predicted target bounding box and the true target bounding box is the true target bounding box, is the upper - left coordinate of the target bounding box, is the representation of the lower - right coordinate of the target bounding box, represents the upper - left coordinate of the smallest bounding box that can enclose both the true bounding box and the predicted bounding box, represents the lower - right coordinate of the smallest bounding box that can enclose both the true bounding box and the predicted bounding box, and min(·, ·) represents taking the minimum value.
[0057] Step (4) Construct a channel pooling module. The input is the video spatio - temporal feature map and the target bounding box tuples, and the output is the targets associated with the action and the action categories at the current moment; specifically:
[0058] (4 - 1) Construct a channel pooling module composed of spatial max - pooling and temporal max - pooling. Based on the target bounding box tuples, use bilinear interpolation operation to encode the target features for the video frame feature maps at different moments and perform channel concatenation to obtain the target context features through a two - dimensional convolutional layer Conv2D 3 (·) represents a two - dimensional convolutional layer with an input channel of C′ = 2·N·C, an output channel of C, and a convolutional kernel size of 1×1×C', and concat(·, ·) represents channel concatenation;
[0059] (4 - 2) Concatenate the target context features and the video clip spatio - temporal feature map along the channel dimension, input the concatenated features into a two - dimensional convolutional layer and then perform a spatial global pooling operation to obtain the target classification scores GAP(·) represents global average pooling in the spatial dimension;
[0060] (4 - 3) Use the Softmax function to process the target classification scores to obtain the output probability that the i - th target in the video clip V t belongs to the action category u as M t is the video clip Vt The number of targets in; calculate the cross-entropy loss function wherein is the true label, represents the video segment V t The i-th target contains an action with the action category u;
[0061] (4-4) Intersection over Union metric Stitch together the video segments with the same action category to obtain a complete action instance.
[0062] Step (5) Use the stochastic gradient descent algorithm to optimize the spatio-temporal action detection model composed of the peak region mining module, the Gestalt mean diffusion module, and the channel pooling module, and sequentially execute steps (1) to (4) on the new video sequence to obtain the target bounding boxes and action categories of all action-associated targets at different times; specifically:
[0063] (5-1) Construct a spatio-temporal action detection model composed of the peak region mining module, the Gestalt mean diffusion module, and the channel pooling module; use the stochastic gradient descent algorithm to optimize the above spatio-temporal action model, detect and iteratively train the model until convergence to obtain an optimized spatio-temporal action detection model;
[0064] (5-2) For the new video, obtain a video frame sequence through sampling and video segments where T″ is the number of video frames, and input the above optimized spatio-temporal action detection model, and sequentially execute according to steps (1) to (4) to output the spatial positions of all action-associated targets at the current time of the video segment and their current segment action categories respectively represent the abscissa and ordinate of the upper left corner of the i'-th target bounding box in the n'-th frame of the t'-th video segment, respectively represent the abscissa and ordinate of the lower right corner of the i'-th target bounding box in the n'-th frame of the t'-th video segment, represents the probability value that the i'-th target in the t'-th video segment belongs to the action category u. When the action probability result exceeds the preset threshold δ 3 then, it is judged as the action category u, 0.5 ≤ δ 3 ≤ 1. In this embodiment, δ 3 = 0.6.
[0065] The content described in this embodiment is only a list of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms stated in the embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art according to the inventive concept of the present invention.
Claims
1. A video spatio-temporal action detection method based on peak region adaptive diffusion, characterized in that, the following operations are sequentially performed on a video data set with a given action category and spatio-temporal action markers: Step (1) Preprocess the video to obtain a video frame sequence, and use two-dimensional and three-dimensional convolutional neural networks and faster region convolutional neural networks to extract initial target bounding box tuples, video frame features, and video spatio-temporal feature maps; specifically: (1-1) Sample the video at a sampling rate of N frames per second, where 5 ≤ N ≤ 10, to obtain a set of frame sequences containing T' frames Denote the real number field as U s Denote the frame sequence of the s-th frame, where H', W', and 3 represent the height, width, and RGB three channels of the video frame respectively; (1-2) Divide the video frame sequence into T video segments The length of a single video segment is 2·N frames, V t represents the t-th video segment, and then input V t into the 3D convolutional neural network to generate the spatio-temporal feature map of the t-th video segment H, W, and 2·N are the height, width, and number of channels of the feature map respectively, and thus obtain the spatio-temporal feature maps of all video segments; (1-3) Use a faster regional convolutional neural network to perform object detection on the middle frame of a video clip to obtain an initial set of object bounding box tuples i = 1, 2, ..., N t,N , where the middle frame is the Nth frame of the video clip; N t,N represents the number of objects existing in the middle frame of video clip V t represents the bounding box of the ith object in the middle frame of video clip V t respectively represent the abscissa and ordinate of the upper left corner of the bounding box of the ith object in the middle frame of video clip V t respectively represent the abscissa and ordinate of the lower right corner of the bounding box of the ith object in the middle frame of video clip V; Input the video frames of video clip V t into a two-dimensional convolutional neural network to obtain video frame features t C is the number of channels, 1 < n < 2·N; Step (2) Construct a peak region mining module, with the input being the initial target bounding box tuples and video frame features, and the output being the peak region and its central position coordinates; specifically: (2-1)Construct a peak region mining module to obtain the central position coordinates and sizes of the target bounding boxes of all targets. The central position coordinates of the i-th target bounding box The size of the target bounding box of the i-th target According to Calculate the Gaussian kernel variance value to adjust the Gaussian kernel size, σ 0 Is a preset variance, 0 < σ 0 < 1, calculate the Gaussian value relative to the i-th target at the coordinates (x, y) Obtain the Gaussian heat map distribution of target i And the Gaussian heat map distributions of other targets, through Obtain the Gaussian heat map distribution matrix of the N-th frame of the t-th video segment; (2-2) Obtain peak region features ⊙ represents the element-wise multiplication operation, maxpool(·) represents the max pooling operation, and the parameter max(·,·) represents taking the maximum value; then calculate the cosine similarity score = cossim(F t,N,peak ·F t,N+1,can ) extracts region features by means of a sliding window; selects the top-k regions with the highest similarity and score > δ 0 where the preset threshold 0 < δ 0 < 1, and selects the intersection of the top-k regions as the peak region tuple of the current frame respectively represent the abscissa and ordinate of the upper left corner of the i-th peak region in the (N + 1)-th frame of the video clip V t ; respectively represent the abscissa and ordinate of the lower right corner of the i-th target peak region in the (N + 1)-th frame of the video clip V t , and calculate the center position coordinates of the peak region of the current frame Thus, the peak regions and their center position coordinates of all frames in the current clip are obtained; (2-3) Calculate the localization offset loss using the center position of the ground truth target bounding box and the center position of the peak region where represents the video clip V t the ground truth target bounding box of the i-th target in the n-th frame, ||·|| 1 represents l 1 norm, N t,n represents the video clip V t the number of targets in the n-th frame; Step (3) Establish a Gestalt mean diffusion module, with the input being the original video frame sequence and the central position coordinates of the peak region, and the output being the target bounding box tuples of all targets at the current moment; Step (4) Construct a channel pooling module, with the input being the video spatio-temporal feature map and the target bounding box tuples, and the output being the targets associated with the action and the action category at the current moment; Step (5) Use the stochastic gradient descent algorithm to optimize the spatio-temporal action detection model composed of the peak region mining module, the Gestalt mean diffusion module, and the channel pooling module, and sequentially execute Steps (1) to (4) on a new video sequence to obtain the target bounding boxes and action categories of all action-associated targets at different times.
2. The video spatio-temporal action detection method based on peak region adaptive diffusion according to claim 1, characterized in that, Step (3) is specifically: (3-1) Construct a Gestalt mean diffusion module composed of a target tracking sub-module and a spatial gradient sub-module. The target tracking sub-module uses color probability distribution to make a coarse-grained discrimination of target localization, and the spatial gradient sub-module is used to extract texture features to refine target localization; map the original video frame sequence to the HSV color space, arrange it into a vector in the form of [Hue, Saturation, Value], where 0 < Hue < 180, 0 < Saturation < 255, 0 < Value < 255, Hue represents hue, Saturation represents saturation, and Value represents brightness; use histogram backprojection, that is, replace all pixel values of the video frame with the probability of their color appearance to generate a color probability distribution map, and obtain the probability of target pixels. m j represents the number of pixel points of color j, and obtain the color distribution matrix Thus, obtain the color histogram probability distribution matrix within the target bounding box. (3-2) Convert the original video frame into a grayscale image, and the grayscale value of the point (x,y) R(·,·) represents the R-channel component, G(·,·) represents the G-channel component, B(·,·) represents the B-channel component, 0 < x < H′, 0 < y < W′; use a preset spatial gradient operator Obtain the texture matrix of the current frame According to the obtained peak region tuple And the center position coordinates of the peak region Initialize the size of the target bounding box (3-3) Perform grid segmentation on the probability distribution matrix within the target bounding box in the manner of length and width being , and aggregate the grids within the closed loop according to the texture feature matrix ; Measure the distance between the grid distributions in different closed loops through . The numerator represents the minimum Manhattan distance between two grids, and the denominator represents the Manhattan distance between the center coordinates of the two grids. left, right, up, down, and cen respectively represent the leftmost, rightmost, topmost, bottommost, and center positions. x and y respectively represent the abscissa and ordinate. a and b respectively represent grid a and grid b. Aggregate adjacent grids with a value less than the preset threshold δ 1 , where 0 < δ 1 < 1; Measure the probability distribution matrix of the color histogram within the target bounding box at adjacent time instants. If it is less than the preset threshold δ 2 , then output the target bounding box tuples of all targets at the current time instant , where 0 < δ 2 < 1. Otherwise, decay the preset threshold δ 1 to half of its original value and re-execute step (3-3); (3-4) Distance IoU Loss Function of the Calculation Model Intersection over Union of the Predicted Target Bounding Box and the Ground Truth Target Bounding Box is the ground truth target bounding box, is the top-left coordinate of the target bounding box, is the representation of the bottom-right coordinate of the target bounding box, represents the top-left coordinate of the smallest bounding box that can enclose both the ground truth bounding box and the predicted bounding box, represents the bottom-right coordinate of the smallest bounding box that can enclose both the ground truth bounding box and the predicted bounding box, and min(·,·) represents taking the minimum value.
3. The video spatio-temporal action detection method based on peak region adaptive diffusion according to claim 2, characterized in that, Step (4) is specifically: (4-1) Construct a channel pooling module consisting of spatial max pooling and temporal max pooling, and use bilinear interpolation operation to encode the target features for the video frame feature maps at different times based on the target bounding box tuple And perform channel concatenation to obtain the target context features through a two-dimensional convolutional layer Conv2D 3 (·) represents a two-dimensional convolutional layer with an input channel of C′ = 2·N·C, an output channel of C, and a convolutional kernel size of 1×1×C'. concat(·,·) represents channel concatenation; (4-2) Target context features and the spatio-temporal feature map of the video clip are concatenated along the channel dimension, and the concatenated features are input into a two-dimensional convolutional layer and then a spatial global pooling operation is performed to obtain the target classification score GAP(·) represents global average pooling in the spatial dimension; (4-3) Process the target classification scores using the Softmax function to obtain the video segment V t The output probability that the i-th target belongs to the action category u is M t is the number of targets in the video segment V t ; calculate the cross-entropy loss function where is the true label indicating that the i-th target in the video segment V t contains the action with the action category u (4-4) Intersection over Union metric Splice video segments that are consistent with the action category to obtain a complete action instance.
4. The video spatio-temporal action detection method based on peak region adaptive diffusion according to claim 3, characterized in that, Step (5) is specifically: (5-1) Construct a spatio-temporal action detection model composed of a peak region mining module, a Gestalt mean diffusion module, and a channel pooling module; use the stochastic gradient descent algorithm to optimize the above spatio-temporal action detection model, detect and iteratively train the model until convergence, and obtain an optimized spatio-temporal action detection model; (5-2) Sample the new video to obtain a video frame sequence and video clips where T″ is the number of video frames, and input the above-optimized spatio-temporal action detection model, and execute it sequentially according to steps (1) to (4) to output the current spatial positions of all action-related targets in the video clip and their current clip action categories 1 < n′ < 2·N, respectively represent the abscissa and ordinate of the upper left corner of the bounding box of the i′-th target in the n′-th frame of the t′-th video clip, respectively represent the abscissa and ordinate of the lower right corner of the bounding box of the i′-th target in the n′-th frame of the t′-th video clip, represents the probability value that the i′-th target in the t′-th video clip belongs to the action category u. When the action probability result exceeds the preset threshold δ 3 then, it is judged as the action category u, 0.5 ≤ δ 3 ≤ 1.
Citation Information
Patent Citations
Vehicle-mounted image de-noising method and system based on adaptive diffusion filtering
CN106651804A
End-to-end video action detecting and positioning system
CN113158723A