A target identification method based on monitoring video

By optimizing inter-frame difference and secondary clustering to extract keyframes, combining scale-invariant feature transformation and random sampling consensus algorithm for video stitching and fusion, and using dual-layer routing attention and deformable convolution for target recognition, the problem that a single camera cannot fully capture the target motion process is solved, and cross-view collaborative monitoring field of view expansion and target recognition accuracy are achieved.

CN121053606BActive Publication Date: 2026-04-17WUHAN CITY VOCATIONAL COLLEGE +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUHAN CITY VOCATIONAL COLLEGE
Filing Date
2025-09-04
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing technologies, the fixed perspective of a single camera cannot fully capture the movement and overall view of a target within the area covered by multiple cameras, resulting in inaccurate target recognition.

Method used

Keyframes are extracted by optimizing the inter-frame difference method and the secondary clustering method. Video stitching and image fusion are performed by combining scale-invariant feature transformation and random sampling consensus algorithm. Target recognition is performed by using two-layer routing attention and deformable convolution. Finally, the target detection result is determined by the adaptive weighted fusion method.

Benefits of technology

It enables cross-view collaborative monitoring field of view expansion, improves the reliability and accuracy of target identification, reduces the false negative rate and the risk of misjudgment, dynamically adapts to changes in target posture and lighting anomalies, and ensures a high level of reliability in target category and location judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053606B_ABST
    Figure CN121053606B_ABST
Patent Text Reader

Abstract

The application discloses a target identification method based on monitoring video and relates to the technical field of target identification. The method comprises the following steps: acquiring a plurality of regional monitoring videos, extracting a preliminary monitoring image key frame sequence by optimizing an interframe difference method, and screening an optimized monitoring image key frame sequence by a secondary clustering method; adopting a scale invariant feature transformation algorithm to perform feature matching on the optimized monitoring image key frame sequence, obtaining a monitoring image sequence with the same feature, and performing image splicing on the monitoring image sequence by a video splicing method based on a random sample consensus algorithm, thereby obtaining a spliced monitoring image sequence; performing fusion on the spliced monitoring image sequence by a video fusion algorithm based on an adaptive threshold, thereby obtaining a fused monitoring image sequence; and performing identification on the optimized key frame sequence and the fused monitoring image sequence respectively by a target identification algorithm, obtaining a first confidence degree and a second confidence degree, and obtaining a corrected confidence degree by using an adaptive weighted fusion method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of target recognition technology, specifically a target recognition method based on surveillance video. Background Technology

[0002] With the development of smart cities and security systems, the coverage of video surveillance networks continues to expand. However, the fixed perspective of a single camera has inherent limitations. When a target moves in the area where multiple cameras are covered, a single view often cannot capture the complete trajectory.

[0003] Current mainstream methods mainly use videos captured by a single camera for target recognition. By analyzing continuous footage captured by the same camera over time, the target in a single surveillance video can be identified.

[0004] However, traditional methods mainly rely on video captured by a single camera for target recognition. In scenarios where the target is moving, the target will undergo dynamic changes such as position and posture within the camera's field of view. The shooting range of a single camera is limited, making it difficult to fully capture the target's movement process and overall appearance, resulting in the loss of key features and the inability to accurately identify the target category and behavior. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a target recognition method based on surveillance video to solve the problems existing in the background technology.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a target recognition method based on surveillance video, comprising the following steps:

[0007] S1: Acquire several area monitoring videos, extract the several area monitoring videos by optimizing the inter-frame difference method to obtain a preliminary monitoring image key frame sequence, and filter the preliminary monitoring image key frame sequence by the secondary clustering method to obtain an optimized monitoring image key frame sequence.

[0008] S2: The optimized monitoring image keyframe sequence is subjected to feature matching by the scale-invariant feature transformation algorithm to obtain a monitoring image sequence with the same features. The monitoring image sequence with the same features is then spliced ​​by the video splicing method based on the random sampling consensus algorithm to obtain a spliced ​​monitoring image sequence.

[0009] S3: The spliced ​​surveillance image sequence is fused using a video fusion algorithm based on an adaptive threshold to obtain a fused surveillance image sequence;

[0010] S4: The optimized monitoring image keyframe sequence is subjected to target recognition using a target recognition algorithm with dual-layer routing attention and deformable convolution to obtain a first confidence level; the fused monitoring image sequence is subjected to target recognition using a target recognition algorithm with dual-layer routing attention and deformable convolution to obtain a second confidence level.

[0011] S5: The first confidence level and the second confidence level are fused by an adaptive weighted fusion method to obtain a corrected confidence level. The target detection result of the surveillance video is determined based on the corrected confidence level, thereby realizing target recognition of the surveillance video.

[0012] Preferably, the step of extracting the monitoring video of the several regions by optimizing the inter-frame difference method to obtain the preliminary monitoring image keyframe sequence includes the following specific steps:

[0013] Local extrema detection is performed on the regional surveillance video using the difference of Gaussian operator. N is the number of keyframes to be extracted. The video frames corresponding to the first N local extrema are the keyframes. All the obtained keyframes are arranged in chronological order of shooting time to obtain the preliminary keyframe sequence of the surveillance image.

[0014] Preferably, the step of filtering the initial keyframe sequence of the monitoring image using a secondary clustering method to obtain an optimized keyframe sequence of the monitoring image includes the following steps:

[0015] The monitoring image sequence corresponding to the initial monitoring image keyframe sequence is: The formula for calculating the joint probability of grayscale values ​​of two adjacent monitoring images is:

[0016]

[0017] in, express and The joint probability of grayscale values ​​from two surveillance images. Two frames of grayscale values and The frequency of the joint histogram This represents the total number of pixels in the monitored image.

[0018] The formula for calculating the edge probability of grayscale values ​​between two adjacent monitoring images is:

[0019]

[0020]

[0021] in, For monitoring images grayscale value The marginal probability, For monitoring images grayscale value The marginal probability, Indicates the monitoring image When the grayscale value i is fixed, traverse the monitoring images. Sum the probabilities of all grayscale values ​​j. Indicates the image When the grayscale value j is fixed, traverse the monitoring images. Sum the probabilities of all grayscale values ​​i;

[0022] The mutual information of adjacent monitoring images The calculation formula is:

[0023]

[0024] in, This represents the mutual information value between the nth frame and the (n+1)th frame of the surveillance image. express and The joint probability of grayscale values ​​from two surveillance images. For monitoring images grayscale value The marginal probability, For monitoring images grayscale value The marginal probability, They represent and The grayscale values ​​of the monitored images are in the image, where n is the frame index;

[0025] The formula for calculating the variance of the mutual information values ​​between high-similarity and low-similarity surveillance images is as follows:

[0026]

[0027] in, The variance of the mutual information values ​​between the two classes of surveillance images with high similarity and low similarity is given. The mean mutual information of surveillance images of high and low similarity. To monitor the number of frames in the image, The optimal split point;

[0028] The formula for calculating the optimal split point is:

[0029]

[0030]

[0031] in, As the optimal split point, the mutual information sequences are divided into two categories: high similarity and low similarity. The total number of candidate keyframes is obtained by optimizing the frame difference method. This indicates the number of surveillance images in the first T frames, i.e., the number of highly similar surveillance images. This indicates the number of surveillance images in the remaining frames, i.e., the number of surveillance images with low similarity. and The variance of the mutual information between the two classes of surveillance images with high similarity and low similarity is given.

[0032] The first candidate keyframe is selected as the first cluster. The mutual information value between adjacent frames is determined, and the mutual information value is greater than the optimal segmentation point. If a node is added to the current cluster, and the mutual information value is less than or equal to the optimal split point, then it is considered a new node. This generates a new cluster. Once all frames have been compared, the process terminates, resulting in an ordered cluster. ;

[0033] The formula for calculating the average mutual information of neighboring clusters is:

[0034]

[0035] in, The average value of the mutual information of the monitoring images of neighboring clusters. express The number of monitored image frames for The number of monitoring image frames in the data. express The surveillance images in the video, for The surveillance images in the video, express Mutual information between two surveillance images Indicates a clustered index;

[0036] The result clusters obtained from the first clustering Based on this, a preset threshold T1 is established. If the value is greater than the threshold T1, then merge. and ,like If the cluster size is less than or equal to the threshold T1, it is not merged. The process continues to check the next neighboring cluster and repeats the above process to merge clusters until... After all cluster comparisons have been completed, the formula for calculating the final selected keyframe of the monitoring image is as follows:

[0037]

[0038] in, These are the key frames of the final selected surveillance images. This indicates the number of monitored image frames in the final merged cluster. and Represents any two frames within a cluster and Surveillance images, This represents the average mutual information between the current frame monitoring image and other frame monitoring images within the cluster. Indicates the index of other frames that are different from the current frame's monitored image.

[0039] Preferably, the step of performing feature matching on the optimized keyframe sequence of the surveillance images using a scale-invariant feature transformation algorithm to obtain a sequence of surveillance images with the same features includes the following steps:

[0040] Construct a Gaussian pyramid. The optimized surveillance image is magnified by one time and used as the first layer of the first group of the Gaussian pyramid. The magnified optimized surveillance image is then convolved with a Gaussian layer and used as the second layer of the first group. Multiplying by a scaling factor k yields a smoothing factor k. This process is repeated to smooth the second layer of the first group of images. The smoothed second layer of the first group becomes the third layer of the same group. This process is repeated until M layers of images are obtained, with corresponding smoothing coefficients of 0. , , , ... ;

[0041] The third-to-last layer of the first group is downsampled with a scaling factor of 2. The resulting image is used as the first layer of the second group. Then, the first layer of the second group is convolved with Gaussian to obtain the second layer of the second group. This process is repeated to obtain O groups, each with M layers, for a total of O*M images.

[0042] A Gaussian difference pyramid (DOG) is constructed for the keyframes. The Gaussian difference pyramid contains O groups of M layers. The first layer of the first group of the Gaussian difference pyramid is obtained by subtracting the first layer of the first group of the Gaussian pyramid from the second layer of the first group of the Gaussian pyramid. The above steps are repeated to obtain several difference images. Combining the several difference images together forms the difference pyramid.

[0043] Feature points exist at local extrema in the difference of Gaussian pyramid. A pixel is considered an extremum when its value is greater than or less than the values ​​of 8 neighboring pixels in its 3×3 neighborhood and 18 neighboring pixels in the corresponding 3×3 neighborhoods of the adjacent upper and lower layers, totaling 26 pixels.

[0044] The extreme points are filtered and precisely located to obtain feature points. The gradients of the feature points are calculated and a direction histogram is established to obtain the direction of the feature points.

[0045] A descriptor is generated for each feature point. The feature points are matched with each other using the descriptors, and potential matching point pairs are identified to obtain a sequence of monitoring images with the same features.

[0046] Preferably, the step of stitching together the surveillance image sequences with the same characteristics using a video stitching method based on a random sampling consensus algorithm to obtain a stitched surveillance image sequence includes the following steps:

[0047] The surveillance image sequence with the same characteristics is stitched together using a video stitching method based on a random sampling consensus algorithm to obtain a stitched surveillance image sequence:

[0048] homography matrix The solution to the system of linear equations is as follows:

[0049]

[0050] in, and Let p and p' be the homogeneous coordinates of corresponding feature points p and p' in the two monitoring images. It is a homography matrix;

[0051] After obtaining the homography transformation matrix H, a continuous panoramic view sequence is constructed through dynamic projection transformation. For each monitoring image frame to be stitched, the source image pixel coordinates are mapped to a unified panoramic canvas coordinate system using the homography matrix. Spatial alignment is achieved through homogeneous coordinate transformation, and finally, panoramic images arranged continuously according to timestamps are output to form a stitched monitoring image sequence.

[0052] Preferably, the step of fusing the stitched surveillance image sequence using a video fusion algorithm based on an adaptive threshold to obtain a fused surveillance image sequence includes the following specific steps:

[0053] Source monitoring images to be fused and target surveillance images To merge:

[0054]

[0055] in, The pixel values ​​of the merged, stitched surveillance images. Source monitoring images The pixel value at coordinates (x, y) For target monitoring images The pixel value at coordinates (x, y) Source monitoring images Adaptive fusion weights, For target monitoring images The adaptive fusion weights determine the position of each pixel. and The degree of contribution;

[0056] The adaptive fusion weight function for the source and target surveillance images is expressed as:

[0057]

[0058]

[0059] in, Source monitoring images Adaptive fusion weights, For target monitoring images Adaptive fusion weights, For adaptive fusion weight calculation function, An adaptive threshold;

[0060] Preferably, the target recognition algorithm with dual-layer routing attention and deformable convolution includes the following specific steps:

[0061] The target recognition algorithm with two-layer routing attention and deformable convolution includes an input end, a backbone network, a neck feature fusion module, and a prediction output end;

[0062] The input first passes through the Focus module, which divides the feature map of the input monitoring image into four complementary sub-images by slicing it into pixels. Then, a convolution operation is performed on each sub-image, and finally the sub-images are stitched together. This enhances the feature representation ability of small targets while preserving spatial information. Subsequently, the CBS module uses a 1×1 convolution kernel to compress the channels of the sub-images.

[0063] The backbone network mainly includes the DCBS module, which adopts the latest DCNv3. DCNv3 can be formally expressed as:

[0064]

[0065] Where G represents the total number of aggregate groups, Indicates the first Group projection weights Let represent the modulation scalar of the d-th sampling point in the s-th group, which is normalized by the softmax function along dimension D. This represents the input feature map of the s-th group. It corresponds to the sampling position in the s-th group. The offset, The coordinates of the convolution center are... For grouped indexes, For sampling point index, The coordinates of the standard convolution kernel sampling points. The coordinates are represented as obtained after DCNv3 convolution.

[0066] The neck feature fusion module adopts an FPN+PAN structure, which achieves multi-scale feature fusion through upsampling and stitching. The improved BADD adds three double-layer routing attention mechanism BRA modules after the FPN output.

[0067] For the input surveillance image feature map Divide it into S×S non-overlapping regions, each region containing The feature vectors, i.e., the input feature map Remodeling Then, the query vector, key vector, and value vectors Q, K, V are obtained through linear projection. :

[0068]

[0069] in, For querying the matrix, The key matrix, For value matrices, The input features are those after region partitioning. To query the projected weights of the matrix, The projection weights of the key matrix are... The projection weights of the value matrix;

[0070] The query vector and key vector at the region level of the monitoring image are obtained by averaging the query vector and key vector at the region level. and ,in , ∈ Then, the adjacency matrix of the region-to-region affinity graph is obtained. Where is the affinity matrix at the monitoring image region level. For monitoring image region-level query matrix, To monitor the key matrix at the image region level, the directed graph is pruned by retaining the first k connections for each region and discarding the rest. Perform a row-level top-k operation, where k defaults to 4, to obtain a routing index matrix. ,in, For the Top-k region routing index matrix, an attention operation is finally performed on the region routing index matrix. For each query vector in region i, it will focus on all key-value pairs in the set of k routing regions. The BAR attention mechanism first aggregates all key-value tensors, and then performs attention calculation on the aggregated key-value pairs:

[0071]

[0072]

[0073] in, and It consists of the aggregated monitoring image key tensor and monitoring image value tensor. LCE is an introduced local context enhancement term. This represents the attention calculation process. The attention value for the final calculated surveillance image;

[0074] The decoupling head structure uses a 1×1 convolutional layer to transform the number of channels in the input monitoring image feature map, reducing it to 256 dimensions. It then connects to two parallel branches, each using two 3×3 convolutional layers: one for classification and the other for regression. An additional SIoU branch is added to the regression branch for classification, determining whether the target belongs to the foreground or background. Through the three branches of the decoupling head, three prediction structures are obtained: first, category prediction, determining the category of the target within the bounding box; second, the two-dimensional position information of the target box, represented by the center coordinates and width and height of the predicted box; and finally, foreground / background judgment of the target box, represented by the foreground and background confidence scores after processing with the Sigmoid function. The decoupling head then concatenates the feature information from the three branches, outputting 85x8400 two-dimensional feature information.

[0075] The formula for calculating the target angle loss in a surveillance image is:

[0076]

[0077] in, To monitor the angle loss of the target in the image, the angle between the line connecting the center points and the coordinate axis is measured. To monitor the Euclidean distance between the center points of the target in the image, To monitor the vertical distance difference between the center points of the target in the image;

[0078] The formula for calculating the target distance loss in a surveillance image is:

[0079]

[0080] in, To monitor the target distance loss in the monitoring image, this is used to measure the planar distance between the predicted bounding box and the corresponding ground truth bounding box in the monitoring image. As a distance loss weighting factor, , This represents the coordinate offset between the predicted bounding box and the ground truth bounding box in the normalized surveillance image. Indicates direction index;

[0081] The formula for calculating the target shape loss in a surveillance image is:

[0082]

[0083] in, To monitor the shape loss of the target in the image, To monitor the aspect ratio of the target in the image, To monitor the target width difference ratio in the image, To monitor the target height difference ratio in the image, The shape sensitivity parameter represents the shape loss for each dataset, and its value is unique. (Setting...) =4, For shape indexing, including height and width;

[0084] The formula for calculating the regression loss function of the target bounding box in a surveillance image is as follows:

[0085]

[0086] in, This is a regression loss function for monitoring image target bounding boxes, used to measure the difference between the model-predicted monitoring image bounding boxes and the ground truth bounding boxes. To monitor the intersection-union ratio (IoU) between predicted bounding boxes and ground truth bounding boxes in an image, To monitor image distance loss, To monitor image shape loss, The bounding box coordinate loss weights are set to 5 by default. To monitor the number of grid cells in the feature map of the image, =7, The number of anchor boxes for each monitoring image grid. =1, This is a function indicating the presence of a monitored image; it is set to 1 when a target is matched. For monitoring image feature map grid indexing, Anchor box index for monitoring the image grid;

[0087] The formula for calculating the loss function for target recognition category prediction in surveillance images is as follows:

[0088]

[0089] in, To predict the category of target recognition in monitored images, This refers to the number of positive samples in the monitored image, i.e., the total number of predicted boxes that match the ground truth boxes. To monitor the number of grid cells in the target feature map of the image, =7, The number of anchor boxes in the feature map grid for each monitored image. =1, This is a target indication function for monitoring images; it is 1 when a target is matched. Here, C represents the target identification category index, and C represents the total number of target identification categories. To predict class probabilities, For real category labels, The binary cross-entropy loss function;

[0090] The formula for calculating the probability loss function of target presence in a surveillance image is as follows:

[0091]

[0092] in, There exists a probability loss function for monitoring image targets. This represents the total number of historical surveillance image samples. To predict the probability of a target's presence in a surveillance image, To ensure that the monitored image actually contains a label, the determination is based on the IoU threshold;

[0093] The final target recognition loss function for surveillance images is:

[0094]

[0095] in, The final target recognition loss function for surveillance images is... This is a loss function for class prediction in monitoring image target recognition, used to measure the difference between the class predicted by the model and the true class. The target existence probability loss function in the monitored image is used to measure the difference between the model's predicted probability of target existence and the actual probability of target existence. The loss function is the bounding box regression function for the monitoring image.

[0096] Preferably, the target recognition algorithm using dual-layer routing attention and deformable convolution to perform target recognition on the optimized keyframe sequence of the surveillance image and obtain a first confidence score includes the following specific steps:

[0097] The formula for calculating the first confidence level is:

[0098]

[0099] in, The first confidence level is the value that comprehensively reflects the confidence level of the probability of a target's presence in a surveillance video image, and its value ranges from [0,1]. This represents the Sigmoid activation function. To predict the probability of a target's presence in a surveillance image, This represents the total number of historical surveillance image samples. This represents the normalized average of the probabilities of all predictions.

[0100] Preferably, the target recognition algorithm using two-layer routing attention and deformable convolution to perform target recognition on the fused surveillance image sequence to obtain a second confidence score includes the following specific steps:

[0101] The BADD algorithm is used to perform target recognition on the fused surveillance image sequence to obtain the second confidence level. .

[0102] Preferably, the step of fusing the first confidence level and the second confidence level using an adaptive weighted fusion method to obtain the corrected confidence level includes the following specific steps:

[0103] Construct the confidence weight allocation function:

[0104]

[0105] in, Second confidence level The weights, with values ​​ranging from [0,1], First confidence level The weights, and Complementary This is an S-curve adjustment factor used to control the steepness of the transition region. This value can quickly respond to changes in occlusion rate. It will jump rapidly from near 0 to near 1. This is the offset; a weight jump is triggered when the occlusion rate is >50%. Light uniformity is used to dynamically adjust the weights. ;

[0106] The final formula for calculating the revised confidence level is:

[0107]

[0108] in, To adjust the confidence level, As the first confidence level, As the second confidence level, Second confidence level The weight, First confidence level The weight.

[0109] This invention provides a target recognition method based on surveillance video, involving machine learning and deep learning technologies, which has the following beneficial effects:

[0110] (1) By using the scale-invariant feature transformation algorithm and video stitching technology based on random sampling consensus algorithm, multiple area monitoring images are integrated into a coherent panoramic view. Based on dynamic threshold, stitching gaps and ghosting phenomena are eliminated, so that the monitoring images of adjacent cameras are naturally connected. This cross-view collaboration significantly expands the monitoring field of view, so that targets such as vehicles and pedestrians always maintain a complete trajectory during cross-regional movement, solving the problem of field of view fragmentation in traditional single-camera monitoring.

[0111] (2) Improve the reliability of target recognition by using a target detection algorithm with dual-layer routing attention and deformable convolution. Introduce an attention routing mechanism so that the network can autonomously focus on the effective local features of the target and avoid interference from the occluded area. Adopt a flexible feature sampling strategy to dynamically adapt to changes in posture such as vehicle turning and pedestrian crouching. This design enables the system to accurately infer the target category through visible parts even in strong occlusion scenarios. It also keeps the positioning box tightly attached when the target is deformed, which greatly reduces the false detection rate and the risk of misjudgment.

[0112] (3) The confidence fusion mechanism can automatically balance the dual-path recognition results according to the environmental conditions. When the target is severely occluded, the system will prioritize the recognition signal of the cross-camera stitched view and use the complementary information of multiple perspectives to restore the full picture of the target. In abnormal lighting scenarios such as backlight and reflection, the system will automatically suppress the interference that may be introduced by the stitched image and instead rely on the stable analysis results of the original video stream. This dynamic balance strategy not only gives full play to the advantages of multi-view collaboration, but also avoids the risk of single-path information distortion, so that the target position and category judgment in the final output always maintain the optimal confidence level. Attached Figure Description

[0113] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0114] Figure 1 This is a flowchart of the target recognition method based on surveillance video proposed in this invention;

[0115] Figure 2 This is a step hierarchy diagram of obtaining the first confidence level and the second confidence level in a target recognition method based on surveillance video proposed in this invention;

[0116] Figure 3 This is a hierarchical diagram of the steps in obtaining target recognition results in a target recognition method based on surveillance video proposed in this invention;

[0117] Figure 4This is a performance comparison chart of the target recognition method based on surveillance video proposed in this invention and the traditional single-camera target recognition method under different degrees of occlusion.

[0118] Figure 5 This diagram shows a performance comparison between the target recognition method based on surveillance video proposed in this invention and the traditional single-camera target recognition method under different lighting conditions. Detailed Implementation

[0119] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0120] Please see Figures 1-5 The present invention provides a technical solution: a target recognition method based on surveillance video.

[0121] S1: Acquire several regional surveillance videos, extract the data from the several regional surveillance videos using the optimized inter-frame difference method to obtain a preliminary surveillance image key frame sequence, and filter the preliminary surveillance image key frame sequence using a secondary clustering method to obtain an optimized surveillance image key frame sequence.

[0122] To acquire surveillance videos from several areas, the videos are first preprocessed to remove the effects of lighting, noise, and other external factors on the images during capture. The mean filtering method is used to denoise the images by averaging the gray values ​​of all pixels in the neighborhood. This method is then applied to each point in the surveillance image to obtain the mean filtered image of the complete picture.

[0123] Histogram equalization is used to enhance surveillance images. First, the grayscale histogram and the total number of pixels in the surveillance image are calculated. Then, the grayscale distribution frequency of the surveillance image is calculated, and the cumulative grayscale distribution frequency of the surveillance image is calculated. Each grayscale value in the target surveillance image is a mapping of the grayscale value calculated from the surveillance image.

[0124] By optimizing the inter-frame difference method, the monitoring videos of the several regions are extracted to obtain a preliminary keyframe sequence of the monitoring images. The specific detection steps are as follows:

[0125] Import the surveillance video file, read the surveillance video and set the window size, and process the surveillance videos of the several areas frame by frame simultaneously; smoothing processing based on Hanning signal window convolution can effectively remove noise; calculate the inter-frame difference value of the surveillance video frame by frame and construct the difference sequence, perform double Gaussian filtering and difference calculation on the difference sequence to obtain the Gaussian difference operator, detect the local maxima in the Gaussian difference operator, select the frames corresponding to the first N local maxima as keyframes and extract them, and arrange all the obtained keyframes in chronological order of shooting time to form a preliminary surveillance image keyframe sequence.

[0126] It should be noted that when using the optimized inter-frame difference method to extract keyframes, the first few frames with the largest inter-frame differences may be misidentified as keyframes. Therefore, the keyframe sequence of the preliminary monitoring image is extracted by a secondary clustering method to obtain the optimized monitoring image keyframe sequence.

[0127] The monitoring image sequence corresponding to the initial monitoring image keyframe sequence is: To calculate the mutual information of adjacent monitoring images, it is necessary to first calculate the joint probability of the gray values ​​of two adjacent monitoring images and the edge probability of the gray values ​​of two adjacent monitoring images.

[0128] The formula for calculating the joint probability of grayscale values ​​of two adjacent monitoring images is:

[0129]

[0130] in, express and The joint probability of grayscale values ​​from two surveillance images. Two frames of grayscale values and The frequency of the joint histogram This represents the total number of pixels in the monitored image.

[0131] The formula for calculating the edge probability of grayscale values ​​between two adjacent monitoring images is:

[0132]

[0133]

[0134] in, For monitoring images grayscale value The marginal probability, For monitoring images grayscale value The marginal probability, Indicates the monitoring image When the grayscale value i is fixed, traverse the monitoring images. Sum the probabilities of all grayscale values ​​j. Indicates the image When the grayscale value j is fixed, traverse the monitoring images. Sum the probabilities of all grayscale values ​​i.

[0135] The mutual information of adjacent monitoring images The calculation formula is:

[0136]

[0137] in, This represents the mutual information value between the nth frame and the (n+1)th frame of the surveillance image. express and The joint probability of grayscale values ​​from two surveillance images. For monitoring images grayscale value The marginal probability, For monitoring images grayscale value The marginal probability, They represent and The grayscale value of the monitored image, where n is the frame index.

[0138] The mutual information of all generated surveillance images is sorted in descending order, and the mutual information value of the surveillance image in frame T is given. That is the threshold value we are looking for.

[0139] The formula for calculating the mean mutual information of high-similarity and low-similarity surveillance images is as follows:

[0140]

[0141] in, The mean mutual information of highly similar surveillance images. The mean mutual information of low-similarity surveillance images. The mutual information between the nth frame and the adjacent monitoring frames. To monitor the number of frames in the image, This is the optimal split point.

[0142] The formula for calculating the variance of the mutual information values ​​between high-similarity and low-similarity surveillance images is as follows:

[0143]

[0144] in, The variance of the mutual information values ​​between the two classes of surveillance images with high similarity and low similarity is given. The mean mutual information of surveillance images of high and low similarity. To monitor the number of frames in the image, This is the optimal split point.

[0145] The formula for calculating the optimal split point is:

[0146]

[0147]

[0148] in, As the optimal split point, the mutual information sequences are divided into two categories: high similarity and low similarity. The total number of candidate keyframes is obtained by optimizing the frame difference method. This indicates the number of surveillance images in the first T frames, i.e., the number of highly similar surveillance images. This indicates the number of surveillance images in the remaining frames, i.e., the number of surveillance images with low similarity. and The variance of the mutual information between the two classes of surveillance images with high similarity and low similarity is denoted as .

[0149] The first candidate keyframe is used as the first cluster. The mutual information value between the next two frames is then determined; if it is greater than the optimal segmentation point... If the value is less than or equal to the optimal split point, then add it to the current cluster. This generates a new cluster. Once all frames have been compared, the process terminates, resulting in an ordered cluster. .

[0150] The purpose of secondary clustering is to merge the clusters generated by primary clustering based on the average mutual information of neighboring clusters, further simplifying keyframes and reducing redundancy. The formula for calculating the average mutual information of neighboring clusters is:

[0151]

[0152] in, The average value of the mutual information of the monitoring images of neighboring clusters. express The number of monitored image frames for The number of monitoring image frames in the data. express The surveillance images in the video, for The surveillance images in the video, express Mutual information between two surveillance images This represents a cluster index.

[0153] The result clusters obtained from the first clustering Based on this, the threshold T1 can be manually set, and different thresholds can be used to determine the final number of clusters. If the value is greater than the threshold T1, then merge. and ,like If the cluster size is less than or equal to the threshold, it is not merged. The process continues to check the next neighboring cluster and repeats until the clusters are merged. After all clusters have been compared and processed, and candidate keyframes have been extracted, and after primary and secondary clustering operations, the final calculation formula for the selected keyframes of the surveillance image is as follows:

[0154]

[0155] in, These are the key frames of the final selected surveillance images. This indicates the number of monitored image frames in the final merged cluster. and Represents any two frames within a cluster and Surveillance images, This represents the average mutual information between the current frame monitoring image and other frame monitoring images within the cluster. Indicates the index of other frames that are different from the current frame's monitored image.

[0156] It should be noted that the candidate keyframe sequence extracted by the optimized inter-frame difference method must contain at least one frame to ensure... >1, During the secondary clustering process of merging neighboring clusters, the number of frames in the new cluster after the merge. = + This ensures that the denominator is not zero.

[0157] This method can maintain the temporal sequence of the content in the original surveillance video. After multiple screenings, it reduces the number of redundant frames in the surveillance video. The extracted key frames convert multi-frame video information into a compact key frame image, thereby finding key information in the video and improving the efficiency of surveillance video image processing.

[0158] The frame with the highest average mutual information with other frames in the final cluster is selected as the key frame. All key frames are arranged in chronological order to obtain the optimized monitoring image key frame sequence.

[0159] S2: The optimized monitoring image keyframe sequence is feature-matched using a scale-invariant feature transformation algorithm to obtain a monitoring image sequence with the same features. The monitoring image sequence with the same features is then stitched together using a video stitching method based on a random sampling consensus algorithm to obtain a stitched monitoring image sequence.

[0160] Scale-Invariant Feature Transform (SIFT) is an algorithm used to detect and extract key point features in images. This algorithm utilizes local maxima in image space and scale space and uses the Laplacian filter response to extract feature points that are scale-, rotation-, and illumination-invariant.

[0161] First, construct a Gaussian pyramid. Optimize the surveillance image by doubling its size and use it as the first layer of the first group of the Gaussian pyramid. Then, convolve the enlarged image using a Gaussian convolution and use it as the second layer of the first group. Multiplying by a scaling factor k yields a smoothing factor k. This smoothing factor is used to smooth the second layer of the first group of images. The smoothed image becomes the third layer of the same group. This process is repeated until an M-layer image is obtained. Each layer within the same group has a different resolution due to the addition of the smoothing factor, and the corresponding smoothing coefficients are 0. , , , ... .

[0162] The third-to-last layer image in the first group is downsampled (its size is halved) with a scaling factor of 2. The resulting image is used as the first layer of the second group. Then, Gaussian convolution is performed on the first layer image of the second group to obtain the second layer of the second group. This process is repeated to obtain O groups, each with M layers, for a total of O*M images.

[0163] A Gaussian difference pyramid (DOG) is constructed from the keyframes. The Gaussian difference pyramid contains O groups and M layers. The first layer of the first group of the Gaussian difference pyramid is obtained by subtracting the first layer of the first group of the Gaussian pyramid from the second layer of the first group of the Gaussian pyramid. By repeating this process, several difference images can be generated from the known groups and layers. All the difference images are combined to form the difference pyramid.

[0164] Feature points exist as local extrema in the difference of Gaussian pyramid. A pixel is considered an extremum only when its value is greater than (or less than) the values ​​of 8 pixels in its 3×3 neighborhood and 18 pixels in the corresponding 3×3 neighborhoods of the adjacent upper and lower layers, for a total of 26 pixels.

[0165] After initially obtaining some extreme points, it is necessary to filter and precisely locate them. The purpose is to remove some extreme points with low contrast and unstable edge response that are detected incorrectly, thereby improving the anti-interference ability of these points and obtaining stable extreme points, i.e. feature points. Feature points have scale invariance. By obtaining the neighborhood information of their scale space, calculating the gradient and establishing the orientation histogram, the orientation of the feature points can be obtained.

[0166] A unique descriptor is generated for each feature point. By dividing the pixel region around the extreme point into blocks and calculating the orientation histogram, a vector can be generated to abstractly describe the extreme point.

[0167] These descriptors are used to match feature points. By comparing the descriptors of feature points in different images, potential matching point pairs can be identified, thus obtaining a sequence of surveillance images with the same features.

[0168] However, due to factors such as changes in lighting, noise interference, and differences in viewing angle, mismatches may occur during the matching process. This technical solution uses RANSAC-based video stitching technology to stitch together the monitoring image sequences with the same characteristics to obtain a stitched monitoring image sequence.

[0169] The improved RANSAC algorithm proceeds as follows: Guided sampling is performed on the sampled data by calculating Euclidean distance and combining it with prior information to sample feature points; four pairs of feature points are selected to calculate the homography matrix H*; based on the calculated homography matrix H*, all feature point pairs in the image are examined, dividing them into inliers (normal data) and outliers (abnormal data). When the number of inliers is greater than 4, the largest set of inliers is saved, and the homography matrix H is recalculated.

[0170] This algorithm achieves more comprehensive feature matching and classification by fusing global and local information. By synergistically utilizing these two types of information, it can more accurately identify targets, understand scenes, and improve overall performance.

[0171] Homography transformation can convert a point on one image plane to its corresponding position on another image plane. These two image planes typically capture views of the same scene from different perspectives. The key to mapping between different image planes is the homography matrix. The solution to the system of linear equations is as follows:

[0172]

[0173] in, and Let p and p' be the homogeneous coordinates of corresponding feature points p and p' in the two monitoring images. It is a homography matrix.

[0174] It should be noted that in a homogeneous coordinate system, scaling the coordinates of a point does not change its position; typically, the last parameter is changed. Set to 1, therefore With only 8 effective degrees of freedom, given a set of corresponding points (i.e., matching feature points) in two images, the homography matrix can be estimated by solving a system of linear equations. According to the above formula, a pair of characteristic points yields two equations. The solution is required. Therefore, 4 pairs of feature points are required.

[0175] After obtaining the homography transformation matrix, the system constructs a continuous panoramic view sequence through dynamic projection transformation. For each monitoring image frame to be stitched, the homography matrix is ​​used to map the source image pixel coordinates to a unified panoramic canvas coordinate system, and spatial alignment is achieved through homogeneous coordinate transformation.

[0176] Specifically, a panoramic coordinate system is established based on the first keyframe. The precise projection matrix is ​​recalculated and the cache is updated for subsequent keyframes. For non-keyframes, the mean matrix in the buffer queue is directly called to achieve fast projection. Finally, the panoramic images are output in a continuous sequence of timestamps to form a stitched monitoring image sequence. The stitched monitoring images show obvious cracks and ghosting, which need to be eliminated by a fusion algorithm.

[0177] S3: The spliced ​​surveillance image sequence is fused by a video fusion algorithm based on an adaptive threshold to obtain a fused surveillance image sequence.

[0178] Image fusion technology can fuse image information from multiple video sequences through certain algorithms and strategies to generate images with greater visual appeal and practical value. This technical solution proposes a video fusion algorithm based on adaptive threshold, which can automatically adjust the image fusion threshold according to the changes in local features between frames. By dynamically adjusting the threshold, the fusion weight is optimized to eliminate stitching gaps, ghosting, and color difference.

[0179] There are two source monitoring images to be merged. and target surveillance images The fused image is The calculation process of the image fusion algorithm is represented as follows:

[0180]

[0181] in, The pixel values ​​of the merged, stitched surveillance images. Source monitoring images The pixel value at coordinates (x, y) For target monitoring images The pixel value at coordinates (x, y) Source monitoring images Adaptive fusion weights, For target monitoring images The adaptive fusion weights determine the position of each pixel. and The extent of their contribution.

[0182] The adaptive fusion weight function for the source and target surveillance images is expressed as:

[0183]

[0184]

[0185] in, Source monitoring images Adaptive fusion weights, For target monitoring images Adaptive fusion weights, For adaptive fusion weight calculation function, This is an adaptive threshold.

[0186] It should be noted that the adaptive fusion weight calculation function The design is based on the principle that when the difference in a certain region of the image exceeds an adaptive threshold, the weight of that region in the fused image is increased. It will then update dynamically based on changes in the video content.

[0187] When a significant change in the content of the monitored image is detected, the threshold is recalculated, and the mean square error of the image content is considered. When the rate of change exceeds a set value, the threshold can be increased to improve the fusion weight of that region. The range of the set value is determined according to the specific application scenario. Based on the difference between MSE and the current threshold, the threshold is dynamically adjusted. The algorithm steps are as follows:

[0188] First, calculate the rate of change of the mean square error of the current monitored image relative to the previous frame:

[0189]

[0190] in, This represents the rate of change of the MSE (Mean Estimate) of the current frame's monitored image relative to the previous frame's monitored image. The mean square error between the current keyframe monitoring image and the previous frame monitoring image. This represents the mean square error of the previous frame of the monitoring image.

[0191] Secondly, use sensitivity factors To adjust for the effect of the rate of change of the mean square error. Where is the sensitivity factor, preset =1.5, used to control the threshold update magnitude. The adjusted mean square error rate of change is used to calculate the rate of change of the mean square error. Update threshold:

[0192]

[0193] in, For the updated threshold, The threshold to be updated is currently set. This represents the adjusted mean square error rate of change.

[0194] It should be noted that, to avoid over-updating the threshold, it is necessary to perform an update after the update. Limit it to ensure it is within a reasonable range. ,in, and These are the minimum and maximum values ​​of the threshold, respectively, used to ensure that the threshold does not exceed the preset range. They are set according to the specific application scenario to prevent the threshold from being updated to an inappropriate level.

[0195] Finally, the spliced ​​surveillance image sequence is fused using this adaptive threshold video fusion algorithm to obtain a fused surveillance image sequence.

[0196] S4: Target recognition is performed on the optimized monitoring image keyframe sequence using a target recognition algorithm with dual-layer routing attention and deformable convolution to obtain a first confidence level. Target recognition is then performed on the fused monitoring image sequence using the same algorithm to obtain a second confidence level.

[0197] This technical solution proposes a target recognition algorithm BADD based on two-layer routing attention and deformable convolution. It addresses the difficulties in feature extraction and target size variation in complex environments. Based on YOLOX-s network, the overall network structure consists of an input end, a backbone network, a neck feature fusion module, and a prediction output end.

[0198] The input first passes through the Focus module, which divides the input monitoring image feature map into four complementary sub-images by slicing it into pixels. Then, a convolution operation is performed on each sub-image, and finally the sub-images are stitched together. This enhances the feature representation capability of small targets while preserving spatial information. Subsequently, the CBS module uses a 1×1 convolution kernel to compress the channels of the sub-images.

[0199] The backbone network mainly includes the DCBS module, which adopts the latest DCNv3. DCNv3 can be formally expressed as:

[0200]

[0201] Where G represents the total number of aggregate groups, Indicates the first Group projection weights Let represent the modulation scalar of the d-th sampling point in the s-th group, which is normalized by the softmax function along dimension D. This represents the input feature map of the s-th group. It corresponds to the sampling position in the s-th group. The offset, The coordinates of the convolution center are... For grouped indexes, For sampling point index, The coordinates of the standard convolution kernel sampling points. The coordinates are represented by the coordinates obtained after DCN v3 convolution.

[0202] The neck feature fusion module adopts an FPN+PAN structure, which achieves multi-scale feature fusion through upsampling and splicing. The improved BADD adds three two-layer routing attention mechanism (BRA) modules after the FPN output. BRA is a dynamic sparse attention mechanism that achieves more flexible computing power allocation through two-layer routing. Its core idea is to selectively retain only a small portion of the routing area, while filtering and excluding most key-value pairs with low correlation, thereby eliminating redundant information.

[0203] For the input surveillance image feature map The algorithm first divides it into S×S non-overlapping regions, each region containing The feature vectors, i.e., the input feature map Remodeling Then, the query vector, key vector, and value vectors Q, K, V are obtained through linear projection. :

[0204]

[0205] in, For querying the matrix, The key matrix, For value matrices, The input features are those after region partitioning. To query the projected weights of the matrix, The projection weights of the key matrix are... The projection weights are the values ​​of the matrix.

[0206] By constructing a directed graph, we can identify the regions that each specific area should participate in, thus determining the participation relationships between regions. We then calculate the region-level query vector and key vector of the monitoring image by averaging the query vector and key vector at the region level. and ,in , ∈ Then, the adjacency matrix of the region-to-region affinity graph is obtained. Where is the affinity matrix at the monitoring image region level. For monitoring image region-level query matrix, To monitor the key matrix at the image region level, the directed graph is pruned by retaining the first k connections for each region and discarding the rest. Perform a row-level top-k operation, where k defaults to 4, to obtain a routing index matrix. ,in, For the Top-k region routing index matrix, an attention operation is finally performed on the region routing index matrix. For each query vector in region i, it will focus on all key-value pairs in the set of k routing regions. The BAR attention mechanism first aggregates all key-value tensors, and then performs attention calculation on the aggregated key-value pairs:

[0207]

[0208]

[0209] in, and It consists of the aggregated monitoring image key tensor and monitoring image value tensor. LCE is an introduced local context enhancement term. This represents the attention calculation process. The attention value for the final calculated surveillance image.

[0210] The decoupling head structure first uses a 1×1 convolutional layer to transform the number of channels in the input surveillance image feature map, reducing it to 256 dimensions. Then, it connects two parallel branches, each using two 3×3 convolutional layers: one for classification and the other for regression. An additional SIoU branch is added to the regression branch for classification, determining whether the target belongs to the foreground or background. Through the three branches of the decoupling head, three prediction structures are obtained: first, category prediction, determining the category of the target within the bounding box; second, the two-dimensional position information of the bounding box, represented by the center coordinates and width and height of the predicted bounding box; and finally, the foreground / background judgment of the bounding box, represented by the foreground and background confidence scores processed by the Sigmoid function. Finally, the decoupling head concatenates the feature information from the three branches, outputting 85x8400 two-dimensional feature information.

[0211] This technical solution employs the SIoU calculation method, which addresses the instability issue when processing highly overlapping or approximate bounding boxes. SIoU reduces the sensitivity of IoU by introducing a smoothing term, thereby improving the stability and robustness of the bounding box regression loss function. The formula for calculating the target angle loss in the monitoring image is as follows:

[0212]

[0213] in, To monitor the angle loss of the target in the image, the angle between the line connecting the center points and the coordinate axis is measured. To monitor the Euclidean distance between the center points of the target in the image, To monitor the vertical distance difference between the center points of the target in the image.

[0214] The formula for calculating the target distance loss in a surveillance image is:

[0215]

[0216] in, To monitor the target distance loss in the monitoring image, this is used to measure the planar distance between the predicted bounding box and the corresponding ground truth bounding box in the monitoring image. As a distance loss weighting factor, , This represents the coordinate offset between the predicted bounding box and the ground truth bounding box in the normalized surveillance image. Indicates the direction index.

[0217] The formula for calculating the target shape loss in a surveillance image is:

[0218]

[0219] in, To monitor the shape loss of the target in the image, To monitor the aspect ratio of the target in the image, To monitor the target width difference ratio in the image, To monitor the target height difference ratio in the image, The shape sensitivity parameter represents the shape loss for each dataset, and its value is unique. (Setting...) =4, For shape indexing, including height and width.

[0220] The formula for calculating the regression loss function of the target bounding box in a surveillance image is as follows:

[0221]

[0222] in, This is a regression loss function for monitoring image target bounding boxes, used to measure the difference between the model-predicted monitoring image bounding boxes and the ground truth bounding boxes. To monitor the intersection-union ratio (IoU) between predicted bounding boxes and ground truth bounding boxes in an image, To monitor image distance loss, To monitor image shape loss, The bounding box coordinate loss weights are set to 5 by default. To monitor the number of grid cells in the feature map of the image, =7, The number of anchor boxes for each monitoring image grid. =1, This is a function indicating the presence of a monitored image; it is set to 1 when a target is matched. For monitoring image feature map grid indexing, Anchor box index for monitoring the image grid.

[0223] The formula for calculating the loss function for target recognition category prediction in surveillance images is as follows:

[0224]

[0225] in, To predict the category of target recognition in monitored images, This refers to the number of positive samples in the monitored image, i.e., the total number of predicted boxes that match the ground truth boxes. To monitor the number of grid cells in the target feature map of the image, =7, The number of anchor boxes in the feature map grid for each monitored image. =1, This is a target indication function for monitoring images; it is 1 when a target is matched. Here, C represents the target identification category index, and C represents the total number of target identification categories. To predict class probabilities, For real category labels, This is the binary cross-entropy loss function.

[0226] The formula for calculating the probability loss function of target presence in a surveillance image is as follows:

[0227]

[0228] in, There exists a probability loss function for monitoring image targets. This represents the total number of historical surveillance image samples. To predict the probability of a target's presence in a surveillance image, To ensure that the monitored images actually contain labels, a threshold for IoU (Intersection over Union) is used for determination.

[0229] It should be noted that the IoU threshold is the optimal value obtained through target recognition experiments in backlit and reflective scenes.

[0230] This technical solution uses the UA-DETRAC vehicle detection dataset to train the model, which contains 100,000 frames of highway monitoring video, covering different lighting, weather and occlusion scenarios, and provides rotating bounding boxes. This solution extracts 80% of the samples for training and 20% for validation.

[0231] The final target recognition loss function for surveillance images is:

[0232]

[0233] in, The final target recognition loss function for surveillance images is... This is a loss function for class prediction in monitoring image target recognition, used to measure the difference between the class predicted by the model and the true class. The target existence probability loss function in the monitored image is used to measure the difference between the model's predicted probability of target existence and the actual probability of target existence. The loss function is the bounding box regression function for the monitoring image.

[0234] The BADD algorithm is used to perform target recognition on the optimized keyframe sequence of the surveillance image. At the output of the decoupled head, the regression branch generates the target bounding box coordinates, the classification branch outputs the target category probability, and the foreground / background judgment branch generates the first confidence score using the Sigmoid function. The formula for calculating the first confidence score is as follows:

[0235]

[0236] in, The first confidence level is the value that comprehensively reflects the confidence level of the probability of a target's presence in a surveillance video image, and its value ranges from [0,1]. This represents the Sigmoid activation function. To predict the probability of a target's presence in a surveillance image, This represents the total number of historical surveillance image samples. This represents the normalized average of the probabilities of all predictions.

[0237] Similarly, the BADD algorithm is used to perform target recognition on the fused monitoring image sequence to obtain... arrive Second confidence level .

[0238] S5: The first confidence level and the second confidence level are fused by an adaptive weighted fusion method to obtain a corrected confidence level. The target detection result of the surveillance video is determined based on the corrected confidence level, thereby realizing target recognition of the surveillance video.

[0239] After obtaining the first confidence level (target recognition result of the original surveillance image) and the second confidence level (target recognition result of the fused and stitched image), the core innovation of this solution lies in the design of an adaptive weighted fusion mechanism based on environmental perception. This method dynamically analyzes the occlusion degree and lighting conditions of the scene where the target is located, intelligently allocates the weight ratio of the two types of confidence, and finally generates a more robust corrected confidence level.

[0240] The core lies in establishing a mathematical mapping between scene complexity and weight allocation to avoid scene mismatch problems caused by empirical parameters. Specifically, for each detected target, its occlusion factor is first calculated. , in, Used to quantify the severity of target occlusion. =0 indicates that it is fully visible. =1 indicates complete occlusion. The visible pixel area is directly calculated by summing the mask matrices output from the foreground / background decision branch. The predicted bounding box area is directly output by the detector.

[0241] Next, the uniformity of illumination is calculated as follows: ,in, Illumination uniformity reflects the evenness of light distribution. A value close to 1 indicates uniform illumination. The closer it is to 0, the stronger the change in light intensity. To monitor the variance of pixel intensity within the target region of an image.

[0242] Based on the above parameters, construct the confidence weight allocation function:

[0243]

[0244] in, Second confidence level The weights, with values ​​ranging from [0,1], First confidence level The weights, and Complementary This is an S-curve adjustment factor used to control the steepness of the transition region. This value can quickly respond to changes in occlusion rate. It will jump rapidly from near 0 to near 1. This is the offset; a weight jump is triggered when the occlusion rate is >50%. Light uniformity, dimensionless, used for dynamically adjusting weights. .

[0245] It should be noted that when the vehicle is obscured by more than 50%, the index item... The value rapidly approaches 1, significantly increasing the decision weight for the confidence of stitched images, especially in strong backlighting scenarios such as tunnel exits. To ensure synchronized attenuation and avoid confidence distortion caused by overexposure in the stitched images, Ensure that the weights change continuously near the critical point to eliminate detection box jitter caused by weight jumps.

[0246] The final formula for calculating the revised confidence level is:

[0247]

[0248] in, To adjust the confidence level, As the first confidence level, As the second confidence level, Second confidence level The weight, First confidence level The weight.

[0249] The system achieves final target recognition by setting a dynamic confidence threshold: when the confidence level is adjusted... When the value exceeds a preset threshold of 0.7, the system determines that a valid target exists in the area and simultaneously outputs the target category and location information. This determination process deeply integrates the dual verification advantages of the original video stream and the stitched video stream. Especially when the target is located in the area where cameras meet, the system automatically associates target segments in adjacent frames through a confidence fusion mechanism, automatically stitching together the target trajectories scattered in each independent video into a complete cross-regional behavior chain, completely solving the trajectory breakage problem caused by traditional methods. This mechanism ensures recognition accuracy while achieving seamless tracking of target behavior in full-area monitoring scenarios.

[0250] This technical solution proposes a target recognition method based on surveillance video. First, it employs optimized inter-frame difference and secondary clustering to achieve efficient extraction of keyframes. Second, it constructs a cross-camera stitched view based on SIFT feature matching and an improved RANSAC algorithm, and combines adaptive threshold fusion technology to eliminate stitching artifacts. Finally, it performs target detection through a two-layer routing attention mechanism and the deformable convolutional network BADD algorithm, fuses the confidence scores of the original video and the stitched video to obtain a corrected confidence score, and performs surveillance video recognition based on the corrected confidence score, significantly improving the robustness of recognition in complex surveillance scenarios.

[0251] By using scale-invariant feature transformation algorithms and video stitching technology based on random sampling consensus algorithms, multiple monitoring images from different areas are merged into a coherent panoramic view. Dynamic thresholds are used to eliminate stitching gaps and ghosting, allowing the monitoring images from adjacent cameras to connect naturally. This cross-view collaboration significantly expands the monitoring field of view, ensuring that targets such as vehicles and pedestrians maintain a complete trajectory during cross-regional movement, thus solving the problem of fragmented field of view in traditional single-camera monitoring.

[0252] The reliability of target recognition is improved by using a target detection algorithm with dual-layer routing attention and deformable convolution. The introduction of an attention routing mechanism enables the network to autonomously focus on the effective local features of the target, avoiding interference from occluded areas. A flexible feature sampling strategy is adopted to dynamically adapt to changes in posture such as vehicle turning and pedestrian crouching. This design enables the system to accurately infer the target category through visible parts even in strong occlusion scenarios, and to maintain a tight fit of the localization box when the target is deformed, which greatly reduces the false negative rate and the risk of misjudgment.

[0253] The confidence fusion mechanism can automatically balance the dual-channel recognition results according to the environmental conditions. When the target is severely obscured, the system prioritizes the recognition signal of the cross-camera stitched view and uses complementary information from multiple perspectives to restore the full picture of the target. In abnormal lighting scenarios such as backlighting and reflection, the system automatically suppresses the interference that may be introduced by the stitched image and instead relies on the stable analysis results of the original video stream. This dynamic trade-off strategy not only leverages the advantages of multi-view collaboration but also avoids the risk of single-channel information distortion, ensuring that the final output target position and category judgment always maintain the optimal confidence level.

[0254] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, the phrase "comprising an element defined as..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0255] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A target identification method based on monitoring video, characterized in that: Includes the following steps: S1: Acquire surveillance videos of several areas, extract keyframe sequences from these videos using an optimized inter-frame difference method to obtain a preliminary keyframe sequence, and then filter this preliminary keyframe sequence using a secondary clustering method to obtain an optimized keyframe sequence. This includes the following specific steps: Local extrema detection is performed on the regional surveillance video using the difference of Gaussian operator. N is the number of keyframes to be extracted. The video frames corresponding to the first N local extrema points are the keyframes. All the obtained keyframes are arranged in chronological order of shooting time to obtain the preliminary surveillance image keyframe sequence. The preliminary monitoring image key frame sequence corresponds to a monitoring image sequence The calculation formula of the joint probability of the gray values of two adjacent monitoring images is: ; in, express and The joint probability of grayscale values ​​from two surveillance images. Two frames of grayscale values and The frequency of the joint histogram This represents the total number of pixels in the monitored image. The formula for calculating the edge probability of grayscale values ​​between two adjacent monitoring images is: ; ; in, For monitoring images grayscale value The marginal probability, For monitoring images grayscale value The marginal probability, Indicates the monitoring image When the grayscale value i is fixed, traverse the monitoring images. Sum the probabilities of all grayscale values ​​j. Indicates the image When the grayscale value j is fixed, traverse the monitoring images. Sum the probabilities of all grayscale values ​​i; Mutual information of adjacent monitoring images The formula for calculating mutual information is: ; in, This represents the mutual information value between the nth frame and the (n+1)th frame of the surveillance image. express and The joint probability of grayscale values ​​from two surveillance images. For monitoring images grayscale value The marginal probability, For monitoring images grayscale value The marginal probability, They represent and The grayscale values ​​of the monitored images are in the image, where n is the frame index; The formula for calculating the variance of the mutual information values ​​between high-similarity and low-similarity surveillance images is as follows: ; in, The variance of the mutual information values ​​between the two classes of surveillance images with high similarity and low similarity is given. The mean mutual information of surveillance images of high and low similarity. To monitor the number of frames in the image, The optimal split point; The formula for calculating the optimal split point is: ; ; in, As the optimal split point, the mutual information sequences are divided into two categories: high similarity and low similarity. The total number of candidate keyframes is obtained by optimizing the frame difference method. This indicates the number of surveillance images in the first T frames, i.e., the number of highly similar surveillance images. This indicates the number of surveillance images in the remaining frames, i.e., the number of surveillance images with low similarity. and The variance of the mutual information between the two classes of surveillance images with high similarity and low similarity is given. The first candidate keyframe is designated as the first cluster. The mutual information value between adjacent frames is determined; if the mutual information value is greater than the optimal segmentation point T, the frame is added to the current cluster; otherwise, the mutual information value is less than or equal to the optimal segmentation point T. This generates a new cluster. Once all frames have been compared, the process terminates, resulting in an ordered cluster. ; The formula for calculating the average mutual information of neighboring clusters is: ; in, The average value of the mutual information of the monitoring images of neighboring clusters. express The number of monitored image frames for The number of monitoring image frames in the data. express The surveillance images in the video, for The surveillance images in the video, express Mutual information between two surveillance images Indicates a clustered index; The result clusters obtained from the first clustering Based on this, a preset threshold T1 is established. If the value is greater than the threshold T1, then merge. and ,like If the cluster size is less than or equal to the threshold T1, it is not merged. The process continues to check the next neighboring cluster and repeats the above process to merge clusters until... After all cluster comparisons have been completed, the formula for calculating the final selected keyframe of the monitoring image is as follows: ; in, These are the key frames of the final selected surveillance images. This indicates the number of monitored image frames in the final merged cluster. and Represents any two frames within a cluster and Surveillance images, This represents the average mutual information between the current frame monitoring image and other frame monitoring images within the cluster. This indicates the index of other frames that are different from the current frame's monitored image; S2: The optimized monitoring image keyframe sequence is feature matched using the scale-invariant feature transformation algorithm to obtain a monitoring image sequence with the same features. The monitoring image sequence with the same features is then stitched together using a video stitching method based on the random sampling consensus algorithm to obtain a stitched monitoring image sequence. S3: The spliced ​​surveillance image sequence is fused using a video fusion algorithm based on an adaptive threshold to obtain a fused surveillance image sequence; S4: The optimized monitoring image keyframe sequence is subjected to target recognition using a target recognition algorithm with dual-layer routing attention and deformable convolution to obtain a first confidence level; the fused monitoring image sequence is subjected to target recognition using a target recognition algorithm with dual-layer routing attention and deformable convolution to obtain a second confidence level. S5: The first confidence level and the second confidence level are fused by an adaptive weighted fusion method to obtain a corrected confidence level. The target detection result of the surveillance video is determined based on the corrected confidence level, thereby realizing target recognition of the surveillance video.

2. The target recognition method based on surveillance video according to claim 1, characterized in that: The step of performing feature matching on the optimized keyframe sequence of the surveillance images using a scale-invariant feature transformation algorithm to obtain a sequence of surveillance images with the same features includes the following steps: Construct a Gaussian pyramid. The optimized surveillance image is magnified by one time and used as the first layer of the first group of the Gaussian pyramid. The magnified optimized surveillance image is then convolved with a Gaussian layer and used as the second layer of the first group. Multiplying by a scaling factor k yields a smoothing factor k. This process is repeated to smooth the second layer of the first group of images. The smoothed second layer of the first group becomes the third layer of the same group. This process is repeated until M layers of images are obtained, with corresponding smoothing coefficients of 0. , , , ... ; The third-to-last layer of the first group is downsampled with a scaling factor of 2. The resulting image is used as the first layer of the second group. Then, the first layer of the second group is convolved with Gaussian to obtain the second layer of the second group. This process is repeated to obtain O groups, each with M layers, for a total of O*M images. A Gaussian difference pyramid is constructed for the keyframes. The Gaussian difference pyramid contains O groups of M layers. The first layer of the first group of the Gaussian difference pyramid is obtained by subtracting the first layer of the first group of the Gaussian pyramid from the second layer of the first group of the Gaussian pyramid. The above steps are repeated to obtain several difference images. The several difference images are combined to form the difference pyramid. Feature points exist at local extrema in the difference of Gaussian pyramid. A pixel is considered an extremum point when the value of a pixel is greater than or less than the values ​​of 8 neighboring pixels in its 3×3 neighborhood and 18 neighboring pixels in the corresponding 3×3 neighborhoods of the adjacent upper and lower layers, totaling 26 pixels. The extreme points are filtered and precisely located to obtain feature points. The gradients of the feature points are calculated and a direction histogram is established to obtain the direction of the feature points. A descriptor is generated for each feature point. The feature points are matched with each other using the descriptors, and potential matching point pairs are identified to obtain a sequence of monitoring images with the same features. 3.The target identification method based on monitoring video of claim 2, characterized in that: The method of stitching together the surveillance image sequences with the same characteristics using a video stitching method based on a random sampling consensus algorithm to obtain a stitched surveillance image sequence includes the following steps: The surveillance image sequence with the same characteristics is stitched together using a video stitching method based on a random sampling consensus algorithm to obtain a stitched surveillance image sequence: homography matrix solving the linear equations ; wherein, and are the homogeneous coordinates of corresponding feature points p and p' in the two monitored images, is a homography matrix; After obtaining the homography transformation matrix H, a continuous panoramic view sequence is constructed through dynamic projection transformation. For each monitoring image frame to be stitched, the source image pixel coordinates are mapped to a unified panoramic canvas coordinate system using the homography matrix. Spatial alignment is achieved through homogeneous coordinate transformation, and finally, panoramic images arranged continuously according to timestamps are output to form a stitched monitoring image sequence.

4. The target identification method based on monitoring video according to claim 3, characterized in that: The process of fusing stitched surveillance image sequences using an adaptive threshold-based video fusion algorithm to obtain a fused surveillance image sequence includes the following specific steps: Fusing source and target surveillance images and target surveillance images Fusion: ; in, The pixel values ​​of the merged, stitched surveillance images. Source monitoring images At the pixel value of coordinates (x, y), For target monitoring images At the pixel value of coordinates (x, y), Source monitoring images Adaptive fusion weights, For target monitoring images The adaptive fusion weights determine the position of each pixel. and The degree of contribution; The adaptive fusion weight function for the source and target surveillance images is expressed as follows: ; ; in, Source monitoring images Adaptive fusion weights, For target monitoring images Adaptive fusion weights, For adaptive fusion weight calculation function, This is an adaptive threshold.

5. The target identification method based on monitoring video according to claim 4, characterized in that: The target recognition algorithm based on dual-layer routing attention and deformable convolution includes the following specific steps: The target recognition algorithm with two-layer routing attention and deformable convolution includes an input end, a backbone network, a neck feature fusion module, and a prediction output end; The input first passes through the Focus module, which divides the feature map of the input monitoring image into four complementary sub-images by slicing it into pixels. Then, a convolution operation is performed on each sub-image, and finally the sub-images are stitched together. This enhances the feature representation ability of small targets while preserving spatial information. Subsequently, the CBS module uses a 1×1 convolution kernel to compress the channels of the sub-images. The backbone network mainly includes the DCBS module, which adopts the latest DCNv3. DCNv3 can be formally expressed as: ; Where G represents the total number of aggregate groups, Indicates the first Group projection weights Let represent the modulation scalar of the d-th sampling point in the s-th group, which is normalized by the softmax function along dimension D. This represents the input feature map of the s-th group. It corresponds to the sampling position in the s-th group. The offset, The coordinates of the convolution center are... For grouped indexes, For sampling point index, The coordinates of the standard convolution kernel sampling points. The coordinates are represented as obtained after DCNv3 convolution. The neck feature fusion module adopts an FPN+PAN structure, which achieves multi-scale feature fusion through upsampling and stitching. The improved BADD adds three double-layer routing attention mechanism BRA modules after the FPN output. For the input surveillance image feature map Divide it into S×S non-overlapping regions, each region containing The feature vectors, i.e., the input feature map Remodeling Then, the query vector, key vector, and value vectors Q, K, V are obtained through linear projection. : ; in, For querying the matrix, The key matrix, For value matrices, The input features are those after region partitioning. To query the projected weights of the matrix, The projection weights of the key matrix are... The projection weights of the value matrix; The query vector and key vector at the region level of the monitoring image are obtained by averaging the query vector and key vector at the region level. and ,in , ∈ Then, the adjacency matrix of the region-to-region affinity graph is obtained. Where is the affinity matrix at the monitoring image region level. For monitoring image region-level query matrix, To monitor the key matrix at the image region level, the directed graph is pruned by retaining the first k connections for each region and discarding the rest. Perform a row-level top-k operation, where k defaults to 4, to obtain a routing index matrix. ,in, For the Top-k region routing index matrix, an attention operation is finally performed on the region routing index matrix. For each query vector in region i, it will focus on all key-value pairs in the set of k routing regions. The BAR attention mechanism first aggregates all key-value tensors, and then performs attention calculation on the aggregated key-value pairs: ; ; in, and It consists of the aggregated monitoring image key tensor and monitoring image value tensor. LCE is an introduced local context enhancement term. This represents the attention calculation process. The attention value for the final calculated surveillance image; The decoupling head structure uses a 1×1 convolutional layer to transform the number of channels in the input monitoring image feature map, reducing it to 256 dimensions. It then connects to two parallel branches, each using two 3×3 convolutional layers: one for classification and the other for regression. An additional SIoU branch is added to the regression branch for classification, determining whether the target belongs to the foreground or background. Through the three branches of the decoupling head, three prediction structures are obtained: first, category prediction, determining the category of the target within the bounding box; second, the two-dimensional position information of the target box, represented by the center coordinates and width and height of the predicted box; and finally, foreground / background judgment of the target box, represented by the foreground and background confidence scores after processing with the Sigmoid function. The decoupling head then concatenates the feature information from the three branches, outputting 85x8400 two-dimensional feature information. The formula for calculating the target angle loss in a surveillance image is: ; in, To monitor the angle loss of the target in the image, the angle between the line connecting the center points and the coordinate axis is measured. To monitor the Euclidean distance between the center points of the target in the image, To monitor the vertical distance difference between the center points of the target in the image; The formula for calculating the target distance loss in a surveillance image is: ; in, To monitor the target distance loss in the monitoring image, this is used to measure the planar distance between the predicted bounding box and the corresponding ground truth bounding box in the monitoring image. As a distance loss weighting factor, , This represents the coordinate offset between the predicted bounding box and the ground truth bounding box in the normalized surveillance image. Indicates direction index; The formula for calculating the target shape loss in a surveillance image is: ; in, To monitor the shape loss of the target in the image, To monitor the aspect ratio of the target in the image, To monitor the target width difference ratio in the image, To monitor the target height difference ratio in the image, The shape sensitivity parameter represents the shape loss for each dataset, and its value is unique. (Setting...) =4, For shape indexing, including height and width; The formula for calculating the regression loss function of the target bounding box in a surveillance image is as follows: ; in, This is a regression loss function for monitoring image target bounding boxes, used to measure the difference between the model-predicted monitoring image bounding boxes and the ground truth bounding boxes. To monitor the intersection-union ratio (IoU) between predicted bounding boxes and ground truth bounding boxes in an image, To monitor image distance loss, To monitor image shape loss, The bounding box coordinate loss weights are set to 5 by default. To monitor the number of grid cells in the feature map of the image, =7, The number of anchor boxes for each monitoring image grid. =1, This is a function indicating the presence of a monitored image; it is set to 1 when a target is matched. For monitoring image feature map grid indexing, Anchor box index for monitoring the image grid; The formula for calculating the loss function for target recognition category prediction in surveillance images is as follows: ; in, To predict the category of target recognition in monitored images, This refers to the number of positive samples in the monitored image, i.e., the total number of predicted boxes that match the ground truth boxes. To monitor the number of grid cells in the target feature map of the image, =7, The number of anchor boxes in the feature map grid for each monitored image. =1, This is a target indication function for monitoring images; it is 1 when a target is matched. Here, C represents the target identification category index, and C represents the total number of target identification categories. To predict class probabilities, For real category labels, The binary cross-entropy loss function; The formula for calculating the probability loss function of target presence in a surveillance image is as follows: ; wherein, is a monitoring image target existence probability loss function, is a total historical monitoring image sample number, is a predicted monitoring image target existence probability, is a monitoring image real existence label, determined according to an IoU threshold; The final target recognition loss function for surveillance images is: ; in, The final target recognition loss function for surveillance images is... This is a loss function for class prediction in monitoring image target recognition, used to measure the difference between the class predicted by the model and the true class. The target existence probability loss function in the monitored image is used to measure the difference between the model's predicted probability of target existence and the actual probability of target existence. The regression loss function is used for the bounding box of the monitoring image.

6. The target identification method based on monitoring video according to claim 5, characterized in that: The target recognition algorithm, which uses dual-layer routing attention and deformable convolution to perform target recognition on the optimized keyframe sequence of the surveillance image and obtain the first confidence score, includes the following specific steps: The formula for calculating the first confidence level is: ; in, The first confidence level is the value that comprehensively reflects the confidence level of the probability of a target's presence in a surveillance video image, and its value ranges from [0,1]. This represents the Sigmoid activation function. To predict the probability of a target's presence in a surveillance image, This represents the total number of historical surveillance image samples. This represents the normalized average of the probabilities of all predictions.

7. The target identification method based on monitoring video according to claim 6, characterized in that: The target recognition algorithm using dual-layer routing attention and deformable convolution to perform target recognition on the fused surveillance image sequence and obtain a second confidence score includes the following specific steps: The BADD algorithm is used to identify the target in the fused monitoring image sequence to obtain a second confidence level . 8.The target identification method based on monitoring video of claim 7, characterized in that: The step of fusing the first confidence level and the second confidence level using an adaptive weighted fusion method to obtain the corrected confidence level includes the following specific steps: Construct the confidence weight allocation function: ; in, Second confidence level The weights, with values ​​ranging from [0,1], First confidence level The weights, and Complementary This is an S-curve adjustment factor used to control the steepness of the transition region. It can quickly respond to changes in occlusion rate and It will jump rapidly from near 0 to near 1. This is the offset; a weight jump is triggered when the occlusion rate is >50%. Light uniformity is used to dynamically adjust the weights. ; The final formula for calculating the revised confidence level is: ; wherein, is a first confidence, is a first confidence, is a second confidence, is a second confidence is a weight, is a weight of the first confidence is a weight of the first confidence.

Citation Information

Patent Citations

  • Video analysis method and system of monitoring terminal

    CN118781526A