A method and system for tracking small moving targets in complex backgrounds

Through the combination of scale adaptive residual neural network and Transformer network, the problem of feature extraction and tracking of small target objects in complex background is solved, and accurate detection and real-time tracking of small targets are achieved.

CN116385492BActive Publication Date: 2025-07-08GUANGXI ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310377570.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-10
Publication Date
2025-07-08
Estimated Expiration
2043-04-10

AI Technical Summary

Technical Problem

The existing small-target object detection and tracking algorithms are difficult to effectively segment the target and background in complex backgrounds, and the small-target characteristics are not obvious, resulting in insufficient detection and tracking accuracy.

Method used

The scale adaptive residual neural network and the Transformer network are used for feature extraction and detection, and the encoding and decoding are combined with a multi-layer perceptron to calculate the similarity of adjacent frame images to determine the tracking target.

Benefits of technology

It realizes accurate identification and real-time tracking of small targets in complex contexts, improving the accuracy and efficiency of detection and tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116385492B_ABST
    Figure CN116385492B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for tracking small moving targets in a complex background. The tracking method includes obtaining a video segment of the small target to be tracked and converting the video segment into a sequence of images frame by frame; coarsely extracting features from each frame of image by using a scale-adaptive residual neural network, and then using a Transformer to perform multi-scale and fine-grained feature extraction on the coarsely extracted features to obtain fine-grained features; using a Transformer and a multi-layer perceptron to detect small targets in the fine-grained features to obtain the categories and detection frames of all small targets in the fine-grained features; calculating the similarity between the detection frames of the same-category small targets in two adjacent frames of images, and determining the tracking targets in each frame of image based on the similarity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to object tracking technology, and specifically to a method and system for tracking small moving targets in complex backgrounds. Background Art

[0002] Small moving target objects can be seen everywhere in our daily life. For example, the phenomenon of high-altitude throwing objects often occurs in current high-rise residential areas. Usually, small moving target objects have characteristics such as rotation, deformation, background interference, light change, occlusion, and low resolution. In recent years, with the continuous development and progress of deep learning technology, continuous breakthroughs have been made in the field of computer vision and it has been applied in many fields. The main task of computer vision is to enable the computer to have a pair of eyes like humans. Detecting and tracking small target objects in images is a difficult problem.

[0003] Currently, with the rapid development of deep learning, significant breakthroughs have been made in the field of object detection. However, there are still great deficiencies in detecting and tracking small target objects. Existing algorithms for detecting and tracking small moving target objects cannot well meet the application in actual complex scenarios, and mainly have the following problems:

[0004] 1. The difficulty in tracking small target objects lies in the fact that the features of the target are not obvious and there is less available feature information. If the resolution of the image itself is not high, small targets usually have only a few pixels and the features are not obvious.

[0005] 2. The problem of complex background interference. Detecting and tracking small moving targets in complex environments will be affected by factors such as light and occlusion. Therefore, it is difficult to separate the target from the background or similar objects and effectively achieve the detection and tracking of small targets. Summary of the Invention

[0006] Aiming at the above deficiencies in the prior art, the method and system for tracking small moving targets in complex backgrounds provided by the present invention solve the problem that the prior art cannot accurately identify small targets in complex scenarios.

[0007] To achieve the above invention purpose, the technical solution adopted by the present invention is as follows:

[0008] In the first aspect, a method for tracking small moving targets in complex backgrounds is provided, which includes the following steps:

[0009] S1. Obtain a video segment of the small target to be tracked and convert the video segment into a sequence of images frame by frame;

[0010] S2. Coarsely extract features from each frame of image using a scale-adaptive residual neural network, and then use a Transformer to perform multi-scale and fine-grained feature extraction on the coarsely extracted features to obtain fine-grained features;

[0011] S3. Use a Transformer and a multi-layer perceptron to detect small targets in the fine-grained features, and obtain the categories and detection frames of all small targets in the fine-grained features;

[0012] S4. Calculate the similarity between the detection frames of small targets of the same category in two adjacent frames of images, and determine the tracking targets in each frame of image based on the similarity.

[0013] The beneficial effects of the present invention are as follows: Through the super strong feature extraction ability of the residual neural network and the pixel-by-pixel extraction of features in the image by the Transformer in this solution, the feature information of small targets can be effectively extracted. The multi-layer perceptron is used for encoding and decoding, so as to accurately obtain the position information of the target, and the small target object in motion can be tracked more accurately, realizing the real-time tracking of moving small target objects under complex backgrounds.

[0014] Further, the convolutional kernels in the residual blocks of the scale-adaptive residual neural network adopt 1*3 and 3*1 convolutional kernels. When using smaller convolutional kernels for convolution operations, the small targets in the image can be feature-extracted pixel by pixel, and the extracted features are more targeted, and the number of parameters is reduced.

[0015] Further, the calculation formulas for coarse extraction and extraction of fine-grained features are respectively:

[0016] y = SA-Resnet(p n + bias)+X

[0017] Q = w*transformer(y, θ)+b

[0018] Among them, y is the feature after coarse extraction; SA-Resnet(·) is the scale-adaptive residual operation, bias is the dynamic adaptive operator; X is the area sampled by the convolutional kernel for the feature map; Q is the fine-grained feature; w is the adaptive weight value; transformer(·) is to obtain the picture set in which the moving small target appears in each frame in the video segment p, and then perform feature transformation on each picture in the picture set through the convolutional operator; θ is the learnable weight parameter of the Transformer model; b is the bias term of the network.

[0019] The beneficial effects of the above technical solution are as follows: In this solution, a scale-adaptive residual neural network is proposed through improvement. On the basis of improving the feature extraction ability, the number of parameters of the network model is reduced, and the features in the image are extracted pixel by pixel. Abundant features in the image are extracted during feature coarse extraction, which is convenient for subsequent encoding and decoding to obtain the positioning information of small target objects.

[0020] Further, the step S3 further includes:

[0021] S31. Encode the fine-grained features using the multi-head attention mechanism in the Transformer encoder:

[0022]

[0023] multi-Head(q, k, v) = Concat(head1,......, head i ) * W i

[0024]

[0025] where Q is the fine-grained feature, q is the query vector during encoding; k is the key, v is the value calculated by the multi-head attention, are the weight values corresponding to the key respectively; self-attention(·) is the self-attention calculation, head i is a subspace; n is the total number of subspaces; multi-Head(·) represents merging the multi-heads; W i is the weight of the model; w i is the weight value of the query vector, q transpose is the transpose of the query vector; k transpose is the matrix transpose of the query key, v transpose is the transpose of the vector value of the query;

[0026] S32. Decode the entities queried in the image using the multi-head attention mechanism in the transformer decoder;

[0027] S33. Map the output of the decoder using a multi-layer perceptron to obtain the categories and detection boxes of all small objects in the image:

[0028] class, box m = w1 * (decoder(encoder(multi-Head(q, k, v)))) + b1

[0029] where class is the object category; box n is the detection box of the small object; w1 is the weight of the multi-layer perceptron; encoder(.) is the encoding, decoder(.) is the decoding, and b1 is the bias term.

[0030] The beneficial effects of the above technical solution are as follows: Through improvement, this solution proposes a multi-head self-attention network model. By introducing the self-attention model during encoding and decoding, by obtaining the feature information of multiple subspaces, and then fusing the features of multiple subspaces, the accurate category and position information of small target objects are finally obtained.

[0031] Furthermore, the calculation formula for similarity is as follows:

[0032] A = Jaccard_SIM(box n , box n-1 , T n )

[0033] where A is the similarity; Jaccard_SIM(.) is the similarity calculation function; box n and box n-1 are the detection boxes of small target objects in the nth frame and the (n - 1)th frame images respectively; T n = (t1, t2,..., t n ) is the time series of the video segment.

[0034] Furthermore, the moving small target tracking method further includes real-time tracking the results of moving small target objects and sending out early warning information in real time.

[0035] In the second aspect, this solution provides a moving small target tracking system, which includes:

[0036] A video acquisition module, which is used to acquire the video segment of the small target to be tracked and convert the video segment into a sequence of frame-by-frame images;

[0037] A feature extraction module, which is used to roughly extract features from each frame of image by using a scale-adaptive residual neural network, and then use Transformer to perform multi-scale and fine-grained feature extraction on the roughly extracted features to obtain fine-grained features;

[0038] A small target detection module, which is used to detect small targets in the fine-grained features by using Transformer and a multi-layer perceptron to obtain the categories and detection boxes of all small targets in the fine-grained features;

[0039] A small target tracking module, which is used to calculate the similarity between the detection boxes of small targets of the same category in two adjacent frames of images and determine the tracking targets in each frame of image based on the similarity. Description of the Drawings

[0040] Figure 1 is a flowchart of the moving small target tracking method under complex backgrounds.

[0041] Figure 2Schematic diagram of the scale adaptive residual block in the scale adaptive residual neural network.

[0042] Figure 3 Schematic diagram of Transformer encoding and decoding. Detailed implementation manners

[0043] The following describes the detailed implementation manners of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed implementation manners. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0044] Reference Figure 1 , Figure 1 shows the flowchart of the moving small target tracking method under complex backgrounds; as Figure 1 shown, the method includes steps S1 to S4.

[0045] In step S1, a video segment of the small target to be tracked is obtained, and the video segment is converted into a sequence of images frame by frame; the conversion process is: p n = opencv(p), where, T n represents the image sequence, and p is the obtained video segment; the time sequence T n =(t1, t2,..., t n ).

[0046] In step S2, a scale adaptive residual neural network is used to perform rough feature extraction on each frame of the image;

[0047] y = SA-Resnet(p n + bias)+ X

[0048] where, y is the feature after rough extraction; SA-Resnet(·) is the scale adaptive residual operation; bias is the dynamic adaptive operator; X is the region where the convolution kernel samples the feature map;

[0049] Performing rough feature extraction on each frame of the image feature in the video in the above manner can more effectively and richly extract the feature of each part of the target object in the image.

[0050] After that, Transformer is used to perform multi-scale and fine-grained feature extraction on the rough extracted features to obtain fine-grained features; specifically, Transformer performs feature interaction and fusion on the rough extracted features, and then performs multi-scale feature extraction to obtain globally fine-grained fused features:

[0051] Q = w * transformer(y, θ) + b

[0052] Among them, Q is the fine-grained feature; w is the adaptive weight value; transformer(·) is to obtain the set of pictures where the moving small target appears in each frame in the video segment p, and then perform feature transformation on each picture in the set of pictures through the convolution operator; θ is the learnable weight parameter of the Transformer model; b is the bias term of the network.

[0053] Such as Figure 2 shown, in implementation, the convolution kernels in the residual block of the scale adaptive residual neural network of this solution preferably adopt 1*3 and 3*1 convolution kernels. When using smaller convolution kernels for convolution operations, the small targets in the image can be feature-extracted pixel by pixel, and the number of parameters is reduced. The reason for the reduction in the number of parameters is explained as follows:

[0054] The two-dimensional convolution kernel size is p*q, the number of parameters is V, S represents the number of all feature maps in the previous layers, then the number of parameters for feature extraction in the current layer: V = p*q*S, and the number of parameters after scale adaptation is Y = (p + q)*S. Generally, the convolution kernel is larger than 2*2, and it can be inferred that Y is smaller than V.

[0055] In step S3, the Transformer and the multi-layer perceptron are used to detect the small targets in the fine-grained features, and the categories and detection frames of all small targets in the fine-grained features are obtained;

[0056] In an embodiment of the present invention, the step S3 further includes:

[0057] S31. Use the multi-head attention mechanism in the Transformer encoder to encode the fine-grained features:

[0058]

[0059] multi-Head(q, k, v) = Concat(head1,..., head i ) * W i

[0060]

[0061] Among them, Q is the fine-grained feature, q is the query vector during encoding; k is the key, v is the value calculated by the multi-head attention, are the weight values corresponding to the keys respectively; self-attention(·) is the self-attention calculation, head i is a subspace; n is the total number of subspaces; multi-Head(·) represents the merging of multiple heads; Wi is the weight of the model; w i is the weight value of the query vector, q transpose is the transpose of the query vector; k transpose is the matrix transpose of the query key, v transpose is the transpose of the vector value of the query;

[0062] S32. Use the multi-head attention mechanism in the transformer decoder to decode the entities queried in the image; The encoding and decoding processes of the transformer can be specifically referred to Figure 3 .

[0063] In Figure 3 , first input the features into the encoder improved by this solution. During the encoding process, optimize the feature part of the extracted feature matrix. The encoder is a stack of N identical layers. Each layer consists of two sub-layers. The first is the multi-head self-attention mechanism, and the second is a simple fully-connected feed-forward neural network. Finally, use the residual connection and normalization operation of ADD&Norm to prevent the occurrence of gradient disappearance and overfitting phenomena. The decoder is similar to the encoder and is also a stack of N identical layers, but uses a multi-head self-attention mask in the first layer to make it have richer feature information, thereby improving the accuracy of the network model.

[0064] S33. Use a multi-layer perceptron to map the output of the decoder to obtain the categories and detection boxes of all small targets in the image:

[0065] class, box m = w1 * (decoder(encoder(multi-Head(q, k, v)))) + b1

[0066] where class is the object category; box n is the detection box of the small target object; w1 is the weight of the multi-layer perceptron; encoder(.) is encoding, decoder(.) is decoding, and b1 is the bias term.

[0067] In step S4, calculate the similarity between the detection boxes of the same-category small targets in two adjacent frames of images, and determine the tracking targets in each frame of image based on the similarity. Among them, the calculation formula of the similarity is:

[0068] A = Jaccard_SIM(box n , box n-1 , T n )

[0069] where A is the similarity; Jaccard_SIM(.) is the similarity calculation function; box n and boxn-1 The detection boxes of small target objects in the nth frame and the (n - 1)th frame images respectively; T n =(t1, t2,..., t n ) is the time series of the video segment.

[0070] When this solution is applied to the monitoring of high-altitude parabolic objects, based on the obtained real-time tracking results of dynamic small targets, early warning information can be sent to the monitoring party in real time to timely remind pedestrians below the falling object, thereby reducing the impact caused by high-altitude parabolic objects.

[0071] In a second aspect, this solution provides a moving small target tracking system, which includes:

[0072] A video acquisition module, configured to acquire a video segment of a small target to be tracked and convert the video segment into a sequence of frame-by-frame images;

[0073] A feature extraction module, configured to perform rough feature extraction on each frame of image using a scale-adaptive residual neural network, and then use a Transformer to perform multi-scale and fine-grained feature extraction on the roughly extracted features to obtain fine-grained features;

[0074] A small target detection module, configured to detect small targets in the fine-grained features using a Transformer and a multi-layer perceptron to obtain the categories and detection boxes of all small targets in the fine-grained features;

[0075] A small target tracking module, configured to calculate the similarity between the detection boxes of the same-category small targets in two adjacent frames of images and determine the tracking targets in each frame of image based on the similarity.

[0076] In summary, through the strong feature extraction ability of the residual neural network and the pixel-by-pixel feature extraction of the Transformer in this solution, the feature information of small targets can be effectively extracted, and moving small target objects can be tracked more accurately.

Claims

1. A method for tracking small moving targets in a complex background, characterized in that, It includes the following steps: S1. Obtain the video segment of the small target to be tracked, and convert the video segment into a sequence of images frame by frame; S2. Coarsely extract features from each frame of image using a scale-adaptive residual neural network, and then use Transformer to perform multi-scale and fine-grained feature extraction on the coarsely extracted features to obtain fine-grained features; S3. Use Transformer and a multi-layer perceptron to detect small targets in the fine-grained features, and obtain the categories and detection frames of all small targets in the fine-grained features; S4. Calculate the similarity between the detection frames of small targets of the same category in two adjacent frames of images, and determine the tracking targets in each frame of image based on the similarity; The step S3 further includes: S31. Encode the fine-grained features using the multi-head attention mechanism in the Transformer encoder: multi-Head(q,k,v)=Concat(head1,……,head i )*W i Among them, Q is the fine-grained feature, q is the query vector during encoding; k is the key, v is the value calculated by multi-head attention, which are the weight values corresponding to the keys respectively; self-attention(·) is self-attention calculation, head i is a subspace; n is the total number of subspaces; multi-Head(·) represents the combination of multiple heads; W i is the weight of the model; w i is the weight value of the query vector, q transpose is the transpose of the query vector; k transpose is the matrix transpose of the query key, v transpose is the transpose of the vector value of the query; S32. Decode the entities in the query image using the multi-head attention mechanism in the transformer decoder; S33. Map the output of the decoder using a multi-layer perceptron to obtain the categories and detection frames of all small targets in the image: class, box m = w1 * (decoder(encoder(multi-Head(q, k, v)))) + b1 Among them, class is the object category; box m is the detection box of the small target object; m represents the number of target detection boxes, w1 is the weight of the multi-layer perceptron; encoder(.) is the encoding, decoder(.) is the decoding, and b1 is the bias term; The calculation formulas for coarse extraction and extraction of fine-grained features are respectively: y = SA-Resnet(p n + bias) + X Q = w * transformer(y, θ) + b where y is the feature after coarse extraction; SA-Resnet(·) is the scale-adaptive residual operation, bias is the dynamic adaptive operator; X is the region where the convolution kernel samples the feature map; Q is the fine-grained feature; w is the adaptive weight value; transformer(·) is to obtain the set of pictures of the moving small target in each frame from the video segment p, and then perform feature transformation on each picture in the set of pictures through the convolution operator; θ is the learnable weight parameter of the Transformer model; b is the bias term of the network; The calculation formula for similarity is: A = Jaccard_SIM(box n , box n-1 , T n ) Among them, A is the similarity; Jaccard_SIM(.) is the similarity calculation function; box n and box n-1 are the detection boxes of small target objects in the nth frame and the (n - 1)th frame images respectively; T n =(t1, t2, …, t n ) is the time series of the video segment.

2. The method for tracking small moving targets according to claim 1, wherein The convolution kernels in the residual blocks of the scale-adaptive residual neural network adopt 1*3 and 3*1 convolution kernels.

3. The method for tracking small moving targets according to any one of claims 1-2, characterized in that, It also includes real-time tracking results of small moving objects and real-time warning information is sent out.

4. A moving small target tracking system for the moving small target tracking method in the complex background described in claim 1, characterized in that, It includes: A video acquisition module for obtaining the video segment of the small target to be tracked and converting the video segment into a sequence of images frame by frame; A feature extraction module for coarsely extracting features from each frame of image using a scale-adaptive residual neural network, and then using Transformer to perform multi-scale and fine-grained feature extraction on the coarsely extracted features to obtain fine-grained features; A small target detection module for detecting small targets in the fine-grained features using Transformer and a multi-layer perceptron, and obtaining the categories and detection frames of all small targets in the fine-grained features; A small target tracking module for calculating the similarity between the detection frames of small targets of the same category in two adjacent frames of images and determining the tracking targets in each frame of image based on the similarity.

Citation Information

Patent Citations

  • Multi-target tracking method based on deep neural network

    CN113744316A

  • Transform and dense feature fusion-based remote sensing image change detection method and system

    CN115690002A