A Siamese Network Remote Sensing Target Tracking Method Based on Multi-Channel and Multi-Scale Fusion

Through the twin network remote sensing target tracking method with multi-channel and multi-scale fusion, the residual neural network and attention mechanism are used to solve the anti-interference problem of the remote sensing target tracking algorithm in complex backgrounds, and high-reliability remote sensing target tracking is achieved.

CN115984751BActive Publication Date: 2025-07-18XIAN INSTITUE OF SPACE RADIO TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310072530.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-12
Publication Date
2025-07-18
Estimated Expiration
2043-01-12

AI Technical Summary

Technical Problem

The existing remote sensing target tracking algorithm lacks anti-interference ability under complex backgrounds, especially in the global motion, lighting changes, noise, target scale changes and occlusion of the background in the remote sensing video image sequence.

Method used

A twin network remote sensing target tracking method with multi-channel multi-scale fusion is adopted to extract features through residual neural networks, combining attention mechanism and multi-scale fusion, feature extraction and target positioning are enhanced, and interference effects such as cloud occlusion are reduced.

Benefits of technology

It improves the robustness and accuracy of remote sensing target tracking, and can achieve high-reliability matching tracking in complex backgrounds to meet the real-time reconnaissance requirements on sensitive target stars.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115984751B_ABST
    Figure CN115984751B_ABST
Patent Text Reader

Abstract

The present invention provides a Siamese network remote sensing target tracking method based on multi-channel multi-scale fusion. First, the reference target image, the target image of the previous video frame, and the image to be tracked in the current frame are used as inputs of three channels, and features are extracted respectively through the proposed residual neural network; then the feature information of the three channels generates fusion information through a cross-correlation operation, and an attention mechanism is used to strengthen feature extraction; finally, the multi-scale method is used to analyze the target center position and the scaling factor respectively. The method of the present invention effectively reduces the influence of interferences such as cloud and fog occlusion when the remote sensing target is moving at high speed by using multi-channel input, attention mechanism and multi-scale fusion, has stronger robustness, can realize highly reliable matching and tracking of remote sensing targets under various complex backgrounds, and meets the on-star real-time high-reliability reconnaissance requirements for sensitive targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of aerospace remote sensing, relates to remote sensing target tracking, and specifically relates to a twin network remote sensing target tracking method based on multi-channel and multi-scale fusion. Background Art

[0002] Moving target detection and tracking is an emerging technology formed by the cross-integration of fields such as computer vision, remote sensing image processing, and artificial intelligence. Due to its great application value in fields such as real-time detection, real-time monitoring, and real-time control, it has become a research hotspot in recent years. At the same time, moving target detection and tracking is also an important research direction for space remote sensing earth observation. Especially for key targets such as airplanes, ships, and vehicles, there are extensive application requirements in both military and civilian fields and have extremely high application value.

[0003] Current target tracking algorithms are divided into traditional tracking algorithms and tracking algorithms based on twin networks. The advantage of traditional correlation filtering methods is fast speed, but the accuracy is average. The correlation filtering method combined with deep features uses a deep convolutional network to extract better features, and the accuracy has been greatly improved, but the speed has decreased and it is difficult to achieve real-time. Tracking algorithms based on the twin network structure use a special neural network structure, represented by SiamFC, which is characterized by receiving two images as inputs. This special network structure transforms the target tracking problem into a similarity learning problem, and well balances speed and accuracy.

[0004] High-resolution remote sensing time-series data can continuously record the dynamic information of moving targets in a large area with high precision and near real-time, which creates opportunities and conditions for accurately detecting and tracking moving targets. However, problems such as the global movement of the background, the illumination change of the background, background noise, the scale change of the target, and the occlusion of the target in the video image sequence will all interfere with the accurate tracking of the target for the detection accuracy of moving targets. Therefore, it is very necessary to conduct in-depth research on the tracking problem of remote sensing moving targets under complex backgrounds. Summary of the Invention

[0005] Aiming at the deficiencies existing in the prior art, the purpose of the present invention is to provide a twin network remote sensing target tracking method based on multi-channel and multi-scale fusion to solve the technical problem that the anti-interference ability of the target tracking method in the prior art needs to be further improved.

[0006] To solve the above technical problems, the present invention is implemented by adopting the following technical solutions:

[0007] A twin network remote sensing target tracking method based on multi-channel and multi-scale fusion includes the following steps:

[0008] Step 1: Given a video to be detected, determine the bounding box coordinate information of the remote sensing moving target in the search image to be detected in the first frame. The bounding box coordinate information includes position information and size information; intercept the remote sensing moving target as the target image; find the reference target image from the GF-3 remote sensing dataset; obtain the search image to be detected in the next frame.

[0009] Step 2: Take the reference target image, target image, and search image to be detected in the next frame obtained in Step 1 as inputs, and send them into the residual neural network respectively for siamese network feature extraction to obtain the depth feature Fx of the reference target image x, the depth feature Fy of the target image y, and the depth feature Fz of the search image to be detected z.

[0010] Step 3: Based on the depth features Fx, Fy, and Fz obtained in Step 2, obtain the multi-channel response map S multi , and then send S multi into the channel attention mechanism network to enhance the contour information of the image to obtain the weight w, and fuse the weight w with the multi-channel response map S multi through convolution to obtain the multi-channel response map S weighted by the attention mechanism weight-final .

[0011] Step 4: Perform multi-scale fusion processing on the multi-channel response map S weighted by the attention mechanism obtained in Step 3 to obtain the final response map S weight-final , and determine the position with the maximum response value in the final response map S final as the target center position s final . center

[0012] Taking the target center position s center as the center and using the multi-channel response map S weighted by the attention mechanism obtained in Step 3 weight-final as the input, construct a multi-scale fusion model to calculate the scaling factor λ, and obtain the corresponding side length according to the scaling factor λ, so as to achieve target tracking.

[0013] Compared with the prior art, the present invention has the following technical effects:

[0014] (Ⅰ) The present invention improves the confidence of target information by using the reference target image as the multi-channel target input information, assists in enhancing the feature information of the target image with the deep and shallow information of the reference target, strengthens the extraction of important details and contour information of the remote sensing target, and improves the anti-interference ability of the target tracking against adverse factors such as illumination, cloud and fog occlusion, and deformation caused by the rapid movement of the remote sensing target.

[0015] (II) The present invention improves the ability to extract details of the feature response map through an attention mechanism, and further reduces the impact on localization at different resolutions through multi-scale fusion, so as to quickly and accurately find the target center position. The scaling factor of the target area is solved through multi-scale fusion to improve the reliability of accurate area selection at different resolutions. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 FIG. is a flowchart of a siamese network remote sensing target tracking method based on multi-channel multi-scale fusion provided by an embodiment of the present invention.

[0017] Figure 2 FIG. is a structural diagram of feature extraction of a siamese network provided by an embodiment of the present invention.

[0018] Figure 3 FIG. is a schematic diagram of the fusion of coarse-grained and fine-grained with a target mask provided by an embodiment of the present invention.

[0019] Figure 4 FIG. is a schematic diagram of a feature enhancement network based on a reinforced attention mechanism network provided by an embodiment of the present invention.

[0020] Figure 5 FIG. is a schematic diagram of calculating the target center position based on multi-scale fusion provided by an embodiment of the present invention.

[0021] Figure 6 FIG. is a schematic diagram of calculating the scaling factor based on multi-scale fusion provided by an embodiment of the present invention.

[0022] The following further elaborates on the specific content of the present invention in conjunction with embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] It should be noted that all the devices and algorithms in the present invention, unless otherwise specified, all adopt the devices and algorithms known in the prior art.

[0024] To solve the problems of global motion of the background, illumination changes of the background, background noise, scale changes of the target, and occlusion of the target in remote sensing video image sequences, the present invention provides a twin network remote sensing target tracking method based on multi-channel multi-scale fusion. First, the reference target image, the target image of the previous frame of the video, and the image to be tracked in the current frame are used as the inputs of three channels, and the features are extracted respectively through the proposed residual neural network; then, the feature information of the three channels is used to generate fusion information through cross-correlation operation, and the attention mechanism is used to strengthen feature extraction; finally, the multi-scale method is used to analyze the target center position and the scaling factor respectively. The method of the present invention effectively reduces the influence of interference such as cloud occlusion when the remote sensing target is moving at high speed by using multi-channel input, attention mechanism and multi-scale fusion, has stronger robustness, can realize highly reliable matching and tracking of remote sensing targets under various complex backgrounds, and meets the on-orbit real-time and highly reliable reconnaissance requirements of sensitive targets.

[0025] In accordance with the above technical solutions, the following are specific embodiments of the present invention. It should be noted that the present invention is not limited to the following specific embodiments, and all equivalent transformations made on the basis of the technical solutions of the present application fall within the protection scope of the present invention.

[0026] Embodiment:

[0027] This embodiment provides a twin network remote sensing target tracking method based on multi-channel multi-scale fusion. As Figure 1 shown, the method includes the following steps:

[0028] Step 1: Given a video to be detected, determine the bounding box coordinate information of the remote sensing moving target in the search image to be detected in the first frame. The bounding box coordinate information includes position information and size information; intercept the remote sensing moving target as the target image; find the reference target image from the GF-3 (Gaofen-3 satellite) remote sensing dataset; obtain the search image to be detected in the next frame.

[0029] Step 1 includes:

[0030] Step 101: Read the remote sensing video sequence. In the search image to be detected in the first frame, the target area needs to be manually intercepted according to the scene. However, the size and angle of the target vary in different scenes. For the convenience of unified feature extraction, the target area needs to be appropriately adjusted and scaled, and the target area is adjusted to a standard square of 256×256 as the target image.

[0031] Step 102: Construct a GF-3 remote sensing dataset, manually divide the target area with a pixel size of M×M, and then uniformly adjust it to a square of 256×256 as the reference target image.

[0032] In step 102, the GF-3 remote sensing data set includes a remote sensing data set of fighter planes, passenger planes, transport planes, helicopters, submarines, aircraft carriers, destroyers, frigates, cars, armored vehicles, and tanks.

[0033] In step 102, when the reference target image cannot be found in the GF-3 remote sensing data set, the target image manually intercepted after adjustment in step 101 is used as the reference target image and added to the GF-3 remote sensing data set.

[0034] In step 103, the size of the search image to be detected in the next frame is also adjusted to a square of 512×512. This ensures that on the premise of sufficient clarity of the search image to be detected, the relative area ratio of the reference target and the remote sensing target is as large as possible, which is convenient for target tracking.

[0035] In step 2, the reference target image, the target image, and the search image to be detected in the next frame obtained in step 1 are used as inputs and respectively fed into the residual neural network for siamese network feature extraction to obtain the deep feature Fx of the reference target image x, the deep feature Fy of the target image y, and the deep feature Fz of the search image z to be detected.

[0036] Step 2 includes:

[0037] In step 201, in the residual neural network, the siamese network feature extraction part includes three inputs, a reference target image, a target image, and a search image to be detected in the next frame. The three inputs respectively correspond to three channels. The three channels share the same network model and weight parameters, and there is an interaction between the reference target image channel and the target image channel.

[0038] In this embodiment, since the reference target image and the target image of the previous frame have a strong correlation and can greatly improve the information volume of the target features, a particularly complex network model is not required. The present invention uses a 5-layer residual network for feature extraction to obtain the features of the three images.

[0039] In this embodiment, the network model is as Figure 2As shown in the figure, the network mainly consists of 5 residual modules. Each residual module is composed of a 1×1 convolutional layer, an n×n convolutional layer, and two activation layers relu. The first residual module has a 1×1 convolutional layer (convolutional stride is 1, number of channels is 32), a 11×11 convolutional layer (convolutional stride is 2, number of channels is 32), and two relu layers. The second residual module has a 1×1 convolutional layer (convolutional stride is 1, number of channels is 64), a 9×9 convolutional layer (convolutional stride is 2, number of channels is 64), and two relu layers. The third residual module has a 1×1 convolutional layer (convolutional stride is 1, number of channels is 128), a 7×7 convolutional layer (convolutional stride is 2, number of channels is 128), and two relu layers. The fourth residual module has a 1×1 convolutional layer (convolutional stride is 1, number of channels is 64), a 5×5 convolutional layer (convolutional stride is 2, number of channels is 64), and two relu layers. The fifth residual module has a 1×1 convolutional layer (convolutional stride is 1, number of channels is 32), a 3×3 convolutional layer (convolutional stride is 2, number of channels is 32), and two relu layers.

[0040] In step 201, in the residual neural network, in each residual module, the input and output are as shown in the following formula: M2 = g2(f2(g1(f1(M1))) + M1)

[0041] In the formula:

[0042] M2 is the output of the residual module;

[0043] M1 is the input of the residual module;

[0044] f1 is a 1×1 convolutional function;

[0045] f2 is an n×n convolutional function;

[0046] g1 is a relu function;

[0047] g2 is a relu function.

[0048] In step 202, the image detail features of the reference target mask are prominent, and the detail features of the target are enhanced by using the high-quality image information of the reference target. As Figure 3 shown, the present invention fuses the coarse-grained and fine-grained of the reference target mask with the target mask, and performs weighted summation on the target output features of each layer to obtain the corrected features, which is expressed as

[0049] Fy n (x,y) = Fy n (x,y) + λ·Fx n (x,y)

[0050] In the formula:

[0051] Fy n (x, y) represents the output of the target feature at the coordinate (x, y) after the output of the nth convolutional layer;

[0052] Fx n (x, y) represents the output of the reference target at the coordinate (x, y) after the nth convolutional layer;

[0053] λ represents the influence weight of the reference target on the target mask.

[0054] In this embodiment, the contour features of the reference target have a certain promoting effect on target localization, and the texture features have a more obvious promotion effect on the segmentation accuracy of the target. Generally, the selected reference image is not interfered and retains the image content information well. Therefore, considering the possible interference situation in the video in the present invention, the details of the network outputs at different levels are improved. The weight λ value varies at different levels, and the deeper the level, the greater the influence, as shown in Table 1.

[0055] Table 1 Values of weight λ at different levels

[0056]

[0057] Step 3: Based on the depth features Fx, Fy, and Fz obtained in Step 2, obtain the multi-channel response map S multi , and then input S multi into the channel attention mechanism network to enhance the contour information of the image to obtain the weight w, and multiply the weight w with the multi-channel response map S multi through convolution fusion to obtain the multi-channel response map S weighted by the attention mechanism weight-final .

[0058] Step 3 includes:

[0059] Step 301: The object tracking algorithm based on the Siamese network formulates the object tracking problem as a template matching problem using the cross-correlation operation.

[0060] Step 30101: After the reference target image x and the search image z to be detected in the current frame pass through the feature extraction network to extract the corresponding depth features Fx and Fz, perform the cross-correlation operation on the obtained depth features Fx and Fz, and then output a response map to accurately locate the position of the tracking target.

[0061] In this embodiment, it is expressed as:

[0062] S multi-x = Fx * Fz + b * I

[0063]

[0064]

[0065] In the formula,

[0066] represents the feature extraction network;

[0067] b represents the offset value

[0068] I represents the standard matrix, and each element in the matrix takes the value of 1;

[0069] * represents the cross-correlation operation, which refers to the convolution operation in the present invention;

[0070] S multi-x represents the cross-correlation response map of the depth features Fx and Fz.

[0071] Step 30102, similarly, after the target image y of the previous frame and the search image z to be detected in the current frame pass through the feature extraction network to extract the corresponding depth features Fy and Fz, the obtained depth features Fy and Fz are subjected to cross-correlation operation, and then a response map is output to accurately locate the position of the tracking target.

[0072] In this embodiment, it is expressed as:

[0073]

[0074] Step 30103, the two response maps are merged to obtain the final multi-channel response map, which is expressed as:

[0075] S multi =[S multi-x ,S multi-y

[0076] In the formula:

[0077] S multi represents the multi-channel response map;

[0078] S multi-x represents the cross-correlation response map of the depth features Fx and Fz;

[0079] S multi-y represents the cross-correlation response map of the depth features Fy and Fz.

[0080] Step 302, the remote sensing video has a complex background and many interfering targets, which causes the target to drift during the tracking process, and the drift amount accumulates over time, resulting in tracking failure. In order to improve the anti-interference robustness, it is necessary to perform feature enhancement on the merged multi-channel response map through the attention mechanism.

[0081] As Figure 4 shown, construct the channel attention mechanism network H, and use the multi-channel response map S multi ​Enhance the contour information of the image through network H, making the target position not easily deviate, and at the same time, this network is designed to be lightweight, expressed as,

[0082] w = H(S multi )

[0083] In the formula:

[0084] w represents the weight;

[0085] H represents the channel attention mechanism network;

[0086] Step 303, multiply the weight w enhanced by the attention mechanism with the feature response map S multi Through convolution fusion, obtain the multi-channel response map S weighted by the attention mechanism weight-final ,, expressed as,

[0087] S weight-final = w * S multi。

[0088] Step 4, perform multi-scale fusion processing on the multi-channel response map S weighted by the attention mechanism obtained in Step 3 to obtain the final response map S weight-final , and determine the position with the maximum response value in the final response map S final as the target center position s final . center .

[0089] Taking the target center position s center as the center and using the multi-channel response map S weighted by the attention mechanism obtained in Step 3 weight-final as the input, construct a multi-scale fusion model to calculate the scaling factor λ, and obtain the corresponding side length according to the scaling factor λ, so as to achieve target tracking.

[0090] In Step 4, it includes:

[0091] Step 401, perform multi-scale fusion on the multi-channel response map S weighted by the attention mechanism obtained in Step 3 to generate the final feature response map S weight-final , calculate the maximum value in the final response map S final , and the corresponding point is the target center position s final . center .

[0092] Specifically in this embodiment, such as Figure 5As shown in the figure, the multi-scale fusion localization network G consists of a lightweight neural network. In the present invention, fusion networks with multi-scale factors of 3×3, 5×5, 7×7, and 9×9 are used respectively. Since the convolutional kernel sizes of each convolutional module are different and the output dimensions of the convolutional modules are different, the outputs of the sub-networks with 5×5, 7×7, and 9×9 need to be upsampled to the dimension of the 3×3 sub-network. After the outputs of the 4 scales are added together, they are extracted through 2 convolutional layers of 3×3, and finally, after Softmax processing, the final feature response map S is obtained. final . Such processing can improve the accuracy of the target center position from multiple dimensions, and this processing can be expressed as

[0093] S final = G(S weight-final ).

[0094] Step 402, after obtaining the target center position, use the multi-scale method to determine the range of the target area.

[0095] Specifically in this embodiment, as Figure 6 shown in the figure, the present invention designs a multi-scale network to obtain the scaling factor λ. Considering the influence of the actual size of the target area on the area scale, the present invention uses fusion networks with multi-scale factors of 3×3, 5×5, 7×7, and 9×9 respectively for the final feature extraction. Among them, the 3×3 convolution is used 4 times, the 5×5 convolution is used 3 times, the 7×7 convolution is used 2 times, and the 9×9 convolution is used 1 time. After the outputs of the 3×3, 5×5, 7×7, and 9×9 sub-networks, the image feature dimensions are the same and no upsampling is required. After the outputs of the 4 scales are added together, they are extracted through 1 convolutional layer of 3×3, and then through a reshape operation, they are integrated into 1 two-dimensional vector. Finally, through 1 sigmoid function, a scaling factor λ is obtained. This processing can be expressed as: λ = K(S weight-final ).

[0096] This method is relatively simple and does not require the length, width, or angle information of the target position. Only scaling is required according to the scaling factor with 256 as the reference. The target is uniformly defined as a square, and the side length M is defined as

[0097] M = λ * 256.

[0098] Step 403, perform image patch search for each scaling factor λ. First, find the image patches corresponding to each scaling factor λ, extract them as the original image slices, and then scale them to the size of 256×256. All the target images have been adjusted to the size of 256×256. Only the difference DIFF between the adjusted extracted image patch D and the adjusted target image M needs to be calculated, which is expressed as

[0099]

[0100] In the formula:

[0101] m and n respectively represent the m-th row pixel points and the n-th column pixel points in the image.

[0102] When the minimum difference DIFF min is obtained, the corresponding scaling factor λ min is the scaling value λ of the current frame current .

[0103] Step 5: The training of the model is independent of the tracking process. After the training is completed, the model is directly used in the tracking process.

[0104] Specifically in this embodiment, step 5 includes:

[0105] Step 501: Select 200 videos from the remote sensing dataset. The input of each network model consists of a reference target image, the target image of the previous frame, and the search image of the current frame. The target area determined in the previous frame is used as the target image of the current frame. If there is no reference target image in the dataset, manually intercept it from the first frame image as the reference target image.

[0106] Step 502: In order to improve the generalization ability of the search, the search image is damaged to a certain extent. 5% salt-and-pepper noise is randomly added to the video library images and 3×3 blur processing is performed.

[0107] Step 503: The model is trained iteratively 100 times. The number of samples selected for one training is 16. The learning rate of the training decays exponentially from 10 -2 to 10 -5 . The loss function L of the training is a weighted difference loss function, as shown in the following formula:

[0108] L = ||DIFF||2 + α||W||2

[0109] In the formula:

[0110] DIFF represents the difference between the corresponding pixels of the scaled current frame and the previous frame image;

[0111] W represents the network parameter value;

[0112] α represents the weight of the network parameter.

[0113] Application example:

[0114] This application example presents a Siamese network remote sensing target tracking method based on multi-channel multi-scale fusion according to the above embodiments. In this application example, a remote sensing video database is constructed, which consists of self-owned GF-3 videos and UAV videos. There are a total of 200 videos, including three types of targets: airplanes, ships, and vehicles. A reference remote sensing target dataset is constructed, which includes a total of 1000 sensitive targets, and all the targets in the video library can be found in the reference target dataset. Randomly select 180 videos as the training set, and the remaining 20 videos as the test set. In this embodiment, AUC (Accuracy) and EAO (Expected Average Overlap) are used as evaluation criteria to compare the algorithm performance. AUC represents the accuracy of the tracking and positioning center, and EAO represents the average target coverage rate. Only when both AUC and EAO are high can it indicate that the algorithm can not only accurately track but also has strong robustness. The comparison between this embodiment and the basic method SiamFC is shown in Table 2 as follows.

[0115] Table 2 Comparison results of the embodiment and the basic method in the test set

[0116]

[0117] From the comparison in Table 2, it can be seen that by using the method of this invention's embodiment, in the remote sensing dataset, there is a performance improvement of more than 10% in both AUC and EAO of this invention. This shows that this invention has good guarantees in both accuracy and robustness. The reason is that this invention introduces a reference target image as a guide to ensure that even if interference appears in the video in a timely manner, it can still have a strong reference effect on target tracking. At the same time, a multi-scale fusion network is introduced respectively to achieve positioning and region selection, ensuring that the target box can lock the target.

Claims

1. A twin network remote sensing target tracking method based on multi-channel and multi-scale fusion, characterized in that It includes the following steps: Step 1: Given a video to be detected, determine the bounding box coordinate information of the remote sensing moving target in the search image to be detected in the first frame. The bounding box coordinate information includes position information and size information; intercept the remote sensing moving target as the target image; find the reference target image from the GF-3 remote sensing dataset; obtain the search image to be detected in the next frame; Step 2: Use the reference target image, target image, and search image to be detected in the next frame obtained in Step 1 as inputs, and send them into the residual neural network respectively for siamese network feature extraction to obtain the depth feature Fx of the reference target image x, the depth feature Fy of the target image y, and the depth feature Fz of the search image z to be detected; Step 3: Based on the depth features Fx, Fy, and Fz obtained in Step 2, obtain the multi-channel response map S multi , and then input S multi into the channel attention mechanism network to enhance the contour information of the image to obtain the weight w, and multiply the weight w with the multi-channel response map S multi Through convolutional fusion, obtain the multi-channel response map S weighted by the attention mechanism weight-final ; Step 4: Perform multi-scale fusion processing on the multi-channel response map S weighted by the attention mechanism obtained in Step 3 weight-final to obtain the final response map S final , and determine the position with the maximum response value in the final response map S final as the target center position s center ; With the target center position s center as the center, using the multi-channel response map S weighted by the attention mechanism obtained in step 3 weight-final as the input, a multi-scale fusion model is constructed to calculate the scaling factor λ, and the corresponding side length is obtained according to the scaling factor λ, so as to achieve target tracking.

2. The twin network remote sensing target tracking method based on multi-channel and multi-scale fusion according to claim 1, characterized in that, Step 1 includes: Step 101: Read the remote sensing video sequence. In the search image to be detected in the first frame, manually intercept the target area according to the scene, and adjust the target area to a standard square of 256×256 as the target image; Step 102: Build the GF-3 remote sensing dataset, manually divide the target area with a pixel size of M×M, and then uniformly adjust it to a square of 256×256 as the reference target image; Step 103: Also adjust the size of the search image to be detected in the next frame to a square of 512×512.

3. The method for remote sensing target tracking based on a multi-channel multi-scale fusion Siamese network according to claim 2, wherein In Step 102, the GF-3 remote sensing dataset is a remote sensing dataset including fighter jets, passenger planes, transport planes, helicopters, submarines, aircraft carriers, destroyers, frigates, cars, armored vehicles, and tanks.

4. The method for remote sensing target tracking based on a multi-channel and multi-scale fusion Siamese network according to claim 2, wherein, In Step 102, when the reference target image cannot be found in the GF-3 remote sensing dataset, use the target image manually intercepted and adjusted in Step 101 as the reference target image and add it to the GF-3 remote sensing dataset.

5. The twin network remote sensing target tracking method based on multi-channel multi-scale fusion according to claim 1, characterized in that, Step 2 includes: Step 201: In the residual neural network, the siamese network feature extraction part includes three inputs, a reference target image, a target image, and a search image to be detected in the next frame. The three inputs correspond to three channels respectively. The three channels share the same network model and weight parameters, and there is an interaction between the reference target image and the target image channels; Step 202: Fuse the coarse-grained and fine-grained of the reference target mask with the target mask, and perform weighted summation on the target output features of each layer to obtain the corrected features, expressed as Fy n (x,y) = Fy n (x,y) + λ·Fx n (x,y) In the formula: Fy n (x, y) represents the output of the target feature at the coordinate (x, y) after the output of the nth convolutional layer; Fx n (x, y) represents the output of the reference target at the coordinate (x, y) after the nth convolutional layer; λ represents the influence weight of the reference target on the target mask.

6. The multi-channel multi-scale fusion-based Siamese network remote sensing target tracking method according to claim 5, wherein In Step 201, in the residual neural network, in each residual module, the input and output are as shown in the following formula M2 = g2(f2(g1(f1(M1))) + M1) In the formula: M2 is the output of the residual module; M1 is the input of the residual module; f1 is a 1×1 convolution function; f2 is an n×n convolution function; g1 is the relu function; g2 is the relu function.

7. The twin network remote sensing target tracking method based on multi-channel and multi-scale fusion according to claim 1, characterized in that, Step 3 includes: Step 301: The object tracking algorithm based on the siamese network formulates the object tracking problem as a template matching problem using the cross-correlation operation; Step 30101, after the depth features Fx and Fz corresponding to the reference target image x and the search image z to be detected in the current frame are extracted through the feature extraction network, the obtained depth features Fx and Fz are subjected to cross-correlation operation, and then a response map is output to accurately locate the position of the tracking target; Step 30102, after the depth features Fy and Fz corresponding to the target image y in the previous frame and the search image z to be detected in the current frame are extracted through the feature extraction network, the obtained depth features Fy and Fz are subjected to cross-correlation operation, and then a response map is output to accurately locate the position of the tracking target; Step 30103, the two response maps are merged to obtain the final multi-channel response map, expressed as: S multi = [S multi-x , S multi-y ​ In the formula: S multi represents a multi-channel response map; S multi-x Represents the cross-correlation response map of the depth features Fx and Fz; S multi-y Represents the cross-correlation response map of the depth features Fy and Fz; Step 302: Construct a channel attention mechanism network H, and apply the multi-channel response map S multi to enhance the contour information of the image through network H, so that the target position is not easily deviated, which is expressed as w = H(S multi ) In the formula: w represents the weight; H represents the channel attention mechanism network; Step 303: Multiply the weight w enhanced by the attention mechanism with the feature response map S multi Through convolutional fusion, obtain the multi-channel response map S weighted by the attention mechanism weight-final , denoted as S weight-final = w * S multi .

8. The method for remote sensing target tracking based on a multi-channel multi-scale fusion Siamese network according to claim 1, characterized in that, Step 4 includes: Step 401: Perform multi-scale fusion on the multi-channel response map S weighted by the attention mechanism obtained in Step 3 weight-final to generate the final feature response map S final , and calculate the maximum value in the final response map S final . The corresponding point is the target center position s center ; Step 402, after obtaining the target center position, use the multi-scale method to determine the range of the target area; Step 403, perform image patch search for each scaling factor λ. First, find the image patch corresponding to each scaling factor λ, extract it as the original image slice, and then scale it to a size of 256×256. All target images have been adjusted to a size of 256×256. Only calculate the difference DIFF between the adjusted extracted image patch D and the adjusted target image M, expressed as: In the formula: m and n respectively represent the m-th row pixel point and the n-th column pixel point in the image; When the minimum difference DIFF is obtained min the corresponding scaling factor λ min is the scaling value λ of the current frame current .

Citation Information

Patent Citations

  • Remote sensing video object tracking method based on JCFNet network

    CN109242884A

  • Depth correlation target tracking algorithm based on mutual reinforcement and multi-attention mechanism learning

    CN110120064A