An object tracking method based on a deep learning Siamese network
By adopting a deep learning twin network method in video target tracking, multi-order depth separation convolution network and multi-scale large core attention module enhance feature extraction, the accuracy problem in complex scenarios in video target tracking is solved, and a more efficient and robust target tracking effect is achieved.
Patent Information
- Application Number
- CN202510318927.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-18
AI Technical Summary
In the video target tracking, it is difficult to effectively deal with the pollution superposition problem of scenes such as the target being blocked, severely deformed or out of the field of view. It is difficult to distinguish the target from the background in the chaotic background scene, resulting in low tracking accuracy.
The target tracking method based on deep learning twin network is adopted, and the local and global context information can be captured through multi-order depth separation convolutional networks, combined with multi-scale large-core attention modules and gated space attention modules to enhance feature extraction capabilities, and the classification regression network is used to optimize the prediction of the target bounding box.
It improves the accuracy and robustness of video target tracking in complex scenarios, can more effectively deal with challenges such as target occlusion, deformation and field of view, and maintains good tracking performance in background messy scenarios.
Smart Images

Figure CN119851184B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method in the field of object tracking, and more particularly to an object tracking method based on a deep learning Siamese network. Background Art
[0002] Video object tracking is one of the important basic research issues in the field of computer vision. Video object tracking means that after specifying an object in the first frame of a video sequence (giving the object bounding box, usually represented by a rectangular box), the object specified in the first frame is continuously tracked in subsequent frames, and the object is calibrated using the object bounding box to achieve object positioning and scale estimation. Video object tracking plays a very important role in fields such as intelligent driving and video surveillance. For example, video object tracking can assist in planning the route of an autonomous driving vehicle, can detect objects in video surveillance, and predict future abnormal situations and possible safety hazards, providing theoretical guidance for security prevention.
[0003] In the process of video object tracking, traditional algorithms based on correlation filtering run relatively fast and have high stability. However, compared with a series of deep learning algorithms based on Siamese networks, traditional algorithms are less accurate. Video object tracking is a major research hotspot in the field of computer vision currently. Although significant progress has been made in recent years, this research direction still has great challenges: 1) How to better cope with challenges such as object occlusion, severe deformation, or out-of-field-of-view, and solve the problem of pollution superposition generated by the tracking model in scenarios where the object is occluded, severely deformed, and out-of-field-of-view in video object tracking; 2) How to effectively improve the discriminative ability of Siamese network algorithms, that is, when the network judges similar objects, for example, when multiple people or multiple vehicles appear in the scene, there are often situations where the confidence of similar objects in the feature layer is higher than that of the tracking object, resulting in the model tracking the wrong object, making it difficult to capture the object information for continuous tracking, or causing tracking failure due to misjudging similar objects; 3) In scenes with cluttered backgrounds, the cluttered background often pollutes the tracking model, making it difficult to distinguish the object from the background and effectively extract the object feature information, thereby resulting in the loss of the object in subsequent frames. And the related algorithms based on Siamese networks were first applied to tasks such as template matching and similarity measurement. In recent years, with the powerful performance shown by deep learning algorithms in the field of computer vision, the algorithms based on deep learning Siamese networks have a relatively high computational speed compared to other deep learning algorithm frameworks. Therefore, researching object tracking algorithms based on deep learning Siamese networks has important value and significance. Summary of the Invention
[0004] The object of the present invention is to solve the deficiencies existing in the prior art that the tracking model will produce pollution superposition in scenarios such as when the target is occluded, severely deformed, or out of the field of view, it is difficult to distinguish similar targets, and the cluttered background is prone to pollute it. A target tracking method based on a deep learning Siamese network is provided, aiming to improve the tracking and positioning accuracy of the target in difficult target tracking scenarios and provide technical support for video target tracking.
[0005] To achieve the above object, the technical solution provided by the present invention is as follows:
[0006] A target tracking method based on a deep learning Siamese network, which is characterized in that it includes the following steps:
[0007] S1. Obtain a video image file, specify a template target in the initial frame of the video image file, and give the ground truth of its target bounding box. The template target is used to provide a discrimination template for subsequent frames;
[0008] S2. Extract the feature information of the initial frame, process the feature information of the current frame through a depthwise separable convolutional network, obtain the corresponding local and global context information, and then perform feature enhancement to obtain the enhanced feature of the template target ;
[0009] S3. Extract the subsequent frame adjacent to the current frame as the current frame, extract its feature information, process the feature information of the current frame through the depthwise separable convolutional network to obtain the corresponding local and global context information, and then perform feature enhancement to obtain the enhanced feature of the subsequent frame ;
[0010] S4. Perform an association operation on the enhanced feature of the subsequent frame and the enhanced feature of the template target to obtain a strongly associated region image and its corresponding correlation response map and loss function;
[0011] S5. Process the strongly associated region image through a classification regression network, iteratively output offset regression, and then obtain the optimized correlation response map, target bounding box and corresponding loss function of the subsequent frame, and complete the target tracking of the current frame;
[0012] S6. Determine whether the current frame is the last frame. If not, return to step S3. If so, output the target bounding boxes of all subsequent frames to complete the video target tracking.
[0013] Furthermore, in steps S2 and S3, the feature enhancement is specifically performed by iteratively enhancing through m multi-scale large kernel modules, where m≥2; the multi-scale large kernel module includes a multi-scale large kernel attention module and a gated spatial attention module arranged in sequence. The process of the multi-scale large kernel attention module is as follows:
[0014]
[0015]
[0016] The process of the gated spatial attention module is as follows:
[0017]
[0018]
[0019] Among them, is the feature obtained by the depthwise separable convolutional network, LN is the normalization operation, and N is the feature is the feature obtained after normalization, and L is the feature is the feature obtained after normalization; are respectively the learnable scaling factors in the current layer, is the outer product of matrices, is the nth pointwise convolution under the invariant dimension, where n takes values of 1, 2, 3, 4, 5, 6, denoted as , is the dynamically adaptive large kernel attention, is the feature after dense transformation and the feature , is the output feature of the multi-scale large kernel attention module, is the enhanced feature output by the multi-scale large kernel module.
[0020] Furthermore, in steps S2 and S3, the is obtained by the following method:
[0021] The feature obtained by processing the feature N through the first pointwise convolution is enhanced by the group multi-scale mechanism for fixed large kernel attention processing to obtain the large kernel attention of the feature , as follows:
[0022]
[0023] Among them, is the depth convolution, is the dilated convolution with a depth of D, is the pointwise convolution;
[0024] For the ith-dimensional feature of the feature the large kernel attention is dynamically adapted to obtain the ith-dimensional dynamic feature , as follows:
[0025]
[0026] Among them, is the space gate;
[0027] Then, aggregate the dynamic features to obtain dynamically adaptive large kernel attention .
[0028] Furthermore, in steps S2 and S3, the is specifically obtained through the following method:
[0029] .
[0030] Furthermore, in steps S2 and S3, the depthwise separable convolutional network includes a feature decomposition module and a multi-order gating aggregation module; input the feature information into the feature decomposition module, and process it through normalization, a convolutional layer , a depth convolutional layer and a first activation function to respectively obtain low-order features , middle-order features and high-order features using different scaling ratios, specifically:
[0031]
[0032]
[0033] Among them, X is the feature information, GAP is global average pooling, is normalization, is a scaling factor initialized to zero, is circular convolution;
[0034] The multi-order gating aggregation module is provided with an aggregation branch and a context branch. The aggregation branch is used to generate gating weights according to the input feature X in . The context branch is used to perform multi-scale feature extraction using circular convolutions of different sizes according to the generated gating weights, thereby capturing context multi-order interactions. The low-order features , middle-order features and high-order features are used as the input feature X in input into the multi-order gating aggregation module. After being processed by a convolutional layer and a first activation function , and guided by the aggregation branch and the context branch at the same time, the feature Z is obtained as the output of the output feature X out .
[0035] Further, in step S5, the loss function is denoted as , and is obtained by the following formula:
[0036]
[0037] where is an indicator function. If holds, then the value of the function is 1; otherwise, it is 0. is the loss of the classification branch, and the binary cross - entropy loss is adopted. is the actual position coordinate of the positive sample. indicates whether it is a positive sample. If it is a positive sample, it takes 1; if it is a negative sample, it takes 0. is the loss of the regression branch, and the IOU loss is adopted. is the ground - truth coordinate of the target position. is the normalized ground - truth coordinate of the target position. is the total number of positive samples. is a modifiable parameter. is the predicted target position.
[0038] Further, in step S5, takes the value of 0.5.
[0039] The present invention also provides a computer program product, including a computer program, characterized in that: when the program is executed by a processor, it implements the steps of the above - mentioned target tracking method based on a deep - learning Siamese network.
[0040] Meanwhile, the present invention also provides a computer - readable storage medium, on which a computer program is stored, characterized in that: when the program is executed by a processor, it implements the steps of the above - mentioned target tracking method based on a deep - learning Siamese network.
[0041] Advantages of the present invention:
[0042] 1. The target tracking method based on a deep - learning Siamese network of the present invention sets up a multi - stage depth - separable convolutional network to capture local and global context information, effectively capturing context information at different levels and improving the limitations of traditional convolutional neural networks in context feature encoding.
[0043] 2. The target tracking method based on a deep - learning Siamese network of the present invention sets up a multi - scale large - kernel attention module and a gated spatial attention module to enhance the feature extraction ability in the target tracking method, so as to achieve better information aggregation and detail restoration.
[0044] 3. The object tracking method based on the deep learning Siamese network of the present invention obtains the relevant response map of the subsequent frame through the classification regression network, and outputs additional offset regression through the classification branch and the regression branch to refine the position prediction of the object bounding box.
[0045] 4. The object tracking method based on the deep learning Siamese network of the present invention adopts the deep learning Siamese network, making the object tracking accurate, efficient and flexible, having better performance and algorithm robustness than the traditional correlation filter algorithm, and better coping with challenges such as object occlusion, severe deformation, out of the field of view, motion blur, etc., providing technical support for video object tracking.
[0046] 5. The present invention also provides a computer program product and a computer-readable storage medium capable of executing the above method steps, which can promote the application of the method of the present invention and realize object tracking on the corresponding hardware device. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is the workflow diagram of the embodiment of the present invention;
[0048] Figure 2 is the structural schematic diagram of the feature decomposition module in the embodiment of the present invention;
[0049] Figure 3 is the structural schematic diagram of the multi-scale large kernel module in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] This method captures local and global context information by designing a multi-stage depthwise separable convolutional network; then introduces a multi-scale large kernel attention module and a gated spatial attention module to enhance the feature extraction ability in the image super-resolution task, so as to achieve better information aggregation and detail restoration; in the subsequent frame images, perform a cross-correlation operation between the template features and the features of the search area, and crop to obtain the strongly correlated area image; finally, input the strongly correlated area image into a classification regression network, and obtain the relevant response map in the subsequent frame through convolutional operations. The following combines specific embodiments to specifically describe a method for object tracking based on a deep learning Siamese network of the present invention. As Figure 1 shown is the workflow of the embodiment of the present invention, which specifically includes the following steps:
[0051] Step 1, acquisition and preprocessing of video image files.
[0052] Video object tracking reads and recognizes each frame of a video image file, which can come from a public dataset or a non-public dataset. In this embodiment, an existing public dataset is adopted, and one of the datasets such as OTB100, dtb70, VOT, UAV123, etc. is selected and cropped. The dataset includes multiple video image files, each of which corresponds to a video sequence. Multiple images in the same video sequence should have the same image resolution, and the image resolutions of different video sequences can be different. At the same time, a template target is specified in the image of the initial frame of each video sequence, and the ground truth of its target bounding box is given for feature extraction of the target in the initial frame. The template target provides a discrimination template for subsequent frames. The target bounding box includes the x coordinates and y coordinates of the four vertices of a rectangle, where x represents the length direction of the image and y represents the height direction of the image.
[0053] In other embodiments of the present invention, for video object tracking using a non-public dataset, like using a public dataset, first, the cropping work of each frame of the picture is carried out, and the template target in the initial frame and the ground truth of its target bounding box are given.
[0054] Step 2: Extract the subsequent frame immediately following the video image file as the current frame, and extract feature information in the current frame. Specifically, a multi-stage depthwise separable convolutional network is designed to capture local and global context information. The depthwise separable convolutional network includes two modules: a feature decomposition module and a multi-stage gated aggregation module. The specific details are as follows.
[0055] 2.1 Feature decomposition
[0056] The feature decomposition module is as Figure 2 shown. To make the network pay more attention to multi-stage interactions, the feature information X of the current frame in each video image file is extracted and input into the feature decomposition module, where , is a real number matrix composed of three dimensions: the length C, height H, and width W of the video image file. Through normalization (Norm), convolutional layer , depth convolutional layer and the first activation function processing, low-order features , middle-order features and high-order features are obtained by processing with different scaling ratios, so as to enhance the expression ability of the features. Among them, low-order features refer to fine-grained local textures in the image, middle-order features refer to complex global shapes, and high-order features refer to the features obtained by convolutional operations on the feature information X, low-order features and middle-order features . The detailed operations are as follows:
[0057]
[0058]
[0059] Among them, GAP is global average pooling, is a scaling factor initialized to zero, is circular convolution.
[0060] 2.2 Multi-order gated aggregation
[0061] An aggregation branch and a context branch are also provided in the multi-order gated aggregation module. As Figure 2 shown, the aggregation branch (Subtract) is used to generate gated weights based on the input feature X in The context branch is used to perform multi-scale feature extraction using circular convolutions of different sizes based on the generated gated weights, thereby capturing context multi-order interactions. The low-order features , middle-order features and high-order features output in step 2.1 are respectively used as the input feature X in input into the multi-order gated aggregation module. After passing through the convolutional layer and the first activation function processing, guided by the aggregation branch and the context branch, the feature Z is obtained as the output feature X out output, as shown specifically below:
[0062]
[0063]
[0064] Among them, is the second activation function, is circular convolution, Concat means merging the low-order features , middle-order features and high-order features and then performing depth convolution to exchange information, is the aggregated feature of the low-order, middle-order and high-order features.
[0065] In the present invention, in order to adaptively aggregate the features extracted from the context branch, the outputs of the aggregation branch and the context branch are processed using the second activation function SiLU. The second activation function SiLU has both the Sigmoid gating effect and stable training characteristics. The processing process is as shown below:
[0066]
[0067] Among them, u is in the above formula or ;
[0068] Step 3: As shown in Figure 3 , perform feature enhancement to improve the feature expression ability of the Siamese network. To better achieve information aggregation and detail restoration, the feature enhancement of the present invention iteratively enhances the feature Z through m multi-scale large kernel modules to improve the feature expression ability in video object tracking, where m≥2 and can be selected according to actual requirements for iterative enhancement processing until convergence to a pre-set algorithm accuracy parameter.
[0069] The multi-scale large kernel module includes a multi-scale large kernel attention module and a gated spatial attention module arranged in sequence, and the specific processing process is as follows:
[0070]
[0071] Among them, LN is the normalization operation, N is the feature obtained after normalization, L is the feature obtained after normalization, are respectively learnable scaling factors in the current layer, is the outer product of matrices, is the nth point-level convolution under the invariant dimension, and n takes values of 1, 2, 3, 4, 5, 6, denoted as in sequence, is the dynamically adaptive large kernel attention, is the feature after dense transformation and the feature , is the output feature of the multi-scale large kernel attention module, is the enhanced feature output by the multi-scale large kernel module.
[0072] is obtained through the following method:
[0073] The attention mechanism adopted by the multi-scale large kernel attention module allows the Siamese network to focus on key information and ignore irrelevant information. The feature obtained by processing the feature N through the first point-level convolution is input. In order to learn the attention map of full-scale information, a group multi-scale mechanism is used to enhance the fixed large kernel attention (LKA, Large kernel attention) to process it, and the large kernel attention of the feature is obtained as follows:
[0074]
[0075] Among them, is the depth convolution, is the dilated convolution with depth D, is the pointwise convolution.
[0076] For the feature of the i-th dimension, perform dynamic adaptation to obtain the dynamic feature of the i-th dimension , and then aggregate the dynamic features to obtain the dynamically adaptive large-kernel attention . Specifically, it is obtained through the following method: To avoid blocking effects and learn more local information, use the spatial gate to dynamically adapt the large-kernel attention of the i-th dimension feature to the dynamic feature as follows:
[0077]
[0078] where, is the spatial gate of, is the outer product of matrices.
[0079] To obtain spatial information more effectively, use the gated spatial attention module (GSAU) to first perform a single-layer depth convolution on the output feature to obtain the normalized feature L, and then perform the 4th pointwise convolution and the 5th pointwise convolution on the feature L to obtain the features and the feature respectively. Then, through the dense transformation, weight the features and the feature to obtain the feature , specifically expressed as:
[0080]
[0081] By applying the gated spatial attention module, the non-linear layer can be removed and local continuity can be captured considering the complexity. Through Step 2 and Step 3, the accuracy of feature extraction for the template target in the initial frame and the target in subsequent frames can be effectively improved.
[0082] Denote the enhanced feature obtained from the initial frame through Step 2 and Step 3 as , and when the current frame is one of the subsequent frames, the obtained enhanced feature is denoted as . After the subsequent frames obtain the enhanced features through Step 2 and Step 3, they enter the next step.
[0083] Step 4, perform the association operation through the Siamese network to obtain the strongly associated region image: For the current frame, through the backbone of the Siamese network, on the enhanced feature of the template target Enhanced features with subsequent frames Share parameters with subsequent frames, embed them into a common feature space, perform cross-correlation between the features of the template region (i.e., the region within the target bounding box of the template target) and the features of the search region (i.e., the subsequent frame image) in the common feature space, judge their correlation, obtain the strongly correlated region according to the correlation, crop the strongly correlated region to obtain the strongly correlated region image, and at the same time output the corresponding correlation response map and loss function, and this loss function is a loss function of similarity measurement. Here it is different from the preprocessing image cropping. The preprocessing is to crop the same video sequence into pictures with the same resolution, while the cropping here is to measure the similarity of features in the subsequent frames and then crop the region with a higher similarity value of the target features in the subsequent frames to obtain the strongly correlated region image. The strongly correlated region image will change arbitrarily in different frames according to the deformation and scale change of the target. It is precisely because of this that the algorithm can be robust to the size change of the target. In this embodiment, the Siamese network is a deep learning Siamese network.
[0084] Video object tracking is regarded as a similarity learning problem under the Siamese network framework. Specifically, perform convolution operations on the enhanced features of the template region u and the search region v respectively, and obtain the corresponding features and features , then establish the corresponding mapping function f(u,v) , then through the mapping function f(u,v) perform similarity evaluation on and to obtain the correlation response map. The mapping function f(u,v) can be expressed as:
[0085]
[0086] Among them, is the similarity measurement function of the features and .
[0087] When f(u,v) has a higher value, it indicates that the targets on the template region and the search region are the same, while f(u,v) obtains a lower value, indicating that the targets on the template region and the search region are different. Crop and output the region corresponding to the higher value in the search region f(u,v) to obtain the strongly correlated region image. In this embodiment, the size of the template region u is: 127×127×3, and the size of the search region v is 255×255×3, f(u,v) The value of
[0088] Step 5: Input the strongly correlated region image into the classification regression network. The strongly correlated region image is processed by the classification branch and the regression branch in the classification regression network, iteratively outputs the offset regression, and then obtains the optimized correlation response map, target bounding box, and the corresponding loss function through convolution operations to complete the target tracking of the current frame.
[0089] For each pixel in the strongly correlated region image, the classification branch classifies the corresponding image patch as a positive sample or a negative sample; the regression branch outputs the offset regression based on the positive and negative samples, refines the position prediction of the target bounding box through the offset regression, and finally the size of the optimized correlation response map is 17×17×1. The classification regression network can use the existing classification regression network publicly disclosed in SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation Guidelines. (AAAI) by Yinda Xu et al. in 2020.
[0090] The strongly correlated region image is directly classified through the classification branch, and the target bounding box is regressed at the pixel position through the regression branch. In other words, each pixel position is directly regarded as a training sample. In this per-pixel prediction method, each pixel on the final feature map has only one prediction result, and each obtained classification score directly represents the confidence of the target at the corresponding pixel.
[0091] In this embodiment, after obtaining the correlation response map in Step 5, the classification score and the position regression score are respectively obtained, and the loss function corresponding to the target bounding box is denoted as , and is obtained through the following formula:
[0092]
[0093] where is the indicator function. If holds, then the value of the function is 1, otherwise it is 0; is the loss of the classification branch, using binary cross-entropy loss, is the actual position coordinate of the positive sample, is whether it is a positive sample. For positive samples, it takes 1, and for negative samples, it takes 0; is the loss of the regression branch, using IOU loss, is the ground truth coordinate of the target position, is the normalized ground truth coordinate of the target position, is the total number of positive samples, is a modifiable parameter, which takes the value of 0.5 in this embodiment, is the predicted target position.
[0094] Step 6: Determine whether the current frame is the last frame. If not, return to Step 2. If so, output the target bounding boxes corresponding to the template target in each subsequent frame, that is, the x and y coordinates of the four vertices thereof, and complete video target tracking.
[0095] To verify the effectiveness of the object tracking method based on the deep learning Siamese network proposed by the present invention, after training on the GOT10K dataset, this embodiment is tested on multiple other datasets, including OTB100, dtb70, VOT, UAV123 and other datasets. These datasets contain various tracking challenges, including scenarios such as object occlusion, object deformation, out-of-field-of-view, motion blur, etc., and at the same time cover image acquisition channels such as fixed cameras, mobile cameras, and drones, all showing good performance, achieving good results in accuracy, and the fps can reach 80, meeting the requirements of real-time video target tracking.
[0096] The object tracking method based on the deep learning Siamese network of the present invention can also form a computer program product. The program product includes a computer program, and when the program is executed by a processor, it implements the steps of an object tracking method based on the deep learning Siamese network. At the same time, the object tracking method of the present invention can also be applied in a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the above object tracking method can be stored in the computer-readable storage medium as a computer program. When the computer program is executed by a processor, it implements the steps of the above object tracking method.
Claims
1. A target tracking method based on deep learning twin network, characterized in that: The steps include: S1, obtain a video image file, specify a template target in the initial frame of the video image file, and give the true value of its target bounding box. The template target is used to provide a discriminant template for subsequent frames; S2, extract the feature information of the initial frame, process the feature information of the current frame through a deep separable convolutional network, obtain the corresponding local and global context information, and then perform feature enhancement to obtain the enhanced feature X of the template target G0 ; The feature enhancement is specifically performed through iterative enhancement processing by m multi-scale large core modules, where m≥2; the multi-scale large core module includes a multi-scale large core attention module and a gated space attention module that are arranged in sequence, and the process of the multi-scale large core attention module is as follows: N=LN(Z) The process of the gated spatial attention module is as follows: L=LN(X M ) X G =X M +λ2f6(GSAU(f4(L),f5(L))) Among them, Z is the feature obtained by the deep separable convolutional network, LN is the normalization operation, N is the feature obtained by normalizing feature Z, and L is the feature X M The normalized features; λ1 and λ2 are the learnable scaling factors in the current layer, is the outer product of the matrix, f n is the nth point-level convolution under the invariant dimension, n takes values of 1, 2, 3, 4, 5, 6, and is denoted as f1, f2, f3, f4, f5, f6, MLKA(f1(N)) is the dynamically adapted large kernel attention, GSAU(f4(L), f5(L)) is the densely transformed feature f4(L) and feature f5(L), X M is the output feature of the multi-scale large kernel attention module, X G It is the enhanced feature output by the multi-scale large kernel module; The deep separable convolutional network includes a feature decomposition module and a multi-order gated aggregation module. The feature information is input into the feature decomposition module and then normalized and convolutional layers are used to aggregate the feature information. 1×1 , deep convolution layer DWConv 3×3 And the first activation function GELU processing, using different scaling rates to obtain low-order features Y l , mid-level features Y m and high-order features Y h , specifically: Y l =Conv 1×1 (X), AND m =GELU(Y l +γ s ⊙(And l -GAP(Y l ))) Y h =GELU(DWConv 3×3 (Conv 1×1 (Norm(X)))) Among them, X is the feature information, GAP is the global average pooling, Norm is the normalization, γ s is a scaling factor initialized to zero, ⊙ is a circular convolution; The multi-stage gated aggregation module is provided with an aggregation branch and a context branch. The aggregation branch is used to m Generate gating weights. The context branch is used to extract multi-scale features using circular convolutions of different sizes based on the generated gating weights, thereby capturing multi-order interactions in the context and converting low-order features Y l , mid-level features Y m and high-order features Y h As input feature X in Input multi-stage gated aggregation module, through the convolution layer Conv 1×1 And the first activation function GELU is processed, and the feature Z is obtained as the output feature X through the guidance of the aggregation branch and the context branch. out Output: S3, extracting the next subsequent frame as the current frame, extracting its feature information, processing the feature information of the current frame through the deep separable convolutional network to obtain the corresponding local and global context information, and then performing the feature enhancement process to obtain the enhanced feature X of the subsequent frame G1 ; S4, enhanced features X of the subsequent frame through the twin network G1 and the enhanced features X of the template target G0 Perform association operation to obtain the strongly associated region image and its corresponding correlation response graph and loss function; S5, processing the strongly correlated region image through a classification regression network, iteratively outputting offset regression, and then obtaining the optimized correlation response map, target bounding box and corresponding loss function of the subsequent frame, and completing the target tracking of the current frame; S6, determine whether the current frame is the last frame, if not, return to step S3, if yes, output the target bounding box of all subsequent frames to complete video target tracking.
2. According to claim 1, a target tracking method based on deep learning twin network is characterized in that: In step S2 and step S3, the MLKA(f1(N)) is obtained by the following method: The feature f1(N) obtained by the first point-level convolution of feature N is processed by the group multi-scale mechanism to enhance the fixed large-core attention process, and the large-core attention LKA(f1(N)) of feature f1(N) is obtained as follows: LKA(f1(N))=f PW (in DWD (in DW (f1(N)))) Among them, f DW is the depth convolution, f DWD is a dilated convolution with a depth of D, f PW It is point-level convolution; For the i-th dimension feature f1(N) of feature f1(N) i Large kernel attention LKA i (f1(N) i ) is dynamically adapted to obtain the dynamic feature MLKA of the i-th dimension i (f1(N) i ), as follows: Among them, G i (f1(N) i ) is f1(N) i Space gate; Then the dynamic feature MLKA i (X i ) are aggregated to obtain the dynamically adapted large kernel attention MLKA(f1(N)).
3. According to claim 2, a target tracking method based on deep learning twin network is characterized in that: In step S2 and step S3, the GSAU (f4 (L), f5 (L)) is obtained specifically by the following method:
4. According to claim 3, a target tracking method based on deep learning twin network is characterized in that: In step S5, the loss function is recorded as L({p x,y }, q x,y , {t x,y }), we get: in, is an indicator function, if If it holds, then the value of the function is 1, otherwise it is 0; L cls is the loss of the classification branch, using binomial cross entropy loss; p x,y is the actual position coordinate of the positive sample; Is it a positive sample? If it is a positive sample, it takes 1, and if it is a negative sample, it takes 0; L reg is the loss of the regression branch, using IOU loss; t x,y is the true value coordinate of the target position, is the normalized true value coordinate of the target position; N pos is the total number of positive samples; λ is a modifiable parameter; q x,y is the predicted target location.
5. According to claim 4, a target tracking method based on deep learning twin network is characterized in that: In step S5, the value of λ is 0.
5.
6. A computer program product, comprising a computer program, characterized in that: When the program is executed by a processor, the steps of a target tracking method based on a deep learning twin network as described in any one of claims 1 to 5 are implemented.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of a target tracking method based on a deep learning twin network as described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Target tracking method based on twin neural network
CN112712546A
Lightweight camouflage target detection method based on deep learning
CN119091117A