A Method for Tracking Weak Dynamic Targets on the Water Surface Based on a Combination of Static and Dynamic Attention
Through the method based on dynamic and static combined attention, the indirect transmission estimation and spatial self-attention feature extraction algorithm are used to solve the problems of performance instability and inaccurate positioning in dynamic weak target tracking on water surface, and real-time and efficient tracking of the target is achieved.
Patent Information
- Application Number
- CN202510186458.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-02-20
AI Technical Summary
In the detection of weak water surface, the existing technology has problems such as unstable and unsustainable dynamic weak target tracking performance, poor target positioning accuracy and large calculation amount.
A method based on dynamic and static combined attention is adopted to obtain clear images of defog through indirect estimation of transmittance, a spatial self-attention feature extraction algorithm is used to obtain feature maps, and a dynamic and static target tracking algorithm is used to locate the target.
Improve the real-time and accuracy of dynamic weak target tracking, reduce the impact of environmental interference, and ensure the sustainability and accuracy of target detection.
Smart Images

Figure CN119672069B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a method for tracking dynamic weak targets on the water surface based on combined static and dynamic attention. Background Art
[0002] In the aspect of weak target detection on the water surface, due to the significant attenuation of light collection caused by sunlight refraction and scattering, the clarity of the water surface imaging significantly decreases. Before target tracking and recognition, the problem of image optimization processing for foggy water surfaces should be focused on. The existing defogging methods mainly include three categories: image enhancement, image restoration, and deep learning. The above traditional algorithms have poor boundary discrimination effects between foggy water surfaces and foggy land surfaces. Due to the lack of publicly available water surface image datasets, it will consume a large amount of computing time when separating water surface targets and interfering backgrounds, and the detection rate and recognition accuracy are poor.
[0003] In the research on dynamic weak target tracking and recognition, algorithms such as deep learning, Kalman filtering, YOLO network, and multi-modal collaborative detection are widely used at present. Their basic idea is to obtain target feature information through sample learning and training, and then predict target location information according to the temporal causal relationship between the front and back frames of the video stream. Although the existing algorithms have made certain progress, there are mainly two problems: one is that in a complex interference environment, the tracking performance of dynamic weak targets is unstable and discontinuous, and the target location accuracy needs to be improved; the other is that due to the large number of target categories, traditional recognition methods rely more on the training effects of previous algorithms and have a large amount of computation, and the recognition effect for targets with unknown types and irregular shapes is poor. Therefore, there is an urgent need for a method for tracking dynamic weak targets on the water surface based on combined static and dynamic attention to solve the deficiencies of the above existing technologies. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for tracking dynamic weak targets on the water surface based on combined static and dynamic attention, so as to solve the problems of unstable and discontinuous tracking performance of dynamic weak targets, poor target location accuracy, and large amount of computation existing in the prior art.
[0005] To achieve the above purpose, the present invention provides a method for tracking dynamic weak targets on the water surface based on combined static and dynamic attention, including the following steps:
[0006] Step 1: Obtain a clear image with successful defogging through a prior defogging algorithm based on indirect estimation of transmittance; on the premise of obtaining prior defogging knowledge, first construct a physical model of foggy imaging, then indirectly estimate the transmittance through an atmospheric light dissipation function, and finally obtain a clear image with successful defogging;
[0007] Step 2: Obtain the spatial self-attention feature map based on the spatial self-attention feature extraction algorithm; adaptively adjust the static image frame through autonomous learning, calculate the depth value of the correlation between features using global max pooling and global average pooling, and obtain the spatial self-attention feature map;
[0008] Step 3: Track and locate the dynamic weak target based on the static and dynamic combined attention target tracking algorithm; obtain the depth feature values of the static and dynamic image frames respectively through the spatial self-attention feature extraction algorithm, then analyze the correlation between the two, and finally perform dynamic update estimation according to the target response scale.
[0009] Preferably, the process of obtaining a clear image with successful defogging based on the prior defogging algorithm indirectly estimated by transmittance in Step 1 is as follows:
[0010] S11: Construct a physical model of foggy imaging, and the expression is as follows:
[0011] (1)
[0012] In the formula, is the original image, is the clear image after defogging, is the global atmospheric light coefficient, is the medium transmittance;
[0013] S12: Calculate the transmittance estimated value ;
[0014] S13: Substitute the transmittance estimated value obtained in S12 into Equation (1) to obtain the defogged image, and the expression is as follows:
[0015] (2).
[0016] Preferably, the process of calculating the transmittance estimated value in S12 based on the transmittance indirect estimation algorithm is as follows:
[0017] S121: Construct an atmospheric dissipation function , and the expression is as follows:
[0018] (3)
[0019] In the formula, is the global atmospheric light coefficient, is the medium transmittance. For any pixel point, always holds;
[0020] S122: For the original image Obtain the minimum value of the three color channels to obtain a grayscale image with the same size as the original image , and the expression is as follows:
[0021] (4)
[0022] In formula (4), the minimum color channel component of the original image always holds;
[0023] S123. Combine formulas (3) and (4) to derive the following formula:
[0024] (5);
[0025] S124. Perform adaptive surface blurring on to obtain the sky region and non-sky region images, and the expression is as follows:
[0026] (6)
[0027] In the formula, is the grayscale image after blurring;
[0028] S125. Since and have the same transformation trend, obtain the transmittance estimate value , and the expression is as follows:
[0029] (7)
[0030] In the formula, is the adjustment parameter.
[0031] Preferably, the process of obtaining the spatial self-attention feature map based on the spatial self-attention feature extraction algorithm in step 2 is as follows:
[0032] S21. Using the static image frame as the input tensor x, set the pooling windows (1, W) and (H, 1) respectively, and calculate the row and column feature means of the input tensor x in the horizontal and vertical directions , , so as to obtain spatial information to suppress background noise interference, and the expression is as follows:
[0033] (8)
[0034] (9)
[0035] In the formula, the input tensor is , the first-order tensor in the horizontal direction is W is the width, the first-order tensor in the vertical direction is , H is the height, is a second-order gradient tensor;
[0036] S22. Construct an attention information fusion network, introduce the number of channels C into the attention information fusion network, and decompose the input tensor into , and respectively obtain the first-order tensor in the horizontal direction and the first-order tensor in the vertical direction after decomposition. Use the add method to perform feature fusion on and after decomposition. The expression is as follows:
[0037] ) (10)
[0038] In the formula, represents the fused feature;
[0039] S23. Perform batch normalization and activation processing on the fused feature to obtain the output value z. The expression is as follows:
[0040] (11)
[0041] S24. Perform global max pooling and global average pooling along the channel axis to capture the correlation relationship between input features, and obtain a spatial self-attention feature map containing color, shape, and texture saliency regions. The expression is as follows:
[0042] (12)
[0043] In the formula, is the global max pooling operation, is the global average pooling operation.
[0044] Preferably, the process of tracking and positioning dynamic weak targets based on the dynamic and static combined attention target tracking algorithm in step 3 is as follows:
[0045] S31. Crop the target area of the static image frame as a reference template; set the length of the target O to w and the length to h. Taking the first frame of the video sequence as an example, crop a square area with a side length of L centered on the position O where the target is located. This area can completely contain all the information of the target O. The expression is as follows:
[0046] (13)
[0047] In the formula, S is the cropped area of the target area;
[0048] S32. Continuously track the target based on the dynamic image frames. When the target moves or is occluded by environmental interference, select the image frame with the highest confidence as the dynamic template, and complete the automatic update of the dynamic template through the self-learning update strategy to obtain the best confidence H of the current frame;
[0049] S33. Continuously track the target, and calculate the similarity between the static / dynamic-static dual-template sub-sequence and the search sub-sequence 、 , and then perform weighted fusion to obtain a more accurate target positioning result. The expression is as follows:
[0050] (14)
[0051] (15)
[0052] (16)
[0053] In the formula, is the self-learning target change value of the static template, is the self-learning target change value of the dynamic-static dual-template, is the spatial self-attention feature map of the static template, is the spatial self-attention feature map of the dynamic template, is the convolution operation, is the learned feature of the background suppression module, is the spatial self-attention feature map of the search area, is the balance parameter of the dynamic-static dual-template, is the response map after weighted fusion of the current frame, and its best confidence is the target prediction positioning;
[0054] S34. Perform dynamic update estimation according to the target response scale. The expression is as follows:
[0055] (17)
[0056] In the formula, is the maximum response scale of the target in the current frame image, is the maximum response scale of the target in the previous frame image, and v is the scale change speed between two frames.
[0057] Preferably, the automatic update of the dynamic template needs to satisfy that the best confidence H of the current frame and the peak value of the response map both satisfy being greater than the historical average level. The expression is as follows:
[0058] (18)
[0059] (19)
[0060] In the formula, is the historical mean, is the maximum mean, , is the dynamic template update frequency.
[0061] Therefore, the present invention adopts the above-mentioned method for tracking dynamic weak targets on the water surface based on the combination of dynamic and static attention, and has the following beneficial effects:
[0062] (1) A physical model of foggy weather imaging is constructed. The transmittance is indirectly estimated through the atmospheric light dissipation function, and finally a clear image with successful defogging is obtained, successfully overcoming the problem of poor clarity of optical imaging in complex foggy weather at sea, creating a basic condition for the subsequent continuous tracking and positioning of dynamic weak targets on the water surface, and further reducing the influence of environmental constraints on water surface target recognition;
[0063] (2) The real-time performance and accuracy of dynamic weak target tracking are improved: First, convolutional channels are added to the attention information fusion network, and feature depth value correlation operations are performed through adaptive maximum pooling and average pooling operations, and spatial self-attention feature maps of dynamic and static image frames can be obtained; Then, the correlation relationship between dynamic and static image frames is calculated through the dynamic and static combined attention target tracking algorithm, and the self-learning update algorithm is used to keep the target detection always in the best state, avoiding the phenomenon of missing and undetected dynamic weak targets caused by movement or occlusion, and further improving the real-time performance and persistence of the whole process of target tracking.
[0064] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. Description of the Drawings
[0065] Figure 1 is the overall flowchart of a method for tracking dynamic weak targets on the water surface based on the combination of dynamic and static attention of the present invention;
[0066] Figure 2 is the principle block diagram of the spatial self-attention feature extraction algorithm of the embodiment of the present invention;
[0067] Figure 3 is the principle block diagram of the dynamic and static combined attention target tracking algorithm of the embodiment of the present invention. Specific Embodiments
[0068] The following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the present invention claimed, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0069] Please refer to Figures 1-3, A method for tracking weak dynamic targets on the water surface based on a combination of dynamic and static attention, comprising the following steps:
[0070] Step 1: Obtain a clear image with successful defogging using a prior defogging algorithm based on indirect transmittance estimation; on the premise of obtaining prior defogging knowledge, first construct a physical model of foggy weather imaging, then indirectly estimate the transmittance through an atmospheric light dissipation function, and finally obtain a clear image with successful defogging; among them, the process of obtaining a clear image with successful defogging using a prior defogging algorithm based on indirect transmittance estimation is as follows:
[0071] S11: Construct a physical model of foggy weather imaging, the expression is as follows:
[0072] (1)
[0073] In the formula, is the original image, is the clear image after defogging, is the global atmospheric light coefficient, is the medium transmittance;
[0074] S12: Calculate the transmittance estimate using an indirect transmittance estimation algorithm; the specific calculation process is as follows:
[0075] S121: Construct an atmospheric dissipation function , the expression is as follows:
[0076] (3)
[0077] In the formula, is the global atmospheric light coefficient, is the medium transmittance, for any pixel point, always holds;
[0078] S122: Obtain the minimum value of the three color channels of the original image to obtain a grayscale image with the same size as the original image, the expression is as follows:
[0079] (4)
[0080] In formula (4), the minimum color channel component of the original image always holds;
[0081] S123: Combine formulas (3) and (4) to derive the following formula:
[0082] (5);
[0083] S124: For Perform adaptive surface blurring to obtain sky region and non-sky region images. The expression is as follows:
[0084] (6)
[0085] In the formula, is the grayscale image after blurring;
[0086] S125. From and having the same transformation trend, obtain the transmittance estimate value . The expression is as follows:
[0087] (7)
[0088] In the formula, is the adjustment parameter.
[0089] S13. Substitute the transmittance estimate value obtained in S12 into Equation (1) to obtain the dehazed image. The expression is as follows:
[0090] (2).
[0091] Step 2. Obtain the spatial self-attention feature map based on the spatial self-attention feature extraction algorithm; adaptively adjust the static image frame through autonomous learning, and calculate the depth value of the correlation between features using global max pooling and global average pooling to obtain the spatial self-attention feature map. Among them, the process of obtaining the spatial self-attention feature map based on the spatial self-attention feature extraction algorithm is as follows:
[0092] S21. Taking the static image frame as the input tensor x, set the pooling windows (1, W) and (H, 1) respectively, and calculate the row and column feature means of the input tensor x in the horizontal and vertical directions to obtain spatial information to suppress background noise interference. The expression is as follows:
[0093] (8)
[0094] (9)
[0095] In the formula, the input tensor is , the first-order tensor in the horizontal direction is W is the width, the first-order tensor in the vertical direction is , H is the height, is the second-order gradient tensor;
[0096] S22. Construct an attention information fusion network. Introduce the number of channels C in the attention information fusion network, and decompose the input tensor into respectively obtain the first-order tensors in the horizontal direction after decomposition and the first-order tensors in the vertical direction and adopt the add method to perform feature fusion on the decomposed and The expression is as follows:
[0097] (10)
[0098] In the formula, represents the fused feature;
[0099] S23. Perform batch normalization and activation processing on the fused feature to obtain the output value z. The expression is as follows:
[0100] (11)
[0101] S24. Perform global max pooling and global average pooling along the channel axis to capture the correlation between input features, and obtain the spatial self-attention feature map containing the significant regions of color, shape, and texture The expression is as follows:
[0102] (12)
[0103] In the formula, is the global max pooling operation, is the global average pooling operation.
[0104] Step 3. Track and locate the dynamic weak target based on the dynamic and static combined attention target tracking algorithm; respectively obtain the depth feature values of the static and dynamic image frames through the spatial self-attention feature extraction algorithm, then analyze the correlation between the two, and finally perform dynamic update estimation according to the target response scale; among them, the process of tracking and locating the dynamic weak target based on the dynamic and static combined attention target tracking algorithm is as follows:
[0105] S31. Crop the target area of the static image frame as the reference template; set the length of the target O to w and the length to h. Taking the first frame of the video sequence as an example, crop a square area with side length L centered at the point O where the target is located. This area can complete the inclusion of all information of the target O. The expression is as follows:
[0106] (13)
[0107] In the formula, S is the cropped area of the target area;
[0108] S32. Continuously track the target based on dynamic image frames. When the target moves or is occluded by environmental interference, select the image frame with the highest confidence as the dynamic template, and complete the automatic update of the dynamic template through a self-learning update strategy to obtain the best confidence H of the current frame. Among them, the automatic update of the dynamic template needs to satisfy that both the best confidence H of the current frame and the peak value of the response map both meet the requirement of being greater than the historical average level. The expression is as follows:
[0109] (18)
[0110] (19)
[0111] In the formula, is the historical mean value, is the maximum mean value, 、 is the dynamic template update frequency;
[0112] S33. Track the target in real time, and calculate the similarity between the static / dynamic-static double-template sub-sequence and the search sub-sequence 、 , and then perform weighted fusion to obtain a more accurate target positioning result. The expression is as follows:
[0113] (14)
[0114] (15)
[0115] (16)
[0116] In the formula, is the self-learning target change value of the static template, is the self-learning target change value of the dynamic-static double-template, is the spatial self-attention feature map of the static template, is the spatial self-attention feature map of the dynamic template, is the convolution operation, is the learned feature of the background suppression module, is the spatial self-attention feature map of the search area, is the balance parameter of the dynamic-static double-template, is the response map after weighted fusion of the current frame, and its best confidence is the target prediction positioning;
[0117] S34. Perform dynamic update estimation according to the target response scale. The expression is as follows:
[0118] (17)
[0119] In the formula, is the maximum response scale of the target in the current frame image, is the maximum response scale of the target in the previous frame image, and v is the scale change speed between two frames.
[0120] Therefore, the present invention adopts the above-mentioned method for tracking dynamic weak targets on the water surface based on the combination of static and dynamic attention. First, a clear image with successful defogging is obtained based on the prior defogging algorithm indirectly estimated by the transmittance, reducing the interference of complex foggy weather at sea on optical target recognition. Then, a spatial self-attention feature map is obtained based on the spatial self-attention feature extraction algorithm. Finally, the dynamic weak target is tracked and positioned based on the algorithm for tracking dynamic weak targets by combining static and dynamic attention, thereby improving the real-time performance and accuracy of tracking dynamic weak targets.
[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements do not make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for tracking weak dynamic targets on the water surface based on a combination of dynamic and static attention, characterized in that, It includes the following steps: Step 1: Obtain a clear image with successful defogging using a prior defogging algorithm based on indirect transmittance estimation. On the premise of obtaining prior defogging knowledge, first construct a physical model of foggy-day imaging, then indirectly estimate the transmittance through an atmospheric light dissipation function, and finally obtain a clear image with successful defogging; Step 2: Obtain a spatial self-attention feature map using a spatial self-attention feature extraction algorithm. Adaptively adjust the static image frame through self-learning, calculate the depth value of the correlation between features using global max pooling and global average pooling, and obtain a spatial self-attention feature map. Among them, calculate the first-order tensor in the horizontal direction and the first-order tensor in the vertical direction through the static frame image, and use the add method to perform feature fusion on the decomposed first-order tensors; Step 3: Track and locate dynamic weak targets using a dynamic and static combined attention target tracking algorithm. Obtain the depth feature values of static and dynamic image frames respectively through a spatial self-attention feature extraction algorithm, then analyze the correlation between the two, and finally perform dynamic update estimation according to the target response scale. Among them, crop the target area of the static image frame as a reference template; for the dynamic image frame, select the image frame with the best confidence as the dynamic template, and complete the automatic update of the dynamic template through a self-learning update strategy to obtain the best confidence of the current frame.
2. The method for tracking weak dynamic targets on the water surface based on a combination of dynamic and static attention according to claim 1, wherein The process of obtaining a clear image with successful defogging using the prior defogging algorithm based on indirect transmittance estimation in Step 1 is as follows: S11: Construct a physical model of foggy-day imaging, and the expression is as follows: (1) In the formula, is the original image, is the clear image after defogging, is the global atmospheric light coefficient, is the medium transmittance; S12. Calculate the estimated transmittance value based on the indirect transmittance estimation algorithm ; S13. Substitute the transmittance estimation value obtained in S12 into Equation (1) to obtain the defogged image, and the expression is as follows: (2)。 3. A method for tracking weak dynamic targets on the water surface based on a combination of static and dynamic attention, characterized in that, The process of calculating the estimated value of transmittance based on the indirect estimation algorithm in S12 is as follows: is as follows: S121. Construct the atmospheric dissipation function , and the expression is as follows: (3) In the formula, is the global atmospheric light coefficient, is the medium transmittance, and for any pixel point, always holds; S122. For the original image Obtain the minimum value of the three color channels to get a grayscale image with the same size as the original image , and the expression is as follows: (4) In formula (4), the minimum color channel component of the original image always holds; S123: Combine equations (3) and (4) to derive the following formula: (5); S124. Adaptive surface blurring is performed on to obtain sky region and non-sky region images, and the expression is as follows: (6) In the formula, is the grayscale image after blurring processing; S125. Obtained from and with the same transformation trend, the estimated transmittance value is obtained, and the expression is as follows: (7) In the formula, is the adjustment parameter.
4. A method for tracking weak dynamic targets on the water surface based on a combination of dynamic and static attention, characterized in that, The process of obtaining a spatial self-attention feature map using a spatial self-attention feature extraction algorithm in Step 2 is as follows: S21. Taking the static image frame as the input tensor x, respectively set the pooling windows (1, W) and (H, 1), and calculate the mean values of the row and column features of the input tensor x in the horizontal and vertical directions. , The expressions are as follows: (8) (9) In the formula, the input tensor is , the first-order tensor in the horizontal direction is , W is the width, and the first-order tensor in the vertical direction is , H is the height, is the second-order gradient tensor; S22. Construct an attention information fusion network. Introduce the number of channels C into the attention information fusion network, and decompose the input tensor into , and respectively obtain the first-order tensor in the horizontal direction and the first-order tensor in the vertical direction . Use the add method to perform feature fusion on the decomposed and . The expression is as follows: (10) In the formula, represents the fusion feature; S23. Perform batch normalization and activation processing on the fused feature to obtain the output value z, and the expression is as follows: (11) S24. Perform global max pooling and global average pooling along the channel axis to capture the correlation between input features, and obtain a spatial self-attention feature map containing color, shape, and texture saliency regions. , and the expression is as follows: (12) In the formula, is the global maximum pooling operation, is the global average pooling operation.
5. A method for tracking weak dynamic targets on the water surface based on a combination of static and dynamic attention, characterized in that, The process of tracking and locating dynamic weak targets using a dynamic and static combined attention target tracking algorithm in Step 3 is as follows: S31: Crop the target area of the static image frame as a reference template; set the length of target O as w and the length as h. Taking the first frame of the video sequence as an example, crop a square area with side length L centered on the position O where the target is located. This area can completely contain all the information of target O, and the expression is as follows: (13) In the formula, S is the cropped area of the target region; S32: Continuously track the target according to the dynamic image frame. When the target moves or is blocked by environmental interference, select the image frame with the highest confidence as the dynamic template, and complete the automatic update of the dynamic template through a self-learning update strategy to obtain the best confidence H of the current frame; S33. Perform real-time tracking on the target, calculate the similarity between the static / moving-static dual-template sub-sequences and the search sub-sequence, , and then perform weighted fusion to obtain the target positioning result. The expression is as follows: (14) (15) (16) Wherein, is the change value of the static template self - learning target, which is obtained by inputting a static image frame into a CNN convolutional network and then performing spatial self - attention feature extraction, is the change value of the dynamic and static dual - template self - learning target, which is obtained by inputting a dynamic image frame into a CNN convolutional network and then performing spatial self - attention feature extraction, is the spatial self - attention feature map of the static template, is the spatial self - attention feature map of the dynamic template, is the convolution operation, is the feature learned by the background suppression module, which is obtained by inputting a learning image frame into a CNN convolutional network and then performing spatial self - attention feature extraction, is the spatial self - attention feature map of the search area, is the balance parameter of the dynamic and static dual - template, is the response map after weighted fusion of the current frame; S34: Perform dynamic update estimation according to the target response scale, and the expression is as follows: (17) Wherein, is the maximum response scale of the target in the current frame image, is the maximum response scale of the target in the previous frame image, and v is the scale change speed between two frames.
6. A method for tracking weak dynamic targets on the water surface based on a combination of dynamic and static attention, characterized in that: The automatic update of the dynamic template needs to meet the best confidence level H of the current frame and the peak of the response map Both need to be greater than the historical average level, and the expression is as follows: (18) (19) Wherein, is the historical mean value, is the maximum mean value, , is the dynamic template update frequency.
Citation Information
Patent Citations
Fishing boat target detection and tracking method based on dynamic video
CN116935332A
Single-target tracking method and device, electronic equipment and computer readable storage medium
CN117218158A