Radar target tracking method, device and electronic equipment

CN122283690BActive Publication Date: 2026-09-18CHENGDU RUIYANXINCHUANG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610729800.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-26
Publication Date
2026-09-18
Estimated Expiration
2046-05-26

AI Technical Summary

Technical Problem

[0003]鉴于上述问题,本申请提供一种雷达目标跟踪方法、装置及电子设备,能够解决现有方法准确性和可靠性低的问题

Benefits of technology

[0023] Fourthly, this application provides a readable storage medium storing a computer program, which, when executed by a processor, performs the radar target tracking method described in any one of the first aspects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122283690B_ABST
    Figure CN122283690B_ABST
Patent Text Reader

Abstract

The application discloses a radar target tracking method and device and electronic equipment. The method comprises the following steps: acquiring an original bimodal image sequence; wherein the original bimodal image sequence comprises a spatially aligned original visible light image sequence and an original radar image sequence; pre-processing the original bimodal image sequence to obtain a pre-processed bimodal image sequence; tracking a preset target in the pre-processed bimodal image sequence through a target tracking model to obtain a bimodal tracking result set of the preset target and a fusion tracking result set; and tracking the preset target according to a preset target motion model, the bimodal tracking result set and the fusion tracking result set to obtain a final tracking result set of the preset target. The method can solve the problem of low accuracy and reliability of existing methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of target tracking technology, specifically to a radar target tracking method, device, and electronic equipment. Background Technology

[0002] With the rapid development of applications such as UAV security and remote sensing monitoring, the need for stable and continuous target tracking in complex environments is becoming increasingly urgent. Utilizing the complementarity of visible light and synthetic aperture radar (SAR) dual-mode images is considered a key technological direction for improving tracking robustness. Current technologies typically use two independent tracking models to process visible light and SAR image sequences, then perform a simple linear weighted fusion of the two tracking results to obtain the final target tracking result. However, this method only performs a simple linear weighted fusion of the two tracking results. When the quality of one mode image suddenly deteriorates due to factors such as illumination or noise, it directly affects the accuracy of the final target tracking result, leading to target tracking drift or loss, thus reducing the reliability of the target tracking system. Summary of the Invention

[0003] In view of the above problems, this application provides a radar target tracking method, apparatus and electronic device that can solve the problems of low accuracy and reliability of existing methods.

[0004] In a first aspect, this application provides a radar target tracking method, including: Obtain the original bimodal image sequence; wherein, the original bimodal image sequence includes a spatially aligned original visible light image sequence and an original radar image sequence; The original bimodal image sequence is preprocessed to obtain a preprocessed bimodal image sequence; The preset target in the preprocessed bimodal image sequence is tracked by the target tracking model to obtain the bimodal tracking result set and the fused tracking result set of the preset target; The preset target is tracked based on the preset target motion model, the dual-modal tracking result set, and the fused tracking result set to obtain the final tracking result set of the preset target.

[0005] In the above technical solution, the method can achieve dynamic adaptive fusion of visible light and radar modes by comprehensively evaluating all dual-mode tracking results and fused tracking results, and optimizing them in combination with the target motion model. This effectively avoids tracking drift when a single mode fails, and significantly improves the accuracy, robustness and system reliability of target tracking in complex environments.

[0006] In some embodiments, the preprocessing of the original bimodal image sequence to obtain a preprocessed bimodal image sequence includes: The original visible light image sequence is subjected to image enhancement processing to obtain an enhanced visible light image sequence; The enhanced visible light image sequence is denoised to obtain a preprocessed visible light image sequence; The original radar image sequence is denoised to obtain a first processed radar image sequence; The first processed radar image sequence is corrected to obtain the second processed radar image sequence. The second processed radar image sequence is corrected according to the imaging parameters of synthetic aperture radar to obtain a preprocessed radar image sequence. The preprocessed dual-modal image sequence includes the preprocessed visible light image sequence and the preprocessed radar image sequence.

[0007] In the above technical solution, the method can effectively eliminate the imaging noise and distortion differences between the two types of images by performing targeted enhancement and denoising on visible light images and sequentially performing denoising and correction processing on radar images, thus ensuring the registration accuracy and information quality of the preprocessed dual-modal images.

[0008] In some implementations, the step of tracking a preset target in the preprocessed bimodal image sequence using a target tracking model to obtain a bimodal tracking result set and a fused tracking result set for the preset target includes: In the initial frame of the preprocessed visible light image sequence, a first initial target box corresponding to the preset target is determined, and in the initial frame of the preprocessed radar image sequence, a second initial target box corresponding to the first initial target box is determined; The preset targets in the preprocessed visible light image sequence are tracked using a target tracking model and the first initial target bounding box to obtain a visible light tracking result set; The preset target in the preprocessed radar image sequence is tracked using a target tracking model and the second initial target bounding box to obtain a radar tracking result set; wherein, the dual-mode tracking result set includes the visible light tracking result set and the radar tracking result set; Feature extraction is performed on the preprocessed bimodal image sequence to obtain a fused feature sequence; The target tracking process of the preset target is performed using the target tracking model and the fused feature sequence to obtain a fused tracking result set.

[0009] In the above technical solution, the method can use the multi-branch structure of the target tracking model to complete single-mode tracking of visible light and radar respectively, and achieve joint tracking based on the fused feature sequence. It effectively integrates the advantages of dual-mode data, avoids the limitations of single-mode information, and significantly improves the accuracy of dual-mode tracking result set and fused tracking result set.

[0010] In some implementations, the target tracking model includes at least a template branch, a search branch, a classification branch, and a regression branch; The step of tracking a preset target in the preprocessed visible light image sequence using a target tracking model and the first initial target bounding box to obtain a visible light tracking result set includes: Visual features within the first initial target box are extracted from the preprocessed visible light image sequence using the template branch; Any frame in the subsequent frames of the preprocessed visible light image sequence is determined as the first target frame, and multiple first candidate features corresponding to the first target frame are extracted based on the visual features through the search branch. A cross-correlation operation is performed on the visual features and the plurality of first candidate features to obtain a plurality of first candidate tracking results; the plurality of first candidate tracking results correspond one-to-one with the plurality of first candidate features. Multiple first confidence scores are obtained by calculating using the classification branch and the multiple first candidate tracking results; each of the multiple first confidence scores corresponds one-to-one with the multiple first candidate tracking results. Based on the plurality of first confidence scores, a first optimal candidate result corresponding to the first target frame is determined from the plurality of first candidate tracking results; After determining all the first optimal candidate results corresponding to the subsequent frames of the preprocessed visible light image sequence, the regression branch is used to process all the first optimal candidate results to obtain the first tracking result set of the preset target in the subsequent frames; The visible light tracking result set includes the initial tracking result obtained based on the first initial target box and the first tracking result set.

[0011] In the above technical solution, the method can efficiently generate candidate regions through the cross-correlation operation of template branch and search branch, and use the confidence score of classification branch to screen the optimal candidate. Finally, the position is refined through regression branch, realizing a refined tracking process for visible light targets from candidate generation, confidence screening to coordinate refinement, which significantly improves the positioning accuracy and stability of visible light tracking results.

[0012] In some implementations, the step of tracking the preset target in the preprocessed radar image sequence using a target tracking model and the second initial target bounding box to obtain a radar tracking result set includes: Through the template branch, microwave scattering features within the second initial target box are extracted from the initial frame of the preprocessed radar image sequence; Any frame in the subsequent frames of the preprocessed radar image sequence is determined as the second target frame, and multiple second candidate features corresponding to the second target frame are extracted based on the microwave scattering features through the search branch. Based on the microwave scattering characteristics and the plurality of second candidate characteristics, a plurality of second candidate tracking results are generated; the plurality of second candidate tracking results correspond one-to-one with the plurality of second candidate characteristics. Multiple second confidence scores are obtained by calculating using the classification branch and the multiple second candidate tracking results; each of the multiple second confidence scores corresponds one-to-one with the multiple second candidate tracking results. Based on the plurality of second confidence scores, a second optimal candidate result corresponding to the second target frame is determined from the plurality of second candidate tracking results; After determining all the second optimal candidate results corresponding to the subsequent frames of the preprocessed radar image sequence, the regression branch is used to process all the second optimal candidate results to obtain the second tracking result set of the preset target in the subsequent frames of the preprocessed radar image sequence. The radar tracking result set includes the initial tracking result obtained based on the second initial target box and the second tracking result set.

[0013] In the above technical solution, the method can accurately extract the microwave scattering features of the target through the template branch, generate candidate regions through the search branch, and combine the confidence screening of the classification branch and the position refinement of the regression branch to fully explore the target characterization information of the radar image and effectively generate a high-precision radar tracking result set, providing a reliable radar modal benchmark for the fusion of dual-modal tracking results.

[0014] In some embodiments, the step of extracting features from the preprocessed bimodal image sequence to obtain a fused feature sequence includes: Based on the first initial target bounding box, obtain the first target region feature map corresponding to each frame in the preprocessed visible light image sequence, and based on the second initial target bounding box, obtain the second target region feature map corresponding to each frame in the preprocessed radar image sequence. The weights of the corresponding first target region feature maps are calculated based on the tracking confidence scores of each frame in the visible light tracking result set, resulting in the first feature weights of the first target region feature maps for each frame. The weights of the corresponding second target region feature maps are also calculated based on the tracking confidence scores of each frame in the radar tracking result set, resulting in the second feature weights of the second target region feature maps for each frame. Based on the first target region feature map and the first feature weight corresponding to each frame, a one-to-one feature weighting process is performed to obtain the first weighted feature map corresponding to each frame. Based on the second target region feature map and the second feature weight corresponding to each frame, a one-to-one feature weighting process is performed to obtain the second weighted feature map corresponding to each frame. The first and second weighted feature maps corresponding to each frame are concatenated one-to-one along the channel dimension to obtain the fused feature map corresponding to each frame. The fusion feature sequence includes fusion feature maps corresponding to all frames.

[0015] In the above technical solution, the method can dynamically adjust the feature weights according to the tracking confidence of each of the two modalities, and achieve deep fusion of the features of the two modalities through weighted summation and channel splicing. This effectively integrates the complementary advantages of the two modalities, avoids the influence of single modal feature bias, thereby improving the representation ability of the fused features and ensuring the accuracy of the subsequent fusion tracking results.

[0016] In some implementations, the step of performing target tracking processing on the preset target using the target tracking model and the fused feature sequence to obtain a fused tracking result set includes: In the initial frame of the fused feature sequence, a third initial target box corresponding to the first initial target box is determined; The target tracking model uses template branches and search branches, along with the third initial target box, to match and track the fused feature sequence, resulting in fused tracking results for multiple frames. The fused tracking results include at least a fused confidence score, fused target box location coordinates, and fused target scale. The fusion tracking result set includes the fusion tracking results for all frames.

[0017] In the above technical solution, the method can perform matching tracking based on the fused feature sequence and combined with the third initial target box, through the template branch and search branch of the target tracking model. It makes full use of the complementary advantages of dual-modal fused features and accurately outputs fused tracking results including confidence, coordinates and scale, further improving the accuracy and stability of target tracking.

[0018] In some implementations, the step of tracking the preset target based on a preset target motion model, the dual-modal tracking result set, and the fused tracking result set to obtain the final tracking result set of the preset target includes: Based on the dual-modal tracking result set and the fused tracking result set, construct multi-source observation vectors corresponding to multiple frames; An initial state vector is generated based on the initial observation vector corresponding to the initial frame in the multi-source observation vector; Recursive prediction processing is performed based on the preset target motion model and the initial state vector to obtain state prediction values ​​for multiple frames. Based on the dual-modal tracking result set and the fusion confidence scores of multiple frames in the fusion tracking result set, the multi-source observation vectors corresponding to the multiple frames are subjected to one-to-one weighted fusion processing to obtain the weighted observation values ​​corresponding to the multiple frames. The state prediction values ​​and weighted observation values ​​corresponding to the multiple frames are processed by a state update algorithm based on Kalman filtering to obtain the optimal estimate of the target state corresponding to the multiple frames. Based on the optimal estimation of the target state corresponding to the multiple frames, the final tracking result corresponding to the preset target in the multiple frames is determined; wherein, the final tracking result includes the final tracking position and the final scale; The final tracking result set includes the final tracking results of the preset target in multiple frames.

[0019] In the above technical solution, the method can use the Kalman filter algorithm, combined with the target motion model, to predict and update multi-source tracking observation information. By using confidence-weighted fusion to weaken the interference of low-quality tracking results, the target state can be accurately estimated, thereby effectively solving the tracking drift or loss problem and ensuring the accuracy and continuity of the final tracking results.

[0020] Secondly, this application provides a radar target tracking device, comprising: An acquisition unit is used to acquire a raw bimodal image sequence; wherein the raw bimodal image sequence includes a spatially aligned raw visible light image sequence and a raw radar image sequence; A preprocessing unit is used to preprocess the original bimodal image sequence to obtain a preprocessed bimodal image sequence; The first tracking unit is used to track a preset target in the preprocessed bimodal image sequence using a target tracking model, and obtain a bimodal tracking result set and a fused tracking result set for the preset target. The second tracking unit is used to perform tracking processing on the preset target according to the preset target motion model, the dual-modal tracking result set and the fused tracking result set, to obtain the final tracking result set of the preset target.

[0021] In the above technical solution, the device can achieve dynamic adaptive fusion of visible light and radar modes by comprehensively evaluating all dual-mode tracking results and fused tracking results, and optimizing them in combination with the target motion model. This effectively avoids tracking drift when a single mode fails, and significantly improves the accuracy, robustness and reliability of target tracking in complex environments.

[0022] Thirdly, this application provides an electronic device including a memory and a processor, the memory storing a computer program, and the processor running the computer program to cause the electronic device to perform the radar target tracking method described in any one of the first aspects.

[0023] Fourthly, this application provides a readable storage medium storing a computer program, which, when executed by a processor, performs the radar target tracking method described in any one of the first aspects.

[0024] Fifthly, this application provides a computer program product comprising a computer program that, when executed by a processor, performs the radar target tracking method described in any one of the first aspects.

[0025] The beneficial effects of this application are: it can utilize a three-branch parallel tracking architecture, combined with confidence-weighted feature fusion and Kalman filter state optimization, to achieve deep complementarity and adaptive fusion of dual-modal information. This not only weakens the interference of low-quality modes through dynamic weight allocation, but also effectively suppresses tracking noise using motion models. In this way, it significantly improves the robustness, accuracy and stability of target tracking in complex environments, and avoids target loss or drift. Attached Figure Description

[0026] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a flowchart illustrating a radar target tracking method according to an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a radar target tracking device according to an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0028] The embodiments of the technical solution of this application will now be described in detail with reference to the accompanying drawings. These embodiments are only used to more clearly illustrate the technical solution of this application and are therefore merely examples, and should not be used to limit the scope of protection of this application.

[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.

[0030] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more (including two), similarly, "multiple sets" refers to two or more sets (including two sets), and "multiple pieces" refers to two or more pieces (including two pieces) unless otherwise explicitly defined.

[0031] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0032] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.

[0033] Existing fusion tracking technologies that combine SAR (Synthetic Aperture Radar) and visible light suffer from technical defects such as rigid feature fusion, poor tracker adaptability, disconnect between preprocessing and fusion, and insufficient robustness. These defects prevent the full utilization of the complementary advantages of the two types of images and make it difficult to balance tracking accuracy and continuity in complex scenarios.

[0034] To address the aforementioned technical problems, this application provides a radar target tracking method. This method is based on a three-branch parallel tracking structure, utilizes dynamic confidence to adjust the fusion weights, and combines motion models for filtering. Ultimately, it can achieve effective complementarity of dual-modal information, thereby significantly optimizing the tracking performance.

[0035] like Figure 1 As shown, some embodiments of this application provide a radar target tracking method, which includes: S100. Obtain the original dual-modal image sequence; wherein, the original dual-modal image sequence includes a spatially aligned original visible light image sequence and an original radar image sequence.

[0036] In this embodiment, the original radar image can be a Synthetic Aperture Radar (SAR) image, an Inverse Synthetic Aperture Radar (ISAR) image, a millimeter-wave radar image, etc. The following description uses SAR images as an example.

[0037] In this embodiment, since the frame rate of visible light images is much higher than that of SAR images, and the field of view of SAR images is larger, the dual-modal images in this embodiment do not need to be time-synchronized, allowing one frame of SAR image to be registered to multiple frames of visible light images.

[0038] In this embodiment, the above spatial alignment can be achieved using a pixel-level alignment scheme, which can be any of the following: SIFT (Scale-Invariant Feature Transform) feature matching combined with homography matrix transformation, sparse feature point matching combining SuperPoint (self-supervised interest point detection and descriptor) and SuperGlue (graph neural network for feature matching), or semi-dense pixel point matching based on LoFTR.

[0039] Among them, the LoFTR algorithm is a local feature matching algorithm based on Transformers, and its full name is Detector-Free Local Feature Matching with Transformers.

[0040] As an optional implementation, acquiring the original bimodal image sequence includes: Acquire the initial visible light image sequence and the initial radar image sequence; Spatial alignment processing is performed on the initial visible light image sequence and the initial radar image sequence to obtain the original dual-modal image sequence.

[0041] In this embodiment, spatial alignment processing involves finding the same feature points in the visible light image and the radar image, calculating the positional offset, rotation angle, and size difference between the two, and adjusting the images through coordinate transformation, image correction, and other methods to ensure that the same real-world target in the two images is located at the same pixel position, achieving pixel-level precise registration and laying the foundation for position matching for subsequent dual-modal feature fusion and target tracking.

[0042] S200. Preprocess the original bimodal image sequence to obtain a preprocessed bimodal image sequence.

[0043] In this embodiment, preprocessing is a differentiated process designed for the respective imaging defects of the dual-modal images, aiming to eliminate noise and distortion, improve feature compatibility, and provide high-quality data support for subsequent tracking. Specifically, preprocessing may include, but is not limited to, image denoising, contrast enhancement, radiometric correction, geometric correction, non-uniformity correction, super-resolution reconstruction, and image registration. Through the above preprocessing, the illumination unevenness and blurring of visible light images can be effectively suppressed, and the inherent speckle noise, radiometric distortion, and geometric deformation of SAR images can be corrected.

[0044] S300. The preset target in the preprocessed dual-modal image sequence is tracked by the target tracking model to obtain the dual-modal tracking result set and the fused tracking result set of the preset target.

[0045] In this embodiment, the target tracking model can employ Siamese Region Proposal Network (Siamese-RPN), a faster Siamese Region Proposal Network (SiamRPN++), or a Transformer-based Siamese Tracking Network (SiamTAN), etc. The Siamese-RPN network will be used as an example for the following introduction.

[0046] In this embodiment, the training data of the model includes visible light and SAR dual-modal samples, covering three relationships: target motion / equipment stationary, target stationary / equipment motion, and both motion simultaneously. Target types include people, vehicles, drones, buildings, and specific areas. All samples are generated after manually annotating the target coordinates.

[0047] In this embodiment, the model fully leverages the representational advantages of single-modal data and the complementarity of cross-modal data through a three-branch architecture of bimodal independent tracking and fusion tracking, thereby ensuring the tracking adaptability of bimodal image sequences.

[0048] S400: Based on the preset target motion model, the dual-modal tracking result set, and the fused tracking result set, the preset target is tracked and processed to obtain the final tracking result set of the preset target.

[0049] In this embodiment, the target motion model can adopt a constant velocity (CV) motion model, an interactive multiple model (IMM), or other similar models. The interactive multiple model can include various motion models such as constant acceleration (CA) and coordinated turn (CT). The constant velocity motion model will be used as an example below.

[0050] In this embodiment, the tracking process is implemented using the Kalman filter algorithm. This algorithm combines motion model prediction and observation updates, which can effectively suppress noise and tracking drift, and output a stable and accurate final trajectory.

[0051] In the above embodiments, the method can achieve dynamic adaptive fusion of visible light and radar modes by comprehensively evaluating all dual-mode tracking results and fused tracking results, and optimizing them in combination with the target motion model. This effectively avoids tracking drift when a single mode fails, and significantly improves the accuracy, robustness and system reliability of target tracking in complex environments.

[0052] In some embodiments, step S200 may include: S210. Perform image enhancement processing on the original visible light image sequence to obtain an enhanced visible light image sequence.

[0053] In this embodiment, the method can employ the Adaptive Histogram Equalization (CLAHE) algorithm to perform image enhancement operations. This algorithm effectively improves the contrast between the preset target and the background while accurately preserving key visual details in the image.

[0054] In this embodiment, the method can also employ a multi-scale retinal enhancement algorithm (MSRCR) or a single-scale retinal enhancement algorithm (SSR) for image enhancement processing.

[0055] S220. Denoise the enhanced visible light image sequence to obtain a preprocessed visible light image sequence.

[0056] In this embodiment, the enhanced visible light image sequence is denoised. Specifically, noise suppression can be achieved through Gaussian bilateral filtering, while preserving the core visual features such as target texture and color during denoising.

[0057] In this embodiment, noise reduction processing of enhanced visible light image sequences can also be performed using non-local means (NLM) filtering or guided filtering.

[0058] S230. The original radar image sequence is denoised to obtain the first processed radar image sequence.

[0059] In this embodiment, the original radar image sequence is denoised by using an improved Lee filter to remove speckle noise, balancing denoising effectiveness with the preservation of target edge details.

[0060] In this embodiment, the original radar image sequence is denoised, and speckle noise suppression can also be achieved by using Gamma-MAP (Gamma Maximum A Posteriori) filtering or Non-Local Means (NLM) filtering.

[0061] S240. Correct the first processed radar image sequence to obtain the second processed radar image sequence.

[0062] In this embodiment, the correction processing of the first processed radar image sequence can be performed by converting the radar image into a standardized backscattering coefficient image through radiometric correction, thereby unifying the data scale. Alternatively, methods such as absolute radiometric calibration based on a calibrator, scene-based normalized scattering area (Sigma-Naught) correction, and antenna pattern and range attenuation compensation can also be used to correct the first processed radar image sequence.

[0063] S250. The second processed radar image sequence is corrected according to the imaging parameters of the synthetic aperture radar to obtain the preprocessed radar image sequence.

[0064] In this embodiment, the method can perform orthorectification based on SAR imaging parameters to eliminate geometric distortion and ensure spatial consistency with visible light images.

[0065] In this embodiment, the preprocessed dual-modal image sequence includes a preprocessed visible light image sequence and a preprocessed radar image sequence.

[0066] In this embodiment, the preprocessed dual-modal image sequences are denoted as Vis-Proc (visible light) and SAR-Proc (radar), respectively. They maintain pixel-level spatial alignment to provide a unified benchmark for subsequent tracking.

[0067] In the above embodiments, the method can effectively eliminate imaging noise and distortion differences between the two types of images by performing targeted enhancement and denoising on visible light images, and sequentially performing speckle suppression, radiometric correction and geometric correction on radar images, thereby ensuring the registration accuracy and information quality of the preprocessed dual-modal images.

[0068] In some embodiments, step S300 may include: S310. Determine a first initial target box corresponding to a preset target in the initial frame of the preprocessed visible light image sequence, and determine a second initial target box corresponding to the first initial target box in the initial frame of the preprocessed radar image sequence.

[0069] In this embodiment, since the preprocessed visible light image sequence and the preprocessed radar image sequence have been spatially aligned, the position coordinates of the second initial target box are consistent with the position coordinates of the first initial target box. The initial target box is initialized manually or automatically, denoted as ROI-Init (Region of Interest Initialization).

[0070] S320. Track the preset targets in the preprocessed visible light image sequence using the target tracking model and the first initial target box to obtain a visible light tracking result set.

[0071] In this embodiment, the target tracking model includes at least a template branch, a search branch, a classification branch, and a regression branch.

[0072] In this embodiment, in each frame, the model evaluates multiple candidate region features to obtain multiple candidate tracking results, calculates the confidence score of each candidate tracking result, and then selects the most likely optimal candidate result for the target based on these scores (usually the highest score).

[0073] To improve the positioning accuracy and stability of visible light tracking results, step S320 may further include: S321. Extract visual features within the first initial target box from the preprocessed visible light image sequence through template branching.

[0074] In this embodiment, the visual features include the texture features, color features, and edge features of the target, which are used to provide a template basis for target matching in subsequent frames.

[0075] S322. Determine any frame in the subsequent frames of the preprocessed visible light image sequence as the first target frame, and extract multiple first candidate features corresponding to the first target frame based on visual features through the search branch.

[0076] In this embodiment, the search branch can perform feature retrieval in the first target frame by scanning a sliding window, and extract multiple candidate region features related to the visual features of the template.

[0077] S323. Perform cross-correlation operation on visual features and multiple first candidate features to obtain multiple first candidate tracking results; the multiple first candidate tracking results correspond one-to-one with the multiple first candidate features.

[0078] In this embodiment, the method can measure the similarity between visual features and each first candidate feature through cross-correlation operation, generate target region proposals, and then obtain preliminary tracking results corresponding to each candidate region.

[0079] S324. Calculate multiple first confidence scores by classifying branches and multiple first candidate tracking results; the multiple first confidence scores correspond one-to-one with the multiple first candidate tracking results.

[0080] In this embodiment, the classification branch can calculate the confidence score corresponding to each first candidate tracking result based on the feature matching degree. The higher the score, the greater the probability that the corresponding candidate region is the preset target.

[0081] S325. Based on multiple first confidence scores, determine the first optimal candidate result corresponding to the first target frame from multiple first candidate tracking results.

[0082] In this embodiment, the method can select the candidate tracking result with the highest confidence score as the optimal result of the first target frame, thereby ensuring the accuracy of target localization.

[0083] S326. After determining all the first optimal candidate results corresponding to the subsequent frames of the preprocessed visible light image sequence, the first optimal candidate results are processed by the regression branch to obtain the first tracking result set of the preset target in the subsequent frames.

[0084] In this embodiment, the regression branch can refine the coordinates (i.e., the target box position coordinates) and scale (i.e., the target scale, referring to the size of the target box) of the optimal candidate result (i.e., the tracking result that best fits the preset target from multiple candidate results) to eliminate the positioning deviation of the candidate region, thereby outputting the accurate position and scale information of the preset target in each frame and forming the first tracking result set.

[0085] In this embodiment, the visible light tracking result set includes the initial tracking result obtained based on the first initial target box and the first tracking result set.

[0086] In this embodiment, the initial tracking result is the target result directly determined by the first initial target box in the initial frame; the first optimal candidate result is the optimal tracking result selected from many candidate results in each subsequent frame.

[0087] In this embodiment, each frame of the visible light tracking result set (denoted as Res_Vis) corresponds to a set of data (which corresponds to a visible light tracking result), including the visible light confidence score (Score_Vis), the visible light target bounding box position coordinates (x1_vis, y1_vis, x2_vis, y2_vis), and the visible light target scale (s_vis).

[0088] In this embodiment, (x1_vis, y1_vis) are the pixel coordinates of the top left corner of the visible light target box; (x2_vis, y2_vis) are the pixel coordinates of the lower right corner of the visible light target bounding box.

[0089] S330. The preset targets in the preprocessed radar image sequence are tracked using the target tracking model and the second initial target box to obtain a radar tracking result set.

[0090] In this embodiment, the core process of radar tracking is basically the same as that of visible light tracking. The key difference lies in the feature types extracted by the radar tracking process, which are adapted to the imaging characteristics of the radar image (SAR image) to meet the requirements for representing the scattering characteristics of the radar image.

[0091] In this embodiment, the dual-modal tracking result set includes a visible light tracking result set and a radar tracking result set. To effectively generate a high-precision radar tracking result set, step S330 may further include: S331. Through template branching, extract microwave scattering features within the second initial target box from the initial frame of the preprocessed radar image sequence.

[0092] In this embodiment, microwave scattering features include characterization information such as the target's outline and structure, which is adapted to the all-day, all-weather imaging advantages of SAR images.

[0093] S332. Determine any frame in the subsequent frames of the preprocessed radar image sequence as the second target frame, and extract multiple second candidate features corresponding to the second target frame based on microwave scattering characteristics through a search branch.

[0094] In this embodiment, the search branch can slide and scan in the second target frame to extract multiple candidate region features that match the template microwave scattering features.

[0095] S333. Based on the microwave scattering characteristics and multiple second candidate characteristics, generate multiple second candidate tracking results; the multiple second candidate tracking results correspond one-to-one with the multiple second candidate characteristics.

[0096] In this embodiment, the method can generate target region proposals through feature matching metrics and obtain preliminary tracking results for each candidate region.

[0097] S334. Calculate multiple second confidence scores by using classification branches and multiple second candidate tracking results; the multiple second confidence scores correspond one-to-one with the multiple second candidate tracking results.

[0098] In this embodiment, the classification branch can calculate the radar confidence score (denoted as Score_SAR) of the radar mode, quantifying the reliability of the candidate results.

[0099] S335. Based on multiple second confidence scores, determine the second optimal candidate result corresponding to the second target frame from multiple second candidate tracking results.

[0100] In this embodiment, the method can select the candidate tracking result with the highest confidence score as the optimal result of the second target frame, thus ensuring the accuracy of radar tracking.

[0101] S336. After determining all the second optimal candidate results corresponding to the subsequent frames of the preprocessed radar image sequence, the second optimal candidate results are processed by the regression branch to obtain the second tracking result set of the preset target in the subsequent frames of the preprocessed radar image sequence.

[0102] In this embodiment, the radar tracking result set includes the initial tracking result obtained based on the second initial target box and the second tracking result set.

[0103] In this embodiment, the initial tracking result is the target result directly determined by the second initial target box in the initial frame; the second optimal candidate result is the optimal tracking result selected from many candidate results in each subsequent frame.

[0104] In this embodiment, each frame of the radar tracking result set (denoted as Res_SAR) corresponds to a set of data (which corresponds to a radar tracking result), specifically including the radar confidence score (Score_SAR), the radar target box position coordinates (x1_sar, y1_sar, x2_sar, y2_sar), and the radar target scale (s_sar).

[0105] In this embodiment, (x1_sar, y1_sar) are the pixel coordinates of the top left corner of the radar target box; (x2_sar, y2_sar) are the pixel coordinates of the lower right corner of the radar target bounding box.

[0106] In this embodiment, the radar tracking result set includes an initial tracking result set and a second tracking result set. The dual-mode tracking result set integrates the tracking results from visible light and radar.

[0107] The bimodal tracking result set includes a set of data for each frame (each set of data corresponds to a bimodal tracking result), including the confidence scores of each frame for both modalities (the confidence scores of each frame for both modalities are collectively referred to as tracking confidence scores), the target bounding box coordinates, and the target scale.

[0108] In this embodiment, the dual-modal tracking result set serves as a unified input benchmark to support subsequent cross-modal feature fusion and multi-source trajectory optimization processing.

[0109] S340. Extract features from the preprocessed dual-modal image sequence to obtain a fused feature sequence.

[0110] In this embodiment, the method uses a joint strategy of confidence weighting and channel concatenation to generate a fused feature sequence, and achieves adaptive complementarity of the feature layer through dynamic weighting of modal confidence, thereby improving the robustness of representation.

[0111] To integrate the complementary advantages of the two modes, step S340 may further include: S341. Obtain the first target region feature map corresponding to each frame in the preprocessed visible light image sequence based on the first initial target box, and obtain the second target region feature map corresponding to each frame in the preprocessed radar image sequence based on the second initial target box.

[0112] In this embodiment, both the first target region feature map (denoted as Feat_Vis) and the second target region feature map (denoted as Feat_SAR) are generated by the template branch of Siamese-RPN, and the output is uniformly a feature tensor of C×H×W dimensions, which ensures the compatibility of subsequent fusion of dual-modal features in terms of dimensionality.

[0113] S342. Calculate the weight of the corresponding first target region feature map based on the tracking confidence score of each frame in the visible light tracking result set, and obtain the first feature weight of the first target region feature map for each frame. Calculate the weight of the corresponding second target region feature map based on the tracking confidence score of each frame in the radar tracking result set, and obtain the second feature weight of the second target region feature map for each frame.

[0114] In this embodiment, the feature weights are normalized and allocated based on the modal confidence scores. The first feature weight (Weight_Vis) is the ratio of the visible light confidence score (Score_Vis) to the total dual-modal confidence score, and the second feature weight (Weight_SAR) is the complement of the first feature weight (Weight_Vis), as shown in the formula: Weight_Vis=Score_Vis / (Score_Vis+Score_SAR); Weight_SAR = 1 - Weight_Vis; The total confidence score for the dual-modal model is the sum of Score_Vis and Score_SAR. Based on this, real-time and reliable dynamic weighting can be achieved.

[0115] S343. Perform one-to-one feature weighting processing based on the first target region feature map and the first feature weight corresponding to each frame to obtain the first weighted feature map corresponding to each frame, and perform one-to-one feature weighting processing based on the second target region feature map and the second feature weight corresponding to each frame to obtain the second weighted feature map corresponding to each frame.

[0116] In this embodiment, the method can perform weighted scaling on the bimodal feature map by element-wise multiplication, amplifying the contribution of the high-confidence mode at the feature level while suppressing the interference of the low-confidence mode.

[0117] S344. Perform a one-to-one correspondence stitching process on the first weighted feature map and the second weighted feature map corresponding to each frame in the channel dimension to obtain the fused feature map corresponding to each frame.

[0118] In this embodiment, the expression for the fused feature map (denoted as Feat_Fuse) is: Weight_Vis×Feat_Vis;Weight_SAR×Feat_SAR; The semicolon represents the concatenation of channel dimensions; Weight_Vis is the weight of the first feature; Feat_Vis is the feature map of the first target region; Weight_SAR is the weight of the second feature; Feat_SAR is the feature map of the second target region; Since the bimodal features output by the Siamese-RPN template branch have maintained consistent dimensions, no additional dimensional adjustment is required after splicing, thus allowing direct adaptation to subsequent target matching processes.

[0119] In this embodiment, the fused feature sequence includes fused feature maps corresponding to all frames.

[0120] S350. Target tracking processing of the preset target is performed through the target tracking model and fused feature sequence to obtain a fused tracking result set.

[0121] In this embodiment, the method can further integrate the advantages of dual-modality tracking based on the fusion feature sequence, thereby improving the robustness of the tracking results.

[0122] To further improve the accuracy and stability of target tracking, step S350 may also include: S351. Determine the third initial target box corresponding to the first initial target box in the initial frame of the fused feature sequence.

[0123] In this embodiment, the coordinates of the third initial target box are consistent with those of the first initial target box to ensure the consistency of tracking initialization.

[0124] S352. The fusion feature sequence is matched and tracked by the template branch and search branch of the target tracking model and the third initial target box to obtain the fusion tracking results corresponding to multiple frames; wherein, the fusion tracking results include at least the fusion confidence score, the fusion target box position coordinates and the fusion target scale.

[0125] In this embodiment, the search branch can perform target matching retrieval on the fused feature map, the classification branch calculates and outputs the fusion confidence score (Score_Fuse), and the regression branch outputs the fused target box location coordinates (x1_fuse, y1_fuse, x2_fuse, y2_fuse) and the fused target scale (s_fuse), thereby forming a single-frame fusion tracking result.

[0126] In this embodiment, the fusion tracking result set includes the fusion tracking results corresponding to all frames.

[0127] In this embodiment, each frame of the fused tracking result set (denoted as Res_Fuse) corresponds to a set of data (which corresponds to a fused tracking result), specifically including the fused confidence score (Score_Fuse), the fused target bounding box position coordinates (x1_fuse, y1_fuse, x2_fuse, y2_fuse), and the fused target scale (s_fuse), providing high-quality observation basis for subsequent optimization of the final tracking state of the preset target.

[0128] In the above embodiments, the method can utilize the multi-branch structure of the target tracking model to complete single-mode tracking of visible light and radar respectively, and achieve joint tracking based on the fused feature sequence. It effectively integrates the advantages of dual-mode data, avoids the limitations of single-mode information, and significantly improves the accuracy of dual-mode tracking result set and fused tracking result set.

[0129] In some embodiments, step S400 may include: S410. Based on the dual-modal tracking result set and the fused tracking result set, construct multi-source observation vectors corresponding to multiple frames.

[0130] In this embodiment, the multi-source observation vector (Z_k) of the k-th frame is defined as follows: (Z_k)=[Res_Vis,Res_SAR,Res_Fuse]^T= [Score_Vis,x_vis,y_vis,s_vis,Score_SAR,x_sar,y_sar,s_sar,Score_Fuse,x_fuse,y_fuse,s_fuse]^T; Where x and y are the center coordinates of the target bounding box, s is the target scale, Res_Vis is the visible light tracking result set, Res_SAR is the radar tracking result set, and Res_Fuse is the fused tracking result set; Among them, Res_Vis (visible light tracking result set) includes the visible light confidence score (Score_Vis), the center coordinates of the visible light target box (x_vis, y_vis), and the visible light target scale (s_vis). Res_SAR (radar tracking results set) includes radar confidence score (Score_SAR), radar target box center coordinates (x_sar, y_sar), and radar target scale (s_sar). Res_Fuse (fusion tracking result set) includes fusion confidence score (Score_Fuse), fusion target bounding box center coordinates (x_fuse, y_fuse), and fusion target scale (s_fuse).

[0131] In this embodiment, Z_k can be split into three independent data parts: The visible light observation component Z_vis corresponding to the visible light tracking result of the kth frame; The radar observation component Z_sar corresponding to the radar tracking result of the kth frame; The fused observation component Z_fuse corresponding to the fused tracking result of the k-th frame.

[0132] S420. Generate an initial state vector based on the initial observation vector corresponding to the initial frame in the multi-source observation vector.

[0133] In this embodiment, the state vector (X_k) of the k-th frame is defined as follows: (X_k)=[x,y,s,v_x,v_y,v_s]^T; Where x and y are the center coordinates of the target bounding box, s is the target scale, v_x and v_y are the motion velocities of the center coordinates, and v_s is the scale change rate.

[0134] In this embodiment, the initial state vector is initialized with the position and scale information of the initial observation vector, and the initial values ​​of v_x, v_y, and v_s are uniformly set to 0.

[0135] S430. Perform recursive prediction processing based on the preset target motion model and initial state vector to obtain state prediction values ​​corresponding to multiple frames.

[0136] In this embodiment, the method uses a uniform motion model (i.e., a target motion model) for recursive prediction, wherein the state transition equation is: X_k = F × X_{k-1} + W_{k-1}; Where F is the state transition matrix and W_{k-1} is the process noise that follows a Gaussian distribution.

[0137] In this embodiment, the method can recursively calculate the state prediction value X_{k|k-1} of the current frame (i.e., the k-th frame) and its corresponding prediction error covariance matrix P_{k|k-1} by the previous frame state X_{k-1}.

[0138] S440. Based on the fusion confidence scores of multiple frames in the dual-modal tracking result set and the fusion tracking result set, perform one-to-one weighted fusion processing on the multi-source observation vectors corresponding to multiple frames to obtain the weighted observation values ​​corresponding to multiple frames.

[0139] In this embodiment, the observation weights are calculated as follows: Weight_Vis'=Score_Vis / (Score_Vis+Score_SAR+Score_Fuse); Similarly, for Weight_SAR' and Weight_Fuse', the weighted observation (Z_k') is obtained by weighted summation of the observation vector and the corresponding weight; Among them, Weight_Vis' represents the visible light observation weight; Weight_SAR' represents the radar observation weight; and Weight_Fuse' represents the fused observation weight.

[0140] In this embodiment, the method employs a three-modal normalized weighting strategy, the calculation formula of which is as follows: Weight_Vis'=Score_Vis / (Score_Vis+Score_SAR+Score_Fuse); Weight_SAR'=Score_SAR / (Score_Vis+Score_SAR+Score_Fuse); Weight_Fuse'=Score_Fuse / (Score_Vis+Score_SAR+Score_Fuse); The weighted observation value Z_k' of the current frame (i.e., the k-th frame) is obtained by summing the products of each modal observation component and its corresponding weight, i.e.: Z_k'=Weight_Vis'×Z_vis+Weight_SAR'×Z_sar+Weight_Fuse'×Z_fuse; Where Z_k' represents the weighted observation value of the current frame; Weight_Vis' represents the visible light observation weight; Z_vis represents the visible light observation component corresponding to the visible light tracking result of the kth frame; Weight_SAR' represents the radar observation weight; Z_sar represents the radar observation component corresponding to the radar tracking result of the kth frame; 'Weight_Fuse' represents the weight of the fused observations; Z_fuse represents the fused observation component corresponding to the fused tracking result of the k-th frame.

[0141] S450. Perform state update processing based on the Kalman filter algorithm on the state prediction values ​​and weighted observation values ​​corresponding to multiple frames to obtain the optimal estimate of the target state corresponding to multiple frames.

[0142] In this embodiment, the Kalman filter update process consists of two steps: First, calculate the Kalman gain K_k, using the following formula: K_k=P_{k|k-1}×H^T×(H×P_{k|k-1}×H^T+R)^{-1}; Where H is the observation matrix and R is the observation noise covariance matrix; Then, the state is corrected through gain, and the state update formula is: X_{k|k}=X_{k|k-1}+K_k×(Z_k'-H×X_{k|k-1}); Finally, the optimal estimate of the target state X_{k|k} is obtained.

[0143] S460. Based on the optimal estimation of the target state corresponding to multiple frames, determine the final tracking result of the preset target in multiple frames; wherein, the final tracking result includes the final tracking position and the final scale.

[0144] In this embodiment, the method can extract the optimal target box position coordinates (x_opt, y_opt, w_opt, h_opt) and the optimal target scale (s_opt) from the optimal target state estimate, thereby obtaining the final tracking result of a single frame.

[0145] The final tracking result for a single frame includes the final tracking position and the final scale. The final tracking position includes the optimal target box position coordinates (x_opt, y_opt, w_opt, h_opt), and the final scale includes the optimal target scale (s_opt).

[0146] In this embodiment, the final tracking result set includes the final tracking results of the preset target in multiple frames.

[0147] In this embodiment, the target tracking model is used to generate a visible light tracking result set (Res_Vis), a radar tracking result set (Res_SAR), and a fused tracking result set (Res_Fuse); the target motion model is used to perform weighted fusion and Kalman filtering on the above three tracking result sets to obtain the final tracking result set. The final tracking result set can form a complete target tracking trajectory, thereby achieving stable and continuous tracking in complex scenarios.

[0148] In the above embodiments, the method can use the Kalman filter algorithm, combined with the target motion model, to predict and update multi-source tracking observation information. By using confidence-weighted fusion to weaken the interference of low-quality tracking results, the target state can be accurately estimated, thereby effectively solving the tracking drift or loss problem and ensuring the accuracy and continuity of the final tracking results.

[0149] like Figure 2 As shown, some embodiments of this application provide a schematic diagram of the structure of a radar target tracking device. It should be understood that this device is related to... Figure 1 The method executed in the middle corresponds to the steps involved in the aforementioned method. The specific functions and effects of the device can be found in the description above. To avoid repetition, detailed descriptions are omitted here.

[0150] The radar target tracking device includes: The acquisition unit 510 is used to acquire the original dual-modal image sequence; wherein the original dual-modal image sequence includes a spatially aligned original visible light image sequence and an original radar image sequence; The preprocessing unit 520 is used to preprocess the original bimodal image sequence to obtain a preprocessed bimodal image sequence; The first tracking unit 530 is used to track a preset target in a preprocessed dual-modal image sequence using a target tracking model, and obtain a dual-modal tracking result set and a fused tracking result set of the preset target. The second tracking unit 540 is used to perform tracking processing on the preset target according to the preset target motion model, the dual-modal tracking result set and the fused tracking result set, so as to obtain the final tracking result set of the preset target.

[0151] In some embodiments, the preprocessing unit 520 includes: Image enhancement subunit 521 is used to perform image enhancement processing on the original visible light image sequence to obtain an enhanced visible light image sequence; The noise suppression subunit 522 is used to perform noise reduction processing on the enhanced visible light image sequence to obtain a preprocessed visible light image sequence. The noise suppression subunit 522 is also used to perform noise suppression processing on the original radar image sequence to obtain the first processed radar image sequence; The radiation correction subunit 523 is used to perform correction processing on the first processed radar image sequence to obtain the second processed radar image sequence. The geometric correction subunit 524 is used to correct the second processed radar image sequence according to the imaging parameters of the synthetic aperture radar to obtain a preprocessed radar image sequence. The preprocessed dual-modal image sequence includes a preprocessed visible light image sequence and a preprocessed radar image sequence.

[0152] In some embodiments, the first tracking unit 530 includes: Subunit 531 is configured to determine a first initial target box corresponding to a preset target in the initial frame of the preprocessed visible light image sequence, and to determine a second initial target box corresponding to the first initial target box in the initial frame of the preprocessed radar image sequence. The target tracking subunit 532 is used to track a preset target in the preprocessed visible light image sequence using a target tracking model and a first initial target box, and obtain a visible light tracking result set. The target tracking subunit 532 is also used to track preset targets in the preprocessed radar image sequence through a target tracking model and a second initial target box to obtain a radar tracking result set. The dual-mode tracking result set includes a visible light tracking result set and a radar tracking result set; The feature extraction subunit 533 is used to extract features from the preprocessed dual-modal image sequence to obtain a fused feature sequence; The fusion tracking subunit 534 is used to perform target tracking processing on a preset target using a target tracking model and a fusion feature sequence to obtain a fusion tracking result set.

[0153] The target tracking model includes at least a template branch, a search branch, a classification branch, and a regression branch.

[0154] In some embodiments, the target tracking subunit 532 includes: The first extraction module is used to extract visual features within a first initial target box from a preprocessed visible light image sequence through template branches; The first determining module is used to determine any frame in the subsequent frames of the preprocessed visible light image sequence as the first target frame, and extract multiple first candidate features corresponding to the first target frame based on visual features through a search branch. The computation module is used to perform cross-correlation operations on visual features and multiple first candidate features to obtain multiple first candidate tracking results; the multiple first candidate tracking results correspond one-to-one with the multiple first candidate features. The first calculation module is used to calculate multiple first confidence scores by using classification branches and multiple first candidate tracking results; the multiple first confidence scores correspond one-to-one with the multiple first candidate tracking results; The first filtering module is used to determine the first optimal candidate result corresponding to the first target frame from multiple first candidate tracking results based on multiple first confidence scores; The first processing module is used to process all the first optimal candidate results corresponding to the subsequent frames of the preprocessed visible light image sequence through a regression branch after determining all the first optimal candidate results, so as to obtain the first tracking result set of the preset target in the subsequent frames. The visible light tracking result set includes the initial tracking result obtained based on the first initial target box and the first tracking result set.

[0155] In some embodiments, the target tracking subunit 532 further includes: The second extraction module is used to extract microwave scattering features within the second initial target box from the initial frame of the preprocessed radar image sequence through template branching. The second determining module is used to determine any frame in the subsequent frames of the preprocessed radar image sequence as the second target frame, and to extract multiple second candidate features corresponding to the second target frame based on microwave scattering features through a search branch. The generation module is used to generate multiple second candidate tracking results based on microwave scattering characteristics and multiple second candidate features; the multiple second candidate tracking results correspond one-to-one with the multiple second candidate features; The second calculation module is used to calculate multiple second confidence scores by using classification branches and multiple second candidate tracking results; the multiple second confidence scores correspond one-to-one with the multiple second candidate tracking results; The second filtering module is used to determine the second optimal candidate result corresponding to the second target frame from multiple second candidate tracking results based on multiple second confidence scores. The second processing module is used to process all the second optimal candidate results corresponding to the subsequent frames of the preprocessed radar image sequence through a regression branch after determining all the second optimal candidate results, so as to obtain the second tracking result set of the preset target in the subsequent frames of the preprocessed radar image sequence. The radar tracking result set includes the initial tracking result obtained based on the second initial target box and the second tracking result set.

[0156] In some embodiments, the feature extraction subunit 533 includes: The acquisition module is used to acquire a first target region feature map corresponding to each frame in the preprocessed visible light image sequence based on a first initial target box, and to acquire a second target region feature map corresponding to each frame in the preprocessed radar image sequence based on a second initial target box; The third calculation module is used to calculate the weight of the corresponding first target region feature map based on the tracking confidence score of each frame in the visible light tracking result set, to obtain the first feature weight of the first target region feature map for each frame, and to calculate the weight of the corresponding second target region feature map based on the tracking confidence score of each frame in the radar tracking result set, to obtain the second feature weight of the second target region feature map for each frame. The weighting module is used to perform one-to-one feature weighting processing based on the first target region feature map and the first feature weight corresponding to each frame to obtain the first weighted feature map corresponding to each frame, and to perform one-to-one feature weighting processing based on the second target region feature map and the second feature weight corresponding to each frame to obtain the second weighted feature map corresponding to each frame. The stitching module is used to perform one-to-one stitching of the first weighted feature map and the second weighted feature map corresponding to each frame in the channel dimension to obtain the fused feature map corresponding to each frame. The fusion feature sequence includes fusion feature maps corresponding to all frames.

[0157] In some embodiments, the fusion tracking subunit 534 includes: The setting module is used to determine the third initial target box corresponding to the first initial target box in the initial frame of the fused feature sequence; The tracking module is used to match and track the fused feature sequence through the template branch and search branch of the target tracking model and the third initial target box to obtain the fused tracking results corresponding to multiple frames; wherein, the fused tracking results include at least the fused confidence score, the fused target box position coordinates and the fused target scale; The fusion tracking result set includes the fusion tracking results for all frames.

[0158] In some embodiments, the second tracking unit 540 includes: Subunit 541 is constructed to construct multi-source observation vectors corresponding to multiple frames based on the dual-modal tracking result set and the fused tracking result set. The generation subunit 542 is used to generate an initial state vector based on the initial observation vector corresponding to the initial frame in the multi-source observation vector; Prediction subunit 543 is used to perform recursive prediction processing based on a preset target motion model and initial state vector to obtain state prediction values ​​for multiple frames. The weighted fusion subunit 544 is used to perform one-to-one weighted fusion processing on the multi-source observation vectors corresponding to multiple frames based on the fusion confidence scores of multiple frames in the dual-modal tracking result set and the fusion tracking result set, to obtain the weighted observation values ​​corresponding to multiple frames. The update subunit 545 is used to perform state update processing based on the Kalman filter algorithm on the state prediction values ​​and weighted observation values ​​corresponding to multiple frames to obtain the optimal estimate of the target state corresponding to multiple frames. The determination subunit 546 is used to determine the final tracking result of the preset target in multiple frames based on the optimal estimation of the target state corresponding to multiple frames; wherein, the final tracking result includes the final tracking position and the final scale; The final tracking result set includes the final tracking results of the preset target in multiple frames.

[0159] like Figure 3 As shown, this application provides an electronic device 600, which includes a processor 601 and a memory 602. The processor 601 and the memory 602 are interconnected and communicate with each other through a communication bus 603 and / or other forms of connection mechanism (not shown). The memory 602 stores a computer program that can be executed by the processor 601. When the computing device is running, the processor 601 executes the computer program to perform the method in any of the aforementioned optional implementations.

[0160] This application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the method in any of the aforementioned optional implementations.

[0161] The computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0162] This application provides a computer program product, which includes a computer program that, when run by a processor, executes the method in any of the aforementioned optional implementations.

[0163] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and not to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application, and they should all be covered within the scope of the claims and specification of this application. In particular, as long as there is no conflict, the various technical features mentioned in the embodiments can be combined in any way. This application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A radar target tracking method, characterized in that, include: Obtain the original bimodal image sequence; wherein, the original bimodal image sequence includes a spatially aligned original visible light image sequence and an original radar image sequence; The original bimodal image sequence is preprocessed to obtain a preprocessed bimodal image sequence; wherein, the preprocessed bimodal image sequence includes a preprocessed visible light image sequence and a preprocessed radar image sequence; The preset target in the preprocessed bimodal image sequence is tracked by the target tracking model to obtain the bimodal tracking result set and the fused tracking result set of the preset target; The preset target is tracked based on the preset target motion model, the dual-modal tracking result set, and the fused tracking result set to obtain the final tracking result set of the preset target. The step of tracking a preset target in the preprocessed bimodal image sequence using a target tracking model to obtain a bimodal tracking result set and a fused tracking result set for the preset target includes: In the initial frame of the preprocessed visible light image sequence, a first initial target box corresponding to the preset target is determined, and in the initial frame of the preprocessed radar image sequence, a second initial target box corresponding to the first initial target box is determined; The preset targets in the preprocessed visible light image sequence are tracked using a target tracking model and the first initial target bounding box to obtain a visible light tracking result set; The preset target in the preprocessed radar image sequence is tracked using a target tracking model and the second initial target bounding box to obtain a radar tracking result set; The dual-modal tracking result set includes the visible light tracking result set and the radar tracking result set; Feature extraction is performed on the preprocessed bimodal image sequence to obtain a fused feature sequence; The target tracking process for the preset target is performed using the target tracking model and the fused feature sequence to obtain a fused tracking result set; The step of extracting features from the preprocessed bimodal image sequence to obtain a fused feature sequence includes: Based on the first initial target bounding box, obtain the first target region feature map corresponding to each frame in the preprocessed visible light image sequence, and based on the second initial target bounding box, obtain the second target region feature map corresponding to each frame in the preprocessed radar image sequence. The weights of the corresponding first target region feature maps are calculated based on the tracking confidence scores of each frame in the visible light tracking result set, resulting in the first feature weights of the first target region feature maps for each frame. The weights of the corresponding second target region feature maps are also calculated based on the tracking confidence scores of each frame in the radar tracking result set, resulting in the second feature weights of the second target region feature maps for each frame. Based on the first target region feature map and the first feature weight corresponding to each frame, a one-to-one feature weighting process is performed to obtain the first weighted feature map corresponding to each frame. Based on the second target region feature map and the second feature weight corresponding to each frame, a one-to-one feature weighting process is performed to obtain the second weighted feature map corresponding to each frame. The first and second weighted feature maps corresponding to each frame are concatenated one-to-one along the channel dimension to obtain the fused feature map corresponding to each frame. The fusion feature sequence includes fusion feature maps corresponding to all frames; The step of tracking the preset target based on the preset target motion model, the dual-modal tracking result set, and the fused tracking result set to obtain the final tracking result set of the preset target includes: Based on the dual-modal tracking result set and the fused tracking result set, construct multi-source observation vectors corresponding to multiple frames; An initial state vector is generated based on the initial observation vector corresponding to the initial frame in the multi-source observation vector; Recursive prediction processing is performed based on the preset target motion model and the initial state vector to obtain state prediction values ​​for multiple frames. Based on the dual-modal tracking result set and the fusion confidence scores of multiple frames in the fusion tracking result set, the multi-source observation vectors corresponding to the multiple frames are subjected to one-to-one weighted fusion processing to obtain the weighted observation values ​​corresponding to the multiple frames. The state prediction values ​​and weighted observation values ​​corresponding to the multiple frames are processed by a state update algorithm based on Kalman filtering to obtain the optimal estimate of the target state corresponding to the multiple frames. Based on the optimal estimation of the target state corresponding to the multiple frames, the final tracking result corresponding to the preset target in the multiple frames is determined; wherein, the final tracking result includes the final tracking position and the final scale; The final tracking result set includes the final tracking results of the preset target in multiple frames.

2. The radar target tracking method according to claim 1, characterized in that, The preprocessing of the original bimodal image sequence to obtain a preprocessed bimodal image sequence includes: The original visible light image sequence is subjected to image enhancement processing to obtain an enhanced visible light image sequence; The enhanced visible light image sequence is denoised to obtain a preprocessed visible light image sequence; The original radar image sequence is denoised to obtain a first processed radar image sequence; The first processed radar image sequence is corrected to obtain the second processed radar image sequence. The second processed radar image sequence is corrected according to the imaging parameters of synthetic aperture radar to obtain a preprocessed radar image sequence. The preprocessed dual-modal image sequence includes the preprocessed visible light image sequence and the preprocessed radar image sequence.

3. The radar target tracking method according to claim 1, characterized in that, The target tracking model includes at least a template branch, a search branch, a classification branch, and a regression branch; The step of tracking a preset target in the preprocessed visible light image sequence using a target tracking model and the first initial target bounding box to obtain a visible light tracking result set includes: Visual features within the first initial target box are extracted from the preprocessed visible light image sequence using the template branch; Any frame in the subsequent frames of the preprocessed visible light image sequence is determined as the first target frame, and multiple first candidate features corresponding to the first target frame are extracted based on the visual features through the search branch. A cross-correlation operation is performed on the visual features and the plurality of first candidate features to obtain a plurality of first candidate tracking results; the plurality of first candidate tracking results correspond one-to-one with the plurality of first candidate features. Multiple first confidence scores are obtained by calculating using the classification branch and the multiple first candidate tracking results; each of the multiple first confidence scores corresponds one-to-one with the multiple first candidate tracking results. Based on the plurality of first confidence scores, a first optimal candidate result corresponding to the first target frame is determined from the plurality of first candidate tracking results; After determining all the first optimal candidate results corresponding to the subsequent frames of the preprocessed visible light image sequence, the regression branch is used to process all the first optimal candidate results to obtain the first tracking result set of the preset target in the subsequent frames; The visible light tracking result set includes the initial tracking result obtained based on the first initial target box and the first tracking result set.

4. The radar target tracking method according to claim 3, characterized in that, The step involves tracking the preset target in the preprocessed radar image sequence using a target tracking model and the second initial target bounding box to obtain a radar tracking result set, including: Through the template branch, microwave scattering features within the second initial target box are extracted from the initial frame of the preprocessed radar image sequence; Any frame in the subsequent frames of the preprocessed radar image sequence is determined as the second target frame, and multiple second candidate features corresponding to the second target frame are extracted based on the microwave scattering features through the search branch. Based on the microwave scattering characteristics and the plurality of second candidate characteristics, a plurality of second candidate tracking results are generated; the plurality of second candidate tracking results correspond one-to-one with the plurality of second candidate characteristics. Multiple second confidence scores are obtained by calculating using the classification branch and the multiple second candidate tracking results; each of the multiple second confidence scores corresponds one-to-one with the multiple second candidate tracking results. Based on the plurality of second confidence scores, a second optimal candidate result corresponding to the second target frame is determined from the plurality of second candidate tracking results; After determining all the second optimal candidate results corresponding to the subsequent frames of the preprocessed radar image sequence, the regression branch is used to process all the second optimal candidate results to obtain the second tracking result set of the preset target in the subsequent frames of the preprocessed radar image sequence. The radar tracking result set includes the initial tracking result obtained based on the second initial target box and the second tracking result set.

5. The radar target tracking method according to claim 1, characterized in that, The target tracking process of the preset target using the target tracking model and the fused feature sequence yields a fused tracking result set, including: In the initial frame of the fused feature sequence, a third initial target box corresponding to the first initial target box is determined; The target tracking model uses template branches and search branches, along with the third initial target box, to match and track the fused feature sequence, resulting in fused tracking results for multiple frames. The fused tracking results include at least a fused confidence score, fused target box location coordinates, and fused target scale. The fusion tracking result set includes the fusion tracking results for all frames.

6. A radar target tracking device, characterized in that, The radar target tracking device includes: An acquisition unit is used to acquire a raw bimodal image sequence; wherein the raw bimodal image sequence includes a spatially aligned raw visible light image sequence and a raw radar image sequence; A preprocessing unit is used to preprocess the original bimodal image sequence to obtain a preprocessed bimodal image sequence; wherein the preprocessed bimodal image sequence includes a preprocessed visible light image sequence and a preprocessed radar image sequence; The first tracking unit is used to track a preset target in the preprocessed bimodal image sequence using a target tracking model, and obtain a bimodal tracking result set and a fused tracking result set for the preset target. The second tracking unit is used to perform tracking processing on the preset target according to the preset target motion model, the dual-modal tracking result set and the fused tracking result set, to obtain the final tracking result set of the preset target; The first tracking unit includes: A subunit is configured to determine a first initial target box corresponding to the preset target in the initial frame of the preprocessed visible light image sequence, and to determine a second initial target box corresponding to the first initial target box in the initial frame of the preprocessed radar image sequence. The target tracking subunit is used to track a preset target in the preprocessed visible light image sequence using a target tracking model and the first initial target bounding box, and to obtain a visible light tracking result set. The target tracking subunit is further configured to track the preset target in the preprocessed radar image sequence using the target tracking model and the second initial target box, to obtain a radar tracking result set; The dual-modal tracking result set includes the visible light tracking result set and the radar tracking result set; The feature extraction subunit is used to extract features from the preprocessed dual-modal image sequence to obtain a fused feature sequence; The fusion tracking subunit is used to perform target tracking processing on the preset target using the target tracking model and the fusion feature sequence to obtain a fusion tracking result set. The feature extraction subunit includes: The acquisition module is used to acquire a first target region feature map corresponding to each frame in the preprocessed visible light image sequence based on the first initial target box, and to acquire a second target region feature map corresponding to each frame in the preprocessed radar image sequence based on the second initial target box; The third calculation module is used to calculate the weight of the corresponding first target region feature map based on the tracking confidence score of each frame in the visible light tracking result set, to obtain the first feature weight of the first target region feature map for each frame, and to calculate the weight of the corresponding second target region feature map based on the tracking confidence score of each frame in the radar tracking result set, to obtain the second feature weight of the second target region feature map for each frame. The weighting module is used to perform one-to-one feature weighting processing based on the first target region feature map and the first feature weight corresponding to each frame to obtain the first weighted feature map corresponding to each frame, and to perform one-to-one feature weighting processing based on the second target region feature map and the second feature weight corresponding to each frame to obtain the second weighted feature map corresponding to each frame. The stitching module is used to perform one-to-one stitching of the first weighted feature map and the second weighted feature map corresponding to each frame in the channel dimension to obtain the fused feature map corresponding to each frame. The fusion feature sequence includes fusion feature maps corresponding to all frames; The second tracking unit includes: A subunit is constructed to construct multi-source observation vectors corresponding to multiple frames based on the dual-modal tracking result set and the fused tracking result set. A generation subunit is used to generate an initial state vector based on the initial observation vector corresponding to the initial frame in the multi-source observation vector; The prediction subunit is used to perform recursive prediction processing based on the preset target motion model and the initial state vector to obtain state prediction values ​​corresponding to multiple frames. The weighted fusion subunit is used to perform one-to-one weighted fusion processing on the multi-source observation vectors corresponding to the multiple frames based on the dual-modal tracking result set and the fusion confidence scores corresponding to the multiple frames in the fusion tracking result set, so as to obtain the weighted observation values ​​corresponding to the multiple frames. The update subunit is used to perform state update processing based on the Kalman filter algorithm on the state prediction values ​​and weighted observation values ​​corresponding to the multiple frames to obtain the optimal estimate of the target state corresponding to the multiple frames. A determining subunit is used to determine the final tracking result of the preset target in the multiple frames based on the optimal estimation of the target state corresponding to the multiple frames; wherein, the final tracking result includes the final tracking position and the final scale; The final tracking result set includes the final tracking results of the preset target in multiple frames.

7. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory being used to store a computer program, and the processor running the computer program to cause the electronic device to perform the radar target tracking method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Unmanned aerial vehicle detection system and method based on vision, laser radar and sound waves

    CN120763880A

  • Multi-modal visual fusion complex scene small target detection tracking method and system

    CN121438218A