An integrated intelligent processing method for radar target detection and tracking on a motion platform
By adding the SPD-Conv layer to the YOLOv8 network and using the ByteTrack algorithm, the problem of inaccurate object detection and tracking in complex environments is solved, and the high-precision and robust object detection and tracking effect is achieved.
Patent Information
- Application Number
- CN202510472424.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-16
AI Technical Summary
When traditional target detection and tracking methods face complex and changing target environments, they are difficult to meet the needs of motion platform radar for high-performance detection and tracking, and are prone to missed detection or loss of targets.
A integrated intelligent processing method for radar target detection and tracking of motion platform is adopted. By adding multiple SPD-Conv layers to the YOLOv8 network, multi-scale features are extracted, and target trajectory is calculated in combination with the ByteTrack algorithm to achieve efficient target tracking.
It improves the accuracy and robustness of target detection and tracking, and can accurately identify and track dynamic targets in complex environments, reducing the situation of missed detection and loss of targets.
Smart Images

Figure CN120013994B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of radar target detection, and particularly to an integrated intelligent processing method for radar target detection and tracking on a moving platform. Background Art
[0002] Target monitoring is the core function of radar. Among them, the detection and tracking of dynamic targets play a crucial role in the monitoring field and have become the focus of academic and engineering research at home and abroad. With the continuous improvement of the intelligent perception ability of mobile platform radar, multiple tasks such as obstacle avoidance, navigation, target detection and tracking can be achieved by combining visual information, which are widely used in fields such as autonomous driving, environmental perception and intelligent transportation. This poses higher requirements for moving platform radar, which not only needs to have a high degree of cognitive and execution capabilities, but also needs to have good generalization capabilities.
[0003] Traditional target detection methods use manually designed scale-invariant features or Histogram of Oriented Gradient (HOG) features to discriminate sliding windows. However, these features lack sufficient robustness in the face of diverse input changes. In traditional target tracking methods, the correlation filtering algorithm and its improved versions improve the tracking efficiency by calculating the correlation of the candidate region feature maps, but only utilize the correlation feature information between adjacent frames, resulting in easy loss of the target when the target deforms or is occluded. Therefore, traditional methods are difficult to meet the requirements of moving platform radar for high-performance detection and tracking, and are prone to target missed detection or loss of tracking. Moreover, these models are usually only applicable to detection and tracking under specific conditions. However, due to the complex and changeable target environment and diverse motion characteristics faced by moving platform radar, it further leads to inaccurate and unstable target detection and tracking. Summary of the Invention
[0004] The purpose of the present application is to provide an integrated intelligent processing method for radar target detection and tracking on a moving platform, which can improve the accuracy of target detection and tracking.
[0005] To achieve the above purpose, the present application provides the following solutions.
[0006] The present application provides an integrated intelligent processing method for radar target detection and tracking on a moving platform, including the following steps.
[0007] Obtain the echo signals of the target at multiple moments of the moving platform radar.
[0008] Preprocess the echo signals of the target at multiple moments to generate multiple frames of P-display images.
[0009] Input multiple frames of P-display images into the target detection model respectively to obtain the target detection results corresponding to each frame of P-display image; the target detection results are the detection results of multiple targets in the P-display image; the target detection model is obtained by transfer training of the improved YOLOv8 network with sample P-display images; the improved YOLOv8 network adds multiple SPD-Conv layers to the YOLOv8 network.
[0010] Based on the target detection results corresponding to each frame of P-display image, use the ByteTrack algorithm to calculate the trajectories of the targets and complete target tracking.
[0011] Optionally, the training process of the target detection model specifically includes: migrating the network parameters in the pre-trained model to the improved YOLOv8 network to obtain an initial target detection model; the structure of the pre-trained model is the same as that of the target detection model; input the sample P-display image into the initial target detection model to obtain the predicted target detection results; construct a loss function according to the predicted target detection results and the sample target detection results corresponding to the sample P-display image, and iteratively optimize the parameters of the initial target detection model according to the loss function until the loss function reaches the minimum value or the number of iterative optimization rounds reaches the maximum value, stop iterative optimization, and obtain the target detection model.
[0012] Optionally, the loss function includes: CIoU loss function and BCEWithLogits loss function.
[0013] Optionally, the calculation formula of the CIoU loss function is as follows.
[0014] 。
[0015] Where, is the CIoU loss function value; is value; is the Euclidean distance; is the center point of the bounding box in the predicted target detection result; is the center point of the bounding box in the sample target detection result; is the diagonal distance of the minimum circumscribed rectangle between the bounding box in the predicted target detection result and the bounding box in the sample target detection result; is the weight parameter; is the parameter for measuring the aspect ratio consistency.
[0016] Optionally, the calculation formula of the BCEWithLogits loss function is as follows.
[0017] 。
[0018] Among them, is the BCEWithLogits loss function value; is a fixed parameter; is the confidence in the sample object detection result; , is the class label vector of the object detection result.
[0019] Optionally, the improved YOLOv8 network includes: an input layer, a backbone network, and a head network; multiple SPD-Conv layers are added to the backbone network and the head network.
[0020] Optionally, the SPD-Conv layer includes: a space-to-depth layer and a non-strided convolution layer.
[0021] Optionally, the object detection result includes: a bounding box, a confidence, and a class label; based on the object detection result corresponding to each frame of P-display image, the ByteTrack algorithm is used to calculate the trajectory of the object to complete object tracking, specifically including: calculating the coordinates of each object in two adjacent frames of P-display images; determining the type of the bounding box according to the magnitude of the confidence of each object in the object detection result corresponding to each frame of P-display image; the types include high-confidence detection boxes and low-confidence detection boxes; for different types of bounding boxes, the Hungarian algorithm is used to match the objects in two adjacent frames of P-display images, and obtain the coordinates of the successfully matched objects and the coordinates of the unsuccessfully matched objects; the Kalman filter is used to predict the coordinates of each object in each frame of P-display image to obtain predicted coordinates; based on the predicted coordinates and the coordinates of the unsuccessfully matched objects, the Hungarian algorithm is used to match the objects in two adjacent frames of P-display images again, and obtain the coordinates of the objects that are successfully matched again; the object trajectory is determined based on the coordinates of the successfully matched objects and the coordinates of the objects that are successfully matched again to achieve object tracking.
[0022] Optionally, determining the type of the bounding box according to the magnitude of the confidence of each object in the object detection result corresponding to each frame of P-display image specifically includes: judging whether the confidence of each object in the object detection result corresponding to each frame of P-display image is greater than or equal to a preset high-frame threshold; if the confidence is greater than or equal to the preset high-frame threshold, determining that the bounding box is a high-confidence detection box; if the confidence is less than the preset high-frame threshold, then judging whether the confidence is greater than or equal to a low-frame threshold; if the confidence is greater than or equal to the low-frame threshold, determining that the bounding box is a low-confidence detection box; if the confidence is less than the low-frame threshold, then the object in the bounding box.
[0023] Optionally, when using the Hungarian algorithm to match the objects in two adjacent frames of P-display images again, if the number of matches reaches the match number threshold, the unsuccessfully matched objects are discarded.
[0024] According to the specific embodiments provided by this application, this application has the following technical effects.
[0025] This application provides an integrated intelligent processing method for radar target detection and tracking on a moving platform. By adding multiple SPD-Conv layers on the basis of the YOLOv8 network, it can extract features at different scales, capture more detailed information, and improve the detection accuracy of targets. By calculating the target detection results corresponding to each frame of P-display image through the ByteTrack algorithm, the trajectory of the target is obtained. Through an efficient data association algorithm, the targets in different frames are matched to ensure the accuracy of tracking. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0027] Figure 1 It is a schematic flowchart of an integrated intelligent processing method for radar target detection and tracking on a moving platform provided by an embodiment of this application.
[0028] Figure 2 It is a specific flowchart of an integrated intelligent processing method for radar target detection and tracking on a moving platform provided by an embodiment of this application.
[0029] Figure 3 It is a schematic diagram of an improved Yolov8 network provided by an embodiment of this application.
[0030] Figure 4 It is a schematic diagram of an SPD-Conv layer provided by an embodiment of this application.
[0031] Figure 5 It is a schematic flowchart of the ByteTrack algorithm provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] The following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the drawings in the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, rather than all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.
[0033] With the development of deep learning, due to its powerful feature learning ability, generalization ability, and interpretability, deep learning has gradually become a highly feasible method in the field of radar detection. Currently, object detection algorithms are mainly divided into one-stage algorithms and two-stage algorithms. One-stage algorithms directly extract features from images and use classifiers or regressors to predict the location and category of objects, such as YOLO and SSD. These algorithms are characterized by fast detection speed and low computational complexity, but the detection accuracy is relatively low. Two-stage algorithms divide object detection into two subtasks: object proposal and object classification, such as Faster R-CNN, etc. Their detection accuracy is higher, but the speed is slower and the computational complexity is higher.
[0034] In terms of object tracking, typical algorithms include SORT, DeepSORT, and ByteTrack. The SORT framework consists of a detector, motion estimation, and data association. It uses a Kalman filter for target motion prediction and matches detection boxes and prediction boxes through the Hungarian algorithm in data association. DeepSORT adds Mahalanobis distance and cosine distance as data association methods for cascade matching on the basis of SORT, improving the long-term association ability. ByteTrack retains low-score detection boxes and performs secondary matching, thereby reducing missed detections and improving tracking consistency. However, SORT only relies on the motion model of the Kalman filter and IoU matching. In the case of occlusion or sudden motion changes, it is prone to frequent ID switches due to the lack of appearance information, and directly discards low-score detection boxes, resulting in missed detections. Although DeepSORT reduces ID switches by introducing appearance features, it relies on a high-dimensional ReID model to calculate similarity, significantly increasing the computational cost, and appearance drift (such as lighting changes or target deformation) will lead to mis-matching, and at the same time, it may still ignore low-confidence real targets. Neither of them systematically processes low-confidence detections, while ByteTrack alleviates missed detections through a two-stage association strategy (high-score detections first and then low-score detections), but SORT and DeepSORT have weak tracking robustness in complex occlusion or target blur scenarios due to design limitations. Although certain progress has been made in the research on moving target detection and tracking, there are still many challenges in the application of moving platform radars, such as occlusion and fast target movement. Especially when it is easier for moving platform radars to lose targets compared to fixed radars, once traditional tracking algorithms lose a target, they usually cannot retrieve the target.
[0035] Therefore, improving the accuracy of object detection and tracking is an urgent problem to be solved.
[0036] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0037] In an exemplary embodiment, as Figure 1 and Figure 2As shown, an integrated intelligent processing method for radar target detection and tracking on a moving platform is provided. This method is executed by a computer device, which can be specifically executed by a computer device such as a terminal or a server alone, or jointly executed by a terminal and a server. In the embodiments of this application, taking this method applied to a server as an example for illustration, it includes the following steps 1 to step 8.
[0038] Step 1: Obtain the echo signals of the target at multiple moments of the moving platform radar. As Figure 2 the echo signals of the target are received by the moving platform radar.
[0039] Step 2: Preprocess the echo signals of the target at multiple moments to generate multiple frames of PPI images.
[0040] Specifically, the echo signals of the target of the moving platform radar are preprocessed to generate PPI images, and the PPI images (Plan Position Indicator, PPI) are drawn. The PPI image is a two-dimensional image that shows the distance and azimuth of the radar relative to its location.
[0041] 。
[0042] 。
[0043] Among them, is the angle index, representing the azimuth angle number in the radar scan data; is the coordinate of the x-axis of the PPI image; is the distance grid point; is the angle data; is the coordinate of the y-axis of the PPI image.
[0044] Step 3: Input the multiple frames of PPI images into the target detection model respectively to obtain the target detection results corresponding to each frame of PPI image; the target detection results are the detection results of multiple targets in the PPI image; the target detection model is obtained by transfer training of the improved YOLOv8 network with sample PPI images; the improved YOLOv8 network adds multiple SPD-Conv layers to the YOLOv8 network. Among them, the improved YOLOv8 network includes: an input layer, a backbone network, and a head network; multiple SPD-Conv layers are added to the backbone network and the head network. The SPD-Conv layer includes: a space-to-depth layer and a non-strided convolutional layer.
[0045] Specifically, the training process of the target detection model specifically includes the following steps.
[0046] Step 31: Transfer the network parameters in the pre-trained model to the improved YOLOv8 network to obtain an initial object detection model. The structure of the pre-trained model is the same as that of the object detection model.
[0047] When performing model transfer, the network parameters in the pre-trained model initialize the weights of the improved YOLOv8 network, replacing the original random initialization operation, and then globally fine-tuning the initial object detection model through training again. As Figure 2 shown, the source domain DATA1 is obtained through preprocessing, denoted as and the target domain DATA2, denoted as , the transfer learning studied in this application focuses on the situation with a source domain and a target domain (i.e., the sample P display image), where the source domain is represented by, , respectively represent the source domain data samples and the corresponding moving object category labels, represents the total number of source domain data samples, and the target domain is represented by , , respectively represent the target domain data samples and the corresponding moving object category labels, represents the total number of target domain data samples. Given the source domain and the learning task , the target domain and the learning task , the purpose of transfer learning is to obtain the knowledge in the source domain and the learning task to help improve the learning of the improved YOLOv8 network in the target domain, where or . The overall process of model transfer is: , where, represents the pre-trained weights of the model in the source domain, represents the training weights of the model in the target domain.
[0048] Step 32: Input the sample P display image into the initial object detection model to obtain a predicted object detection result.
[0049] Step 33: Construct a loss function based on the predicted object detection result and the sample object detection result corresponding to the sample P display image, and iteratively optimize the parameters of the initial object detection model according to the loss function until the loss function reaches the minimum value or the number of iterative optimization rounds reaches the maximum value, then stop the iterative optimization to obtain the object detection model.
[0050] Specifically, the overall network structure of the YOLOv8 network consists of an input layer (Input), a backbone network (Backbone), and a head network (Head). In the input layer, YOLOv8 turns off the Mosaic augmentation operation in the last 10 iteration cycles of the data augmentation part, thus effectively improving the accuracy of the model. In the backbone network part, the YOLOv8 network uses Conv layers, specifically including 2D convolution (Conv2d), 2D batch normalization (BatchNorm2d), and the activation function SiLU. The last layer of the backbone network is the SPPF layer, which consists of two Conv layers before and after and three MaxPooling layers connected in series in the middle, and then feature fusion is performed through a connection layer. The input feature map first passes through a Conv layer, then sequentially through three max pooling operations, and finally through another Conv layer. In the connection layer, four feature maps are connected, namely Conv, Conv and 1 max pooling, Conv and 2 max poolings, and Conv and 3 max poolings, finally achieving feature fusion at different scales. The YOLOv8 network replaces all C3 modules in the YOLOv5 network with C2f modules to obtain richer gradient flow information. In the head part, the original Anchor-based method in the YOLOv5 network is replaced by the Anchor-Free method, thus simplifying the object detection process.
[0051] Performing small target detection in the context of a moving platform radar is more challenging. In this scenario, the relative position between the radar and the target is constantly changing, increasing the complexity of the data and the diversity of the interference background. In addition, the noise and jitter generated by the moving platform further increase the difficulty of small target detection. The resolution of moving targets in the background is low, and their features are easily masked by the complex background and moving noise, thus increasing the possibility of missed detection. The backbone network of the YOLOv8 network increases the receptive field and reduces the computational amount through downsampling in the stride convolution module, which can effectively reduce redundant information in a fixed scenario. However, in the context of a moving platform radar, this operation may inadvertently lose fine-grained information, thus affecting the feature extraction of moving targets and reducing the detection accuracy. Therefore, the detection effect of the YOLOv8 network on moving targets is limited and not suitable for direct application in the moving platform radar scenario. For this reason, an improved YOLOv8 network is proposed to enhance the detection performance of moving targets. By optimizing the stride convolution and pooling layers in the YOLOv8 network, the SPD-Conv layer can complete downsampling while maintaining feature integrity, thus retaining fine-grained information. In addition, to reduce the interference of the complex background, Figure 3In the improved YOLOv8 network, SPD-Conv modules are added at the 3rd, 6th, 9th, and 12th layers of the backbone network, and the 22nd and 28th layers of the head network, replacing the original downsampling modules. This design enhances the learning ability of the moving target features, enabling the model to better adapt to the characteristics of the radar background on the moving platform.
[0052] When the YOLOv8 network inputs a radar image (i.e., a PPI image), it will first perform preliminary feature extraction. Assume that the input radar image contains targets of different scales, and these targets may have different sizes due to the movement of the platform (e.g., a small boat at close range and a large ship at long range). Due to the characteristics of the radar image, there may be a phenomenon of target-background fusion, especially for targets at long range, which may be difficult to distinguish. During the feature extraction process of the YOLOv8 network, the SPD-Conv layer is introduced. This layer is added after the backbone network of the YOLOv8 network and processes the feature map using multi-scale pooling and transposed convolution operations. Multi-scale pooling: Through the pooling operation, the SPD-Conv layer can extract features at different scales, enhancing the sensitivity to targets of different sizes. This is particularly important for the radar images of the moving platform because the size of the target changes over time (e.g., due to different movements and distances), which requires the detection network to be able to adapt to multiple scales. Transposed convolution: The SPD-Conv layer restores the spatial resolution of the feature map through transposed convolution upsampling, thereby enhancing the target features in the low-resolution region. This process helps the YOLOv8 network improve the spatial accuracy when processing long-range targets or small targets, making these targets easier to identify. The SPD-Conv layer performs multi-scale fusion on the feature map of the YOLOv8 network, enhancing the model's detection ability for small and large targets. In radar images, especially in complex backgrounds (such as the sea surface or urban areas), the scale and shape of the targets often vary highly, and the SPD-Conv layer can effectively improve the recognition of targets at multiple scales. By fusing features from different scales, the SPD-Conv layer helps the YOLOv8 network capture more detailed information and improves the detection ability for dynamic targets (such as moving ships).
[0053] A spatial-to-depth layer and a non-strided convolutional layer constitute the spatial-to-depth layer and non-strided convolution (SPD-Conv) layer. In the spatial-to-depth layer and non-strided convolutional layer, the original feature map (with size ) is divided according to the scale factor scale to form two feature submaps with size scale and located at , with size , and realizes the original feature map Downsampling by a scale factor. Next, the feature sub - graphs are concatenated along the channel dimension to obtain the intermediate - layer feature map , which preserves each bit of data in the channel dimension. The following formula describes the calculation process.
[0054] .
[0055] Among them, represents the feature sub - graph located at index , where ; represents the operation of extracting a sub - graph from the original feature map; S represents the stride used for index segmentation in the original feature map; is the scale factor, which determines the downsampling multiple of the space - to - depth operation.
[0056] Figure 4 Taking = 2 as an example, the original feature map is divided into four feature sub - graphs , and the size of all sub - graphs is , and 2 - fold downsampling of X is achieved. Then, these feature sub - graphs are concatenated to obtain the intermediate - layer feature map . At this time, the length and width of the original feature map X are reduced to half of the original, while the channel dimension increases to four times the original.
[0057] Specifically, the loss function includes: CIoU loss function and BCEWithLogits loss function. The loss function of common object detectors consists of coordinate loss, object confidence loss, and object classification loss. This application only focuses on, so the object confidence loss and coordinate loss constitute two parts of the loss function. The object confidence loss uses BCEWithLogitsLoss, which is a variant of the binary cross - entropy loss (BCELoss), combining BCELoss and the sigmoid function, and is numerically more stable than using BCELoss and sigmoid separately. The coordinate loss is based on CIoU loss, which considers the distance and aspect ratio between the center points of the bounding boxes, improving the ability to identify occlusion interference. Aiming at the problem that the traditional IoU metric fails, this application designs a metric based on the extended intersection - over - union (EIoU) region, which constructs spatio - temporal similarity between the initial non - overlapping detection region and the trajectory. It expands the matching space between the two without changing the original center point of the position, azimuth angle, scale, and shape.
[0058] The calculation formula of the CIoU loss function is as follows.
[0059] .
[0060] 。
[0061] 。
[0062] 。
[0063] Among them, among them, is the value of the CIoU loss function; is value; is the Euclidean distance; is the center point of the bounding box in the predicted object detection result; is the center point of the bounding box in the sample object detection result; is the diagonal distance of the minimum circumscribed rectangle between the bounding box in the predicted object detection result and the bounding box in the sample object detection result; is the weight parameter; is the parameter for measuring the aspect ratio consistency; D is the predicted bounding box after expanding the predicted object detection result; C is the size bounding box after expanding the predicted object detection result; is the width of the label bounding box after expanding the sample object detection result; is the height of the label bounding box after expanding the sample object detection result; is the width of the predicted bounding box after expanding the predicted object detection result; is the height of the predicted bounding box after expanding the predicted object detection result.
[0064] The calculation formula of the BCEWithLogits loss function is as follows.
[0065] 。
[0066] Among them, is the value of the BCEWithLogits loss function; is a fixed parameter; is the confidence in the sample object detection result; , is the class label vector of the object detection result.
[0067] Specifically, the Adam gradient descent method is used during iterative optimization training, given the parameters , the gradient , the momentum parameter , the learning rate parameter and the constant . First, perform exponential weighted moving average to estimate the first moment and the second moment .
[0068] 。
[0069] 。
[0070] Among them, is the number of the current iteration; is the first moment of the th iteration; is the first moment of the th iteration.
[0071] Then, update the parameters according to the corrected first moment and second moment.
[0072] 。
[0073] 。
[0074] 。
[0075] Among them, is the updated first moment of the th iteration; is the updated second moment of the th iteration; is the parameter of the th iteration; is the learning rate; is a constant.
[0076] Step 4: Based on the object detection results corresponding to each frame of P-display image, use the ByteTrack algorithm to calculate the trajectory of the object and complete object tracking.
[0077] Specifically, the object detection results include: bounding box, confidence, and class label; Step 4 includes the following steps.
[0078] Step 41: Calculate the coordinates of each object in two adjacent frames of P-display images.
[0079] Step 42: Determine the type of the bounding box according to the confidence of each object in the object detection results corresponding to each frame of P-display image; the types include high-confidence detection boxes and low-confidence detection boxes.
[0080] Specifically, Step 42 includes the following steps.
[0081] Step 421: Judge whether the confidence of each object in the object detection results corresponding to each frame of P-display image is greater than or equal to a preset high-frame threshold.
[0082] Step 422: If the confidence is greater than or equal to the preset high-frame threshold, determine that the bounding box is a high-confidence detection box.
[0083] Step 423: If the confidence level is less than the preset high-frame threshold, then check whether the confidence level is greater than or equal to the low-frame threshold.
[0084] Step 424: If the confidence level is greater than or equal to the low-frame threshold, determine the bounding box as a low-confidence detection box.
[0085] Step 425: If the confidence level is less than the low-frame threshold, the target in the bounding box.
[0086] Step 43: For different types of bounding boxes, use the Hungarian algorithm to match the targets in adjacent two-frame P-display images, and obtain the coordinates of the successfully matched targets and the coordinates of the unsuccessfully matched targets.
[0087] Step 44: Use Kalman filtering to predict the coordinates of each target in each frame of P-display image to obtain the predicted coordinates.
[0088] Specifically, the ByteTrack algorithm uses Kalman filtering to perform state estimation on each target. Assume the state vector of the target is as follows.
[0089] 。
[0090] In the formula, is the position (center point coordinates) of the target in the th frame; is the scale (area) of the target; is the aspect ratio of the target; is the velocity component of the target.
[0091] The Kalman filter prediction formula is as follows.
[0092] 。
[0093] 。
[0094] Among them, is the state prediction vector at the k th moment; is the state transition matrix; is the state prediction error covariance matrix at the k th moment; is the process noise covariance matrix.
[0095] The calculation formula for Kalman filter update is as follows.
[0096] 。
[0097] 。
[0098] 。
[0099] Among them, is the Kalman gain; is the observation matrix; is the observation noise covariance matrix; is the k state estimation vector at time is the k actual observation vector at time is the k state estimation covariance matrix at time
[0100] Step 45: Based on the predicted coordinates and the coordinates of the targets that have not been successfully matched, use the Hungarian algorithm to re-match the targets in two adjacent P-display images, and obtain the coordinates of the targets that have been successfully re-matched.
[0101] Specifically, when using the Hungarian algorithm to re-match the targets in two adjacent P-display images, if the number of matching times reaches the matching times threshold, the targets that have not been successfully matched are discarded.
[0102] Step 46: Determine the target trajectory based on the coordinates of the successfully matched targets and the coordinates of the targets that have been successfully re-matched, and achieve target tracking.
[0103] Specifically, in the context of a moving platform radar, most tracking algorithms may cause the true targets under occlusion or noise conditions to lose their tracking trajectories when directly discarding low-confidence detection boxes. To solve this problem, ByteTrack adopts a data association method, establishes preliminary trajectories through high-score detection boxes, and performs secondary matching on low-confidence detection boxes, thereby effectively mining occluded or low-score moving targets and maintaining the coherence of the trajectories. The specific steps are as follows.
[0104] As Figure 5 shown, first, set a high-confidence threshold T_high and a low-confidence threshold T_low according to the detection results of the improved YOLOv8. If the confidence of the target detection result exceeds the high-confidence threshold, it is judged as a high-confidence detection box; if it is between the low-confidence threshold and the high-confidence threshold, it is judged as a low-confidence detection box; when it is less than the low-confidence threshold T_low, it is directly discarded; at the same time, Kalman filtering will predict each tracking target to predict the next position (the Kalman filtering prediction result of the current frame is matched with the next frame of the current target by the Hungarian algorithm. If the match is successful, the Kalman filtering parameters are updated. The purpose of this is to ensure that when the target is occluded and the next frame position cannot be obtained, Kalman filtering can accurately predict).
[0105] Then, the high-confidence detection boxes and low-confidence detection boxes are respectively subjected to the first and second target matching (Hungarian matching (the high-confidence detection boxes are associated with the existing tracking targets, and the low-confidence detection boxes are associated with the unmatched tracking targets)). When the matching is successful, a tracking trajectory is generated. If the tracking fails, it is saved in the unmatched trajectory set, and the Hungarian matching of the target boxes is performed next time. When the number of failures reaches n, this detection box is deleted.
[0106] Specifically, the speed is calculated by the ratio of the displacement of the target center point to the time interval and smoothed using a moving average.
[0107] 。
[0108] Among them, , is the displacement in the X-axis direction of adjacent frames; is the coordinate in the X-axis direction of the t-th frame; is the coordinate in the X-axis direction of the (t - 1)-th frame; is the displacement in the Y-axis direction of adjacent frames; is the coordinate in the Y-axis direction of the t-th frame; is the coordinate in the Y-axis direction of the (t - 1)-th frame; s is the scale factor (assuming converting pixel distance to actual speed, which needs to be calibrated according to the scene), is the time interval; is the frame rate; the heading angle is calculated by the arctangent angle of the displacement vector and the coordinate system direction is adjusted.
[0109] 。
[0110] Since the y-axis in the image coordinate system is positive downward, −Δy is taken to ensure that the heading is 0 degrees at the top of the image (negative y-axis direction) and increases clockwise.
[0111] The beneficial effects of an integrated intelligent processing method for radar target detection and tracking on a moving platform proposed in this application are mainly manifested in the following three aspects.
[0112] (1) For variable environments and moving targets, the present invention has effective detection and tracking performance. The YOLOv8 network is combined with the spatial depth conversion convolution (SPD-Conv) layer to enhance the detection ability of moving targets under a moving platform radar. The improved YOLOv8 network is significantly superior to the original YOLOv8 network in detection performance. This application further combines the improved YOLOv8 network with the Bytetrack algorithm and optimizes the loss function, thereby improving the tracking performance.
[0113] (2) In this application, the parameters of the pre-trained model using the source domain dataset are used to initialize the improved YOLOv8 instead of random initialization, thereby improving the starting point of model training. Then, through global fine-tuning on the target domain (i.e., the sample P display image), the model can be adapted to new task requirements. The process of transfer learning involves transferring knowledge between the source domain and the target domain to enhance the generalization ability of the model. This method not only improves the accuracy and robustness of detection, but also accelerates the training process, solves the problem of data scarcity, effectively copes with the complexity of the maritime environment, and ensures the effect of real-time target tracking. Through transfer training, the target detection model of this application can shorten the training time of the model and has better generalization ability.
[0114] (3) In a dynamic environment, there may be situations where multiple targets are close in space and the detection confidence is low. At this time, the ByteTrack algorithm can match the targets in different frames through an efficient data association algorithm to maintain the accuracy of tracking. The adaptive threshold adjustment and background suppression of the improved YOLOv8 network can reduce detection errors and improve the target recognition ability of the ByteTrack algorithm. Combining the detection results output by the improved YOLOv8 network with the efficient data association of the ByteTrack algorithm can perform high-precision target tracking in real time in complex environments (such as multiple targets, occlusion, etc.). Through the precise detection framework provided by YOLOv8, ByteTrack can predict the target trajectory more accurately, avoiding the influence of traditional tracking algorithms being easily affected by drift and false matching. In this application, through integrated tracking and detection processing, the target detection results of the improved YOLOv8 network are combined with the Bytetrack algorithm, and real-time high-precision target detection and tracking can be achieved.
[0115] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0116] Specific examples are used in this article to elaborate on the principle and implementation method of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. An integrated intelligent processing method for radar target detection and tracking on a moving platform, characterized in that: The integrated intelligent processing method for moving platform radar target detection and tracking includes: Obtain the echo signals of targets at multiple times of the moving platform radar; Preprocess the echo signals of the target at multiple times to generate multiple frames of P display images; Input multiple frames of P display images into the target detection model respectively to obtain the target detection result corresponding to each frame of P display image; the target detection result is the detection result of multiple targets in the P display image; the target detection model is obtained by performing migration training on the improved YOLOv8 network through sample P display images; the improved YOLOv8 network adds multiple SPD-Conv layers to the YOLOv8 network; the loss function includes: CIoU loss function and BCEWithLogits loss function; Based on the target detection results corresponding to each frame of P display image, the ByteTrack algorithm is used to calculate the target trajectory and complete target tracking, including: The target detection result includes: a bounding box, a confidence level, and a category label; Calculate the coordinates of each target in two adjacent P display images; Determine the type of the bounding box according to the confidence level of each target in the target detection result corresponding to each frame of the P display image; the type includes a high confidence detection box and a low confidence detection box; For different types of bounding boxes, the Hungarian algorithm is used to match the objects in two adjacent P-display images, and the coordinates of the successfully matched objects and the unmatched objects are obtained; The Kalman filter is used to predict the coordinates of each target in each frame of P display image to obtain the predicted coordinates; Based on the predicted coordinates and the coordinates of the target that was not successfully matched, the Hungarian algorithm is used to rematch the targets in two adjacent frames of P display images, and the coordinates of the target that was successfully matched again are obtained; The target trajectory is determined based on the coordinates of the successfully matched target and the coordinates of the target that is successfully matched again, thereby achieving target tracking.
2. The integrated intelligent processing method for moving platform radar target detection and tracking according to claim 1 is characterized in that: The training process of the target detection model specifically includes: Migrating the network parameters in the pre-trained model to the improved YOLOv8 network to obtain an initial target detection model; the structure of the pre-trained model is the same as that of the target detection model; Inputting the sample P display image into the initial target detection model to obtain a predicted target detection result; A loss function is constructed based on the predicted target detection result and the sample target detection result corresponding to the sample P display image, and the parameters of the initial target detection model are iteratively optimized according to the loss function until the loss function reaches a minimum value or the iterative optimization round reaches a maximum value, and the iterative optimization is stopped to obtain the target detection model.
3. The integrated intelligent processing method for moving platform radar target detection and tracking according to claim 1 is characterized in that: The calculation formula of the CIoU loss function is: ; in, is the CIoU loss function value; for value; is the Euclidean distance; To predict the center point of the bounding box in the target detection result; is the center point of the bounding box in the sample target detection result; The diagonal distance between the bounding box in the predicted target detection result and the bounding box in the sample target detection result; is the weight parameter; is a parameter to measure the consistency of aspect ratio.
4. The integrated intelligent processing method for radar target detection and tracking on a moving platform according to claim 1 is characterized in that: The calculation formula of the BCEWithLogits loss function is: ; in, is the BCEWithLogits loss function value; is a fixed parameter; It is the confidence level in the sample target detection result; , is the category label vector of the target detection result.
5. The integrated intelligent processing method for moving platform radar target detection and tracking according to claim 1 is characterized in that: The improved YOLOv8 network includes: an input layer, a backbone network and a head network; the backbone network and the head network are added with multiple SPD-Conv layers.
6. The integrated intelligent processing method for radar target detection and tracking on a moving platform according to claim 1 is characterized in that: The SPD-Conv layer includes: a space-to-depth layer and a non-strided convolution layer.
7. The integrated intelligent processing method for radar target detection and tracking on a moving platform according to claim 1 is characterized in that: According to the confidence level of each target in the target detection result corresponding to each frame of the P display image, the type of the bounding box is determined, specifically including: Determine whether the confidence of each target in the target detection result corresponding to each frame of the P display image is greater than or equal to a preset high frame threshold; If the confidence is greater than or equal to a preset high frame threshold, determining the bounding box as a high confidence detection box; If the confidence is less than the preset high frame threshold, then whether the confidence is greater than or equal to the low frame threshold; If the confidence is greater than or equal to the low frame threshold, the bounding box is determined to be a low confidence detection box; If the confidence is less than the low frame threshold, the object is in the bounding box.
8. The integrated intelligent processing method for radar target detection and tracking on a moving platform according to claim 1 is characterized in that: When the Hungarian algorithm is used to match the targets in two adjacent P-display images again, if the number of matches reaches the matching number threshold, the unmatched targets are discarded.
Citation Information
Patent Citations
SPD-YOLOv8-based remote sensing image ship detection method
CN119229309A