Motion platform radar target detection and tracking integrated intelligent processing method
By adding the SPD-Conv layer to the YOLOv8 network and combining the ByteTrack algorithm, the problem of inaccuracy in traditional radar target detection methods in complex environments is solved, and high-precision and robust target detection and tracking are achieved.
Patent Information
- Application Number
- CN202510472424.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-16
AI Technical Summary
Traditional radar target detection methods lack robustness in the face of diverse changes and complex environments, resulting in inaccurate and unstable target detection tracking.
A integrated intelligent processing method for radar target detection and tracking of motion platform is adopted. By adding multiple SPD-Conv layers to the YOLOv8 network, multi-scale features are extracted, and target trajectory is calculated in combination with the ByteTrack algorithm to achieve efficient target tracking.
It improves the accuracy and robustness of object detection and tracking, and can accurately identify and track moving targets in complex environments, reducing the situation of missed target detection and loss.
Smart Images

Figure CN120013994A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of radar target detection, and in particular to an integrated intelligent processing method for radar target detection and tracking on a moving platform. Background Art
[0002] Target monitoring is the core function of radar, and the detection and tracking of dynamic targets plays a vital role in the monitoring field. It has become the focus of academic and engineering research at home and abroad. With the continuous improvement of the intelligent perception capability of mobile platform radar, it can achieve multiple tasks such as obstacle avoidance, navigation, target detection and tracking by combining visual information, and is widely used in fields such as autonomous driving, environmental perception and intelligent transportation. This puts higher requirements on the moving platform radar, requiring it to have not only a high level of cognitive and execution capabilities, but also good generalization capabilities.
[0003] Traditional target detection methods use artificially designed scale-invariant features or directional gradient histogram (HOG) features to identify sliding windows. However, these features lack sufficient robustness in the face of changes in input diversity. In traditional target tracking methods, the correlation filtering algorithm and its improved version improve tracking efficiency by calculating the correlation of candidate region feature maps, but only use the relevant feature information between adjacent frames, which makes it easy to lose the target when the target is deformed or occluded. Therefore, traditional methods are difficult to meet the needs of moving platform radars for high-performance detection and tracking, and are prone to missed detection or lost tracking of targets, and these models are usually only applicable to detection and tracking under specific conditions. However, due to the complex and changeable target environment and diverse motion characteristics faced by moving platform radars, target detection and tracking are further inaccurate and unstable. Summary of the invention
[0004] The purpose of this application is to provide an integrated intelligent processing method for moving platform radar target detection and tracking, which can improve the target detection and tracking accuracy.
[0005] To achieve the above objectives, this application provides the following solutions.
[0006] The present application provides an integrated intelligent processing method for radar target detection and tracking on a moving platform, comprising the following steps.
[0007] Acquire the echo signals of targets at multiple times of the moving platform radar.
[0008] The echo signals of the target at multiple moments are preprocessed to generate multiple frames of P display images.
[0009] Multiple frames of P-display images are respectively input into the target detection model to obtain the target detection results corresponding to each frame of the P-display image; the target detection results are the detection results of multiple targets in the P-display image; the target detection model is obtained by performing migration training on the improved YOLOv8 network through sample P-display images; the improved YOLOv8 network is to add multiple SPD-Conv layers to the YOLOv8 network.
[0010] Based on the target detection results corresponding to each frame of P display image, the ByteTrack algorithm is used to calculate the target trajectory and complete target tracking.
[0011] Optionally, the training process of the target detection model specifically includes: migrating the network parameters in the pre-trained model to the improved YOLOv8 network to obtain an initial target detection model; the structure of the pre-trained model is the same as that of the target detection model; inputting the sample P display image into the initial target detection model to obtain a predicted target detection result; constructing a loss function based on the predicted target detection result and the sample target detection result corresponding to the sample P display image, and iteratively optimizing the parameters of the initial target detection model according to the loss function until the loss function reaches a minimum value or the iterative optimization round reaches a maximum value, stopping the iterative optimization, and obtaining the target detection model.
[0012] Optionally, the loss function includes: a CIoU loss function and a BCEWithLogits loss function.
[0013] Optionally, a calculation formula for the CIoU loss function is as follows.
[0014] .
[0015] in, is the CIoU loss function value; for value; is the Euclidean distance; To predict the center point of the bounding box in the target detection result; is the center point of the bounding box in the sample target detection result; The diagonal distance between the bounding box in the predicted target detection result and the bounding box in the sample target detection result; is the weight parameter; is a parameter to measure the consistency of aspect ratio.
[0016] Optionally, the calculation formula of the BCEWithLogits loss function is as follows.
[0017] .
[0018] in, is the BCEWithLogits loss function value; is a fixed parameter; It is the confidence level in the sample target detection result; , is the category label vector of the target detection result.
[0019] Optionally, the improved YOLOv8 network includes: an input layer, a backbone network and a head network; the backbone network and the head network are added with multiple SPD-Conv layers.
[0020] Optionally, the SPD-Conv layer includes: a space-to-depth layer and a non-strided convolution layer.
[0021] Optionally, the target detection result includes: a bounding box, a confidence level and a category label; based on the target detection result corresponding to each frame of the P display image, the ByteTrack algorithm is used to calculate the trajectory of the target to complete target tracking, specifically including: calculating the coordinates of each target in two adjacent frames of the P display image; determining the type of the bounding box according to the size of the confidence level of each target in the target detection result corresponding to each frame of the P display image; the types include high-confidence detection boxes and low-confidence detection boxes; for different types of bounding boxes, the Hungarian algorithm is used to match the targets in two adjacent frames of the P display image, and the coordinates of the successfully matched targets and the coordinates of the unmatched targets are obtained; the Kalman filter is used to predict the coordinates of each target in each frame of the P display image to obtain the predicted coordinates; based on the predicted coordinates and the coordinates of the unmatched targets, the Hungarian algorithm is used to match the targets in two adjacent frames of the P display image again, and the coordinates of the targets that are matched again are obtained; the target trajectory is determined based on the coordinates of the successfully matched targets and the coordinates of the targets that are matched again to achieve target tracking.
[0022] Optionally, the type of the bounding box is determined according to the size of the confidence of each target in the target detection results corresponding to each frame of the P-displayed image, specifically including: judging whether the confidence of each target in the target detection results corresponding to each frame of the P-displayed image is greater than or equal to a preset high frame threshold; if the confidence is greater than or equal to the preset high frame threshold, determining that the bounding box is a high-confidence detection box; if the confidence is less than the preset high frame threshold, whether the confidence is greater than or equal to the low frame threshold; if the confidence is greater than or equal to the low frame threshold, determining that the bounding box is a low-confidence detection box; if the confidence is less than the low frame threshold, the target in the bounding box.
[0023] Optionally, when the Hungarian algorithm is used to match the targets in two adjacent P-display image frames again, if the number of matches reaches a matching number threshold, the targets that have not been successfully matched are discarded.
[0024] According to the specific embodiments provided in this application, this application has the following technical effects.
[0025] The present application provides an integrated intelligent processing method for radar target detection and tracking on a moving platform. By adding multiple SPD-Conv layers on the basis of the YOLOv8 network, features can be extracted at different scales, more detailed information can be captured, and the detection accuracy of the target can be improved. The target detection result corresponding to each frame of the P display image is calculated by the ByteTrack algorithm to obtain the trajectory of the target. The targets in different frames are matched through an efficient data association algorithm to ensure the accuracy of tracking. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0027] Figure 1 A flowchart of an integrated intelligent processing method for moving platform radar target detection and tracking provided in one embodiment of the present application.
[0028] Figure 2 A schematic diagram of a specific process of an integrated intelligent processing method for radar target detection and tracking on a moving platform provided in one embodiment of the present application.
[0029] Figure 3 A schematic diagram of an improved Yolov8 network provided in accordance with an embodiment of the present application.
[0030] Figure 4 A schematic diagram of the SPD-Conv layer provided in one embodiment of the present application.
[0031] Figure 5 A schematic diagram of the ByteTrack algorithm flow provided in one embodiment of the present application. DETAILED DESCRIPTION
[0032] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0033] With the development of deep learning, deep learning has gradually become a highly feasible method in the field of radar detection due to its powerful feature learning ability, generalization ability and interpretability. At present, target detection algorithms are mainly divided into one-stage algorithms and two-stage algorithms. The one-stage algorithm directly extracts features from the image and uses a classifier or regressor to predict the location and category of the object, such as YOLO and SSD. These algorithms have the characteristics of fast detection speed and low computational complexity, but the detection accuracy is relatively low. The two-stage algorithm divides target detection into two subtasks: target proposal and target classification, such as Faster R-CNN, which has higher detection accuracy, but slower speed and higher computational complexity.
[0034] In terms of target tracking, typical algorithms include SORT, DeepSORT and ByteTrack. The SORT framework consists of detectors, motion estimation and data association. It uses Kalman filters to predict target motion and matches detection boxes and prediction boxes through the Hungarian algorithm in data association. DeepSORT adds Mahalanobis distance and cosine distance as data association methods for cascade matching on the basis of SORT, which improves the long-term association capability. ByteTrack retains low-score detection boxes and performs secondary matching to reduce missed detections and improve tracking consistency. However, SORT only relies on the motion model and IoU matching of Kalman filtering. When there is occlusion or motion mutation, it is easy to cause frequent ID switching due to lack of appearance information, and directly discard low-score detection boxes to cause missed detections; although DeepSORT introduces appearance features to reduce ID switching, it relies on high-dimensional ReID models to calculate similarity, which significantly increases the computational cost, and appearance drift (such as illumination changes or target deformation) will lead to mismatching, while low-score real targets may still be ignored. Neither of them systematically handles low-confidence detections, and ByteTrack mitigates missed detections through a two-stage association strategy (first high-score detection and then low-score detection), but SORT and DeepSORT have weak tracking robustness in complex occlusion or target blur scenes due to design limitations. Although research based on motion target detection and tracking has made some progress, there are still many challenges in the application of moving platform radars, such as occlusion and rapid target movement. In particular, when moving platform radars are more likely to lose targets than fixed radars, traditional tracking algorithms are usually unable to recover the target once it loses the target.
[0035] Therefore, improving the accuracy of target detection and tracking is an urgent problem to be solved.
[0036] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0037] In an exemplary embodiment, Figure 1 and Figure 2As shown, a moving platform radar target detection and tracking integrated intelligent processing method is provided, which is executed by a computer device, and can be specifically executed by a computer device such as a terminal or a server alone, or by a terminal and a server together. In an embodiment of the present application, the method is applied to a server as an example for explanation, including the following steps 1 to 8.
[0038] Step 1: Obtain the target echo signals of the moving platform radar at multiple times. Figure 2 The target echo signal is received through the moving platform radar.
[0039] Step 2: Preprocess the echo signals of the target at multiple moments to generate multiple frames of P display images.
[0040] Specifically, the target echo signal of the moving platform radar is preprocessed to generate a P display image, and the P display image (Plan Position Indicator, PPI) is drawn. The P display image is a two-dimensional image that shows the distance and direction of the radar relative to its location.
[0041] .
[0042] .
[0043] in, is the angle index, representing the azimuth number in the radar scanning data; is the x-axis coordinate of the P display image; is the distance to the grid point; is the angle data; is the y-axis coordinate of the P display image.
[0044] Step 3: Input multiple frames of P-display images into the target detection model respectively to obtain the target detection results corresponding to each frame of P-display image; the target detection results are the detection results of multiple targets in the P-display image; the target detection model is obtained by performing migration training on the improved YOLOv8 network through sample P-display images; the improved YOLOv8 network is to add multiple SPD-Conv layers to the YOLOv8 network. Among them, the improved YOLOv8 network includes: input layer, backbone network and head network; the backbone network and the head network add multiple SPD-Conv layers. The SPD-Conv layer includes: space to depth layer and non-strided convolution layer.
[0045] Specifically, the training process of the target detection model includes the following steps.
[0046] Step 31: Migrate the network parameters in the pre-trained model to the improved YOLOv8 network to obtain an initial target detection model. The structure of the pre-trained model is the same as that of the target detection model.
[0047] When migrating the model, the network parameters in the pre-trained model are used to initialize the weights of the improved YOLOv8 network, replacing the original random initialization operation, and the initial target detection model is globally fine-tuned through training again. Figure 2 As shown, the source domain DATA1 is obtained through preprocessing and is recorded as and the target domain DATA2, denoted as , the transfer learning studied in this application focuses on the active domain and a target domain (i.e., sample P display image), where the source domain Use the expression, , Respectively represent the source domain data samples and the corresponding motion target category labels, Represents the total number of source domain data samples, and the target domain uses express, , Respectively represent the target domain data samples and the corresponding moving target category labels, Represents the total number of target domain data samples. Given the source domain and learning tasks , target domain and learning tasks , the purpose of transfer learning is to obtain the source domain and learning tasks The knowledge in helps improve the learning of the improved YOLOv8 network in the target domain, where or The overall process of model migration is: ,in, represents the model's pre-trained weights in the source domain, Represents the model target domain training weights.
[0048] Step 32: Input the sample P display image into the initial target detection model to obtain a predicted target detection result.
[0049] Step 33: Construct a loss function based on the predicted target detection result and the sample target detection result corresponding to the sample P display image, and iteratively optimize the parameters of the initial target detection model according to the loss function until the loss function reaches a minimum value or the iterative optimization round reaches a maximum value, stop the iterative optimization, and obtain the target detection model.
[0050] Specifically, the overall network structure of the YOLOv8 network consists of an input layer (Input), a backbone network (Backbone), and a head network (Head). In the input layer, YOLOv8 turns off the Mosaic enhancement operation in the last 10 iterations of the data enhancement part, thereby effectively improving the accuracy of the model. In the backbone network part, the YOLOv8 network uses a Conv layer, which specifically includes 2D convolution (Conv2d), 2D batch normalization (BatchNorm2d), and activation function SiLU. The last layer of the backbone network is the SPPF layer, which consists of two Conv layers and three MaxPooling layers in series in the middle, and then a connection layer is used for feature fusion. The input feature map first passes through a Conv layer, then passes through three maximum pooling operations in sequence, and finally passes through another Conv layer. In the connection layer, four feature maps are connected, namely Conv, Conv and 1 maximum pooling, Conv and 2 maximum pooling, and Conv and 3 maximum pooling, and finally feature fusion at different scales is achieved. The YOLOv8 network replaces all C3 modules in the YOLOv5 network with C2f modules to obtain richer gradient flow information. In the head part, the original Anchor-based method of the YOLOv5 network is replaced by the Anchor-Free method, which simplifies the object detection process.
[0051] Small target detection in the context of a moving platform radar is more challenging. In this scenario, the relative position of the radar and the target is constantly changing, which increases the complexity of the data and the diversity of the interference background. In addition, the noise and jitter generated by the moving platform further increase the difficulty of small target detection. The resolution of moving targets in the background is low, and their features are easily masked by complex background and motion noise, which increases the possibility of missed detection. The backbone network of the YOLOv8 network increases the receptive field and reduces the amount of calculation by downsampling in the strided convolution module, which can effectively reduce redundant information in fixed scenarios. However, in the context of a moving platform radar, this operation may inadvertently lose fine-grained information, thereby affecting the feature extraction of moving targets and reducing detection accuracy. Therefore, the YOLOv8 network has limited detection effect on moving targets and is not suitable for direct application in moving platform radar scenarios. To this end, an improved YOLOv8 network is proposed to enhance the detection performance of moving targets. By optimizing the strided convolution and pooling layers in the YOLOv8 network, the SPD-Conv layer is able to complete downsampling while maintaining feature integrity, thereby retaining fine-grained information. In addition, to reduce the interference of complex backgrounds, Figure 3The improved YOLOv8 network adds SPD-Conv modules to the 3rd, 6th, 9th and 12th layers of the backbone network and the 22nd and 28th layers of the head network, replacing the original downsampling modules. This design enhances the ability to learn the characteristics of moving targets, allowing the model to better adapt to the characteristics of the moving platform radar background.
[0052] When the YOLOv8 network inputs a radar image (i.e., a P-display image), it first performs preliminary feature extraction. Assume that the input radar image contains targets of different scales, which may appear in different sizes due to the motion of the platform (for example, a small boat at a close distance and a large ship at a long distance). Due to the characteristics of radar images, there may be a fusion of targets and backgrounds, especially targets at a long distance may be difficult to distinguish. In the feature extraction process of the YOLOv8 network, the SPD-Conv layer is introduced. This layer is added after the backbone network of the YOLOv8 network, and multi-scale pooling and deconvolution operations are used to process the feature map. Multi-scale pooling: Through pooling operations, the SPD-Conv layer can extract features at different scales and enhance sensitivity to targets of different sizes. This is particularly important for radar images of moving platforms, because the size of the target changes over time (for example, due to different motion and distance), which requires the detection network to adapt to multiple scales. Deconvolution: The SPD-Conv layer restores the spatial resolution of the feature map through deconvolution upsampling, thereby enhancing the target features in low-resolution areas. This process helps the YOLOv8 network improve spatial accuracy when dealing with distant or small targets, making these targets easier to identify. The SPD-Conv layer fuses the feature maps of the YOLOv8 network at multiple scales, enhancing the model's ability to detect small and large targets. In radar images, especially in complex backgrounds (such as sea surfaces or urban areas), the scale and shape of targets often vary greatly. The SPD-Conv layer can effectively improve the recognition of targets at multiple scales. By fusing features from different scales, the SPD-Conv layer helps the YOLOv8 network capture more detailed information and improves the ability to detect dynamic targets (such as moving ships).
[0053] A spatial-to-depth layer and a non-strided convolution layer constitute a spatial-to-depth layer and non-strided convolution (SPD-Conv) layer. In the spatial-to-depth layer and the non-strided convolution layer, the original feature map (Size is ) is divided according to the scale factor scale, forming two scales located at The feature subgraph of , and realized the original feature map Next, the feature sub-maps are connected along the channel dimension to obtain the intermediate layer feature map , which retains every bit of data in the channel dimension. The following formula describes the calculation process.
[0054] .
[0055] in, Indicates that it is at index The characteristic subgraph at ; Represents the operation of extracting a sub-map from the original feature map; S Stride represents the index segmentation in the original feature map. It is the scaling factor that determines the downsampling multiple of the space-to-depth operation.
[0056] Figure 4 by = 2 as an example, the original feature map Divided into four feature subgraphs , the size of all sub-graphs is , and achieves 2 times downsampling of X. Then, these feature subgraphs are connected to obtain the intermediate layer feature graph At this time, the length and width of the original feature map X are reduced to half of the original, while the channel dimension is increased to four times the original.
[0057] Specifically, the loss function includes: CIoU loss function and BCEWithLogits loss function. The loss function of common target detectors consists of coordinate loss, target confidence loss and target classification loss. This application is only for, so the target confidence loss and coordinate loss constitute two parts of the loss function. The target confidence loss uses BCEWithLogitsLoss, which is a variant of the binary cross entropy loss (BCELoss), combining BCELoss and sigmoid functions, and is numerically more stable than using BCELoss and sigmoid alone. The coordinate loss is based on the CIoU loss, taking into account the distance and aspect ratio between the center points of the bounding box, and improving the recognition ability of occlusion interference. In response to the problem of failure of the traditional IoU metric, this application designs a metric based on the extended intersection-over-union (EIoU) area, which constructs spatiotemporal similarity between the initial non-overlapping detection area and the trajectory. The matching space between the two is expanded without changing the center point, azimuth, scale, and shape of the original position.
[0058] The calculation formula of CIoU loss function is as follows.
[0059] .
[0060] .
[0061] .
[0062] .
[0063] Among them, among them, is the CIoU loss function value; for value; is the Euclidean distance; To predict the center point of the bounding box in the target detection result; is the center point of the bounding box in the sample target detection result; The diagonal distance between the bounding box in the predicted target detection result and the bounding box in the sample target detection result; is the weight parameter; is a parameter to measure the consistency of aspect ratio; D is the predicted box after the expansion of the predicted target detection result; C is the size box after the expansion of the predicted target detection result; The width of the label box after the sample target detection result is expanded; The height of the label box after the sample target detection result is expanded; The width of the prediction box after the target detection result is expanded; The height of the prediction box after the target detection result is expanded.
[0064] The calculation formula of BCEWithLogits loss function is as follows.
[0065] .
[0066] in, is the BCEWithLogits loss function value; is a fixed parameter; It is the confidence level in the sample target detection result; , is the category label vector of the target detection result.
[0067] Specifically, the Adam gradient descent method is used during iterative optimization training. Given the parameters ,gradient , momentum parameter , the learning rate parameter and constant First, an exponentially weighted moving average is performed to estimate the first moment of the gradient and the second moment .
[0068] .
[0069] .
[0070] in, is the number of current iterations; For the The first moment of the iteration; For the The first moment of the iteration.
[0071] Then, the parameters are updated according to the modified first-order and second-order moments.
[0072] .
[0073] .
[0074] .
[0075] in, For the The updated first-order moment of the iteration; For the The updated second-order moment of the iteration; For the The parameters of the iterations; is the learning rate; is a constant.
[0076] Step 4: Based on the target detection results corresponding to each frame of the P display image, the ByteTrack algorithm is used to calculate the target trajectory to complete target tracking.
[0077] Specifically, the target detection result includes: a bounding box, a confidence level, and a category label; step 4 includes the following steps.
[0078] Step 41: Calculate the coordinates of each target in two adjacent P display image frames.
[0079] Step 42: Determine the type of the bounding box according to the confidence level of each target in the target detection result corresponding to each frame of the P display image; the type includes a high-confidence detection box and a low-confidence detection box.
[0080] Specifically, step 42 includes the following steps.
[0081] Step 421: Determine whether the confidence level of each target in the target detection result corresponding to each frame of the P display image is greater than or equal to a preset high frame threshold.
[0082] Step 422: If the confidence is greater than or equal to the preset high frame threshold, the bounding box is determined to be a high confidence detection box.
[0083] Step 423: If the confidence level is less than the preset high frame threshold, then determine whether the confidence level is greater than or equal to the low frame threshold.
[0084] Step 424: If the confidence is greater than or equal to the low frame threshold, determine that the bounding box is a low confidence detection box.
[0085] Step 425: If the confidence score is less than the low frame threshold, then the target is in the bounding box.
[0086] Step 43: For different types of bounding boxes, the Hungarian algorithm is used to match the targets in two adjacent P-display image frames, and the coordinates of the successfully matched targets and the coordinates of the unmatched targets are obtained.
[0087] Step 44: Use Kalman filtering to predict the coordinates of each target in each frame of P display image to obtain predicted coordinates.
[0088] Specifically, the ByteTrack algorithm uses Kalman filtering to estimate the state of each target. Assume that the state vector of the target is as follows.
[0089] .
[0090] In the formula, For the goal The position of the frame (center point coordinates); is the size (area) of the target; is the aspect ratio of the target; is the velocity component of the target.
[0091] The Kalman filter prediction formula is as follows.
[0092] .
[0093] .
[0094] in, For the k The state prediction vector at the moment; is the state transfer matrix; For the k The state prediction error covariance matrix at time ; is the process noise covariance matrix.
[0095] The calculation formula for Kalman filter update is as follows.
[0096] .
[0097] .
[0098] .
[0099] in, is the Kalman gain; is the observation matrix; is the observation noise covariance matrix; For the k The state estimation vector at time t; For the k The actual observation vector at time instant; For the k The state estimation covariance matrix at time .
[0100] Step 45: Based on the predicted coordinates and the coordinates of the target that was not successfully matched, the Hungarian algorithm is used to rematch the targets in two adjacent P display image frames, and the coordinates of the target that was successfully matched again are obtained.
[0101] Specifically, when the Hungarian algorithm is used to match the targets in two adjacent P-display images again, if the number of matches reaches a matching number threshold, the targets that have not been successfully matched are discarded.
[0102] Step 46: Determine the target trajectory based on the coordinates of the successfully matched target and the coordinates of the target that is successfully matched again, so as to achieve target tracking.
[0103] Specifically, in the context of moving platform radar, most tracking algorithms may cause the real target under occlusion or noise conditions to lose its tracking trajectory when directly discarding low-confidence detection frames. To solve this problem, ByteTrack uses a data association method to establish a preliminary trajectory through high-scoring detection frames and perform secondary matching on low-confidence detection frames, thereby effectively mining occluded or low-scoring moving targets and maintaining trajectory continuity. The specific steps are as follows.
[0104] like Figure 5 As shown in the figure, first, the high confidence threshold T_high and the low confidence threshold T_low are set according to the detection results of the improved YOLOv8. If the confidence of the target detection result exceeds the high confidence threshold, it is judged as a high confidence detection frame; if it is between the low confidence threshold and the high confidence threshold, it is judged as a low confidence detection frame; when it is less than the low confidence threshold T_low, it is directly discarded; at the same time, the Kalman filter will predict each tracked target and predict the next position (the Kalman filter prediction result of the current frame is Hungarian matched with the next frame of the current target. If the match is successful, the Kalman filter parameters are updated. The purpose is to accurately predict when the target is blocked and the next frame position cannot be obtained).
[0105] Then the high-confidence detection frame and the low-confidence detection frame are matched with the target for the first time and the second time respectively (Hungarian matching (high-confidence detection frame is associated with the existing tracking target, low-confidence detection frame is associated with the unmatched tracking target)). When the match is successful, a tracking trajectory is generated. If the tracking fails, it is saved in the unmatched trajectory set and the next Hungarian matching target frame is performed. When the number of failures reaches n, the detection frame is deleted.
[0106] Specifically, the velocity is calculated by the ratio of the displacement of the target center point to the time interval, and a sliding average smoothing is applied.
[0107] .
[0108] in, , is the displacement in the X-axis direction of adjacent frames; is the coordinate of the X-axis direction of the tth frame; is the coordinate of the X-axis direction of the t-1th frame; is the displacement in the Y-axis direction of adjacent frames; is the coordinate of the Y axis direction of the tth frame; is the coordinate of the Y axis of the t-1th frame; s is the scale factor (assuming that the pixel distance is converted into actual speed, which needs to be calibrated according to the scene), is the time interval; is the frame rate; heading angle The inverse tangent angle of the displacement vector is calculated and the coordinate system direction is adjusted.
[0109] .
[0110] Since the y-axis in the image coordinate system is positive downward, −Δy is taken to ensure that the heading is 0 degrees at the top of the image (negative direction of the y-axis) and increases clockwise.
[0111] The beneficial effects of the integrated intelligent processing method for moving platform radar target detection and tracking proposed in this application are mainly manifested in the following three aspects.
[0112] (1) The present invention has effective detection and tracking performance for changing environments and moving targets. The YOLOv8 network is combined with the spatial depth conversion convolution (SPD-Conv) layer to enhance the detection capability of moving targets under the moving platform radar. The improved YOLOv8 network is significantly better than the original YOLOv8 network in detection performance. The present application further combines the improved YOLOv8 network with the Bytetrack algorithm and optimizes the loss function, thereby improving the tracking performance.
[0113] (2) This application uses the parameters of the pre-trained model of the source domain dataset to initialize the improved YOLOv8 instead of random initialization, thereby improving the starting point of model training. Then, by performing global fine-tuning on the target domain (i.e., sample P-display images), the model can adapt to new task requirements. The process of transfer learning involves transferring knowledge between the source domain and the target domain to enhance the generalization ability of the model. This method not only improves the accuracy and robustness of detection, but also accelerates the training process, solves the problem of data scarcity, effectively copes with the complexity of the marine environment, and ensures the effect of real-time target tracking. This application enables the target detection model to shorten the training time of the model through transfer training and has better generalization ability.
[0114] (3) In a dynamic environment, there may be multiple targets that are close in space and the detection confidence is low. In this case, the ByteTrack algorithm can match targets in different frames through an efficient data association algorithm to maintain tracking accuracy. The adaptive threshold adjustment and background suppression of the improved YOLOv8 network can reduce detection errors and enhance the ByteTrack algorithm's ability to identify targets. Combining the detection results output by the improved YOLOv8 network with the efficient data association of the ByteTrack algorithm, it is possible to perform high-precision target tracking in real time in complex environments (such as multiple targets, occlusions, etc.). Through the precise detection framework provided by YOLOv8, ByteTrack can more accurately predict the target trajectory and avoid the traditional tracking algorithm being susceptible to drift and mismatching. This application combines the target detection results of the improved YOLOv8 network with the Bytetrack algorithm through integrated tracking and detection processing, and can achieve real-time and high-precision target detection and tracking.
[0115] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0116] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. An integrated intelligent processing method for radar target detection and tracking on a moving platform, characterized in that: The integrated intelligent processing method for moving platform radar target detection and tracking includes: Obtain the echo signals of targets at multiple times of the moving platform radar; Preprocess the echo signals of the target at multiple times to generate multiple frames of P display images; Input multiple frames of P display images into the target detection model respectively to obtain the target detection result corresponding to each frame of P display image; the target detection result is the detection result of multiple targets in the P display image; the target detection model is obtained by performing migration training on the improved YOLOv8 network through the sample P display image; the improved YOLOv8 network is to add multiple SPD-Conv layers to the YOLOv8 network; Based on the target detection results corresponding to each frame of P display image, the ByteTrack algorithm is used to calculate the target trajectory and complete target tracking.
2. The integrated intelligent processing method for moving platform radar target detection and tracking according to claim 1 is characterized in that: The training process of the target detection model specifically includes: Migrating the network parameters in the pre-trained model to the improved YOLOv8 network to obtain an initial target detection model; the structure of the pre-trained model is the same as that of the target detection model; Inputting the sample P display image into the initial target detection model to obtain a predicted target detection result; A loss function is constructed based on the predicted target detection result and the sample target detection result corresponding to the sample P display image, and the parameters of the initial target detection model are iteratively optimized according to the loss function until the loss function reaches a minimum value or the iterative optimization round reaches a maximum value, and the iterative optimization is stopped to obtain the target detection model.
3. The integrated intelligent processing method for radar target detection and tracking on a moving platform according to claim 2 is characterized in that: The loss functions include: CIoU loss function and BCEWithLogits loss function.
4. The integrated intelligent processing method for moving platform radar target detection and tracking according to claim 3 is characterized in that: The calculation formula of the CIoU loss function is: ; in, is the CIoU loss function value; for value; is the Euclidean distance; To predict the center point of the bounding box in the target detection result; is the center point of the bounding box in the sample target detection result; The diagonal distance between the bounding box in the predicted target detection result and the bounding box in the sample target detection result; is the weight parameter; is a parameter to measure the consistency of aspect ratio.
5. The integrated intelligent processing method for radar target detection and tracking on a moving platform according to claim 3 is characterized in that: The calculation formula of the BCEWithLogits loss function is: ; in, is the BCEWithLogits loss function value; is a fixed parameter; It is the confidence level in the sample target detection result; , is the category label vector of the target detection result.
6. The integrated intelligent processing method for radar target detection and tracking on a moving platform according to claim 1 is characterized in that: The improved YOLOv8 network includes: an input layer, a backbone network and a head network; the backbone network and the head network are added with multiple SPD-Conv layers.
7. The integrated intelligent processing method for radar target detection and tracking on a moving platform according to claim 1 is characterized in that: The SPD-Conv layer includes: a space-to-depth layer and a non-strided convolution layer.
8. The integrated intelligent processing method for radar target detection and tracking on a moving platform according to claim 1 is characterized in that: The target detection result includes: a bounding box, a confidence level, and a category label; Based on the target detection results corresponding to each frame of P display image, the ByteTrack algorithm is used to calculate the target trajectory and complete target tracking, including: Calculate the coordinates of each target in two adjacent P display images; Determine the type of the bounding box according to the confidence level of each target in the target detection result corresponding to each frame of the P display image; the type includes a high confidence detection box and a low confidence detection box; For different types of bounding boxes, the Hungarian algorithm is used to match the objects in two adjacent P-display images, and the coordinates of the successfully matched objects and the unmatched objects are obtained; The Kalman filter is used to predict the coordinates of each target in each frame of P display image to obtain the predicted coordinates; Based on the predicted coordinates and the coordinates of the target that was not successfully matched, the Hungarian algorithm is used to rematch the targets in two adjacent frames of P display images, and the coordinates of the target that was successfully matched again are obtained; The target trajectory is determined based on the coordinates of the successfully matched target and the coordinates of the target that is successfully matched again, thereby achieving target tracking.
9. The integrated intelligent processing method for moving platform radar target detection and tracking according to claim 8 is characterized in that: According to the confidence level of each target in the target detection result corresponding to each frame of the P display image, the type of the bounding box is determined, specifically including: Determine whether the confidence of each target in the target detection result corresponding to each frame of the P display image is greater than or equal to a preset high frame threshold; If the confidence is greater than or equal to a preset high frame threshold, determining the bounding box as a high confidence detection box; If the confidence is less than the preset high frame threshold, whether the confidence is greater than or equal to the low frame threshold; If the confidence is greater than or equal to the low frame threshold, the bounding box is determined to be a low confidence detection box; If the confidence is less than the low frame threshold, the object is in the bounding box.
10. The integrated intelligent processing method for moving platform radar target detection and tracking according to claim 8, characterized in that: When the Hungarian algorithm is used to match the targets in two adjacent P-display images again, if the number of matches reaches the matching number threshold, the unmatched targets are discarded.
Citation Information
Patent Citations
Small target detection method for remote sensing image
CN118115893A
Low-illumination unmanned aerial vehicle target detection method combining EnlighttenGAN and improved YOLOv8n
CN118823302A
Radar multi-frame image target detection and tracking integrated processing method
CN118962627A
Highway pavement dynamic small target tracking detection method and system based on improved YOLOv5 and ByteTrack
CN119091394A
Unmanned aerial vehicle aerial photography small target detection method based on improved YOLOv8
CN119152390A
Cited By
Intelligent AI target detection method and system for 3D trend of elevator door
CN120544187A