Target detection optimization method and system based on Center Point model
Through the object detection optimization method based on the Center Point model, the problem of inaccurate detection of fast moving or obstructed targets in complex urban environments is solved, and high-accurate target tracking is achieved, which improves the safety and comfort of autonomous vehicles.
Patent Information
- Application Number
- CN202510446178.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-10
AI Technical Summary
The prior art is difficult to accurately detect rapidly moving or obscured targets in complex urban environments, resulting in interruption of target tracking or incorrect associations.
The object detection optimization method based on the Center Point model is adopted. By obtaining the original picture of the current frame, inputting it into the CenterPoint model for inference, obtaining the target category, center offset position and target size information, and combining the historical target center position set of the previous frame, input it into the pre-trained target center point motion model to generate the target center position set, and finally rendering this information into the original picture for object detection and tracking.
Accurate detection of fast moving or obscured targets is achieved, the accuracy of target tracking is improved, and the stable operation can be ensured in complex environments and the safety and comfort of passengers are improved.
Smart Images

Figure CN119964129A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a target detection optimization method and system based on a Center Point model. Background Art
[0002] At present, with the continuous advancement of science and technology, autonomous driving technology is gradually moving from science fiction to reality. Autonomous vehicles rely on a series of complex systems and technologies to perceive the surrounding environment, make decisions and navigate safely. Among these systems, target detection and tracking are key components to ensure that vehicles can understand and respond to dynamic traffic conditions. In the context of autonomous driving, vehicles must be able to identify various road users, such as pedestrians, cyclists, other motor vehicles, etc., and predict their behavior. To achieve this, vehicles are equipped with a variety of sensors, including cameras, laser radar (LiDAR), radar and ultrasonic sensors, etc., to collect a large amount of environmental data. However, processing this data is not easy, especially in complex urban environments, which are full of fast-moving objects, sudden pedestrians, and occlusions from time to time. Therefore, improving the ability to detect and track targets, especially in response to rapidly changing and complex conditions, has become an important research direction to promote the development of autonomous driving technology. Researchers are committed to developing smarter and more efficient algorithms to ensure that autonomous vehicles can operate stably under various conditions, improve the safety and comfort of passengers, and transportation efficiency.
[0003] In one prior art, a method and system for real-time detection of multiple targets in a vehicle road scene are disclosed. The method comprises the following steps: reading a vehicle road scene image frame, detecting targets in a current image frame using a target detection model; mapping the targets to an original image frame according to the target detection result; tracking targets marked in the current image frame using a target tracking model, visually predicting the targets through the target tracking result, adding the predictions to a prediction result set, and rechecking the prediction result set as a data set; obtaining a union of multiple target detection results and multiple target tracking prediction results; processing the union of multiple target detection results and multiple target tracking prediction results to obtain final multiple target detection and tracking results, and outputting the final multiple target detection frame to a vehicle road scene image captured by a camera through a camera projection matrix for display, thereby achieving the purpose of accurate detection.
[0004] The existing technology cannot accurately detect fast-moving or obscured targets, resulting in tracking interruption or incorrect association. Summary of the invention
[0005] The present invention provides a target detection optimization method and system based on a Center Point model, which can achieve accurate detection of fast-moving or obscured targets to improve the accuracy of target tracking.
[0006] In a first aspect, in order to solve the above technical problems, the present invention provides a target detection optimization method based on a Center Point model, comprising: Get the original picture of the current frame; Input the original image into the CenterPoint model for reasoning to obtain target category, center offset position and target size information; Inputting the target category, the center offset position, the target size information and a historical target center position set pre-stored in a previous frame into a pre-trained target center point motion model to obtain a target center position set; Rendering the target category, the target center position set and the target size information into the original image to obtain a rendered image; Target detection is performed based on the rendered image, and target tracking is performed based on the target detection result.
[0007] In an optional implementation, the inputting the original image into the CenterPoint model for reasoning to obtain the target category, center offset position and target size information includes: Input the original image into the ResNet backbone network of the optimized CenterPoint model to extract image features and obtain a feature map; The feature pyramid network is used to fuse the scene features of the feature map to obtain a fused feature map. The fused feature map is input into CenterBox Head for position prediction to obtain the center offset position and target size information; Input the fused feature map into the RetinaNet network for regression to obtain a target prediction box; The center offset position and the target prediction box are input into the KeyPoint Head for type recognition to obtain the target category.
[0008] In an optional implementation, the optimization process of the CenterPoint model includes: Input the original image into the CNN backbone network of the initial CenterPoint model to extract image features and obtain the initial feature map; Input the feature map into CenterBox Head for position prediction to obtain the initial center offset position; Inputting the initial center offset position into the KeyPoint Head to perform target type recognition to obtain an initial target category; Inputting the initial feature map into the MaskNetHead model to obtain a pixel-level segmentation map, wherein the pixel-level segmentation map includes pixel information of the target category and the category region area; Binarizing the pixel-level segmentation map to obtain a binary image; Performing target judgment according to the binary image, the initial target category, the category area and the initial center offset, and counting the initial target category, the category area and the initial center offset into a target list according to the judgment result; The target truth table in the original image is compared with the target list and the positioning loss, classification loss and segmentation loss are calculated. The initial CenterPoint model is iteratively optimized through back propagation according to the calculation results to obtain the CenterPoint model.
[0009] In an optional implementation, the target category, the center offset position, the target size information, and the historical target center position set pre-stored in the previous frame are input into a pre-trained target center point motion model to obtain the target center position set, including: Use Kalman filter as the basic model to establish the initial motion prediction model; Dynamically adjusting the process noise covariance matrix and the measurement noise covariance matrix according to the pre-stored target speed, pre-stored target acceleration and pre-stored motion position of the previous frame to obtain a motion prediction model; Input the target category, the center offset position, the target size information and the historical target center position set pre-stored in the previous frame into the motion prediction model to perform Kalman filter prediction to generate a preliminary prediction position sequence; According to the preliminary predicted position sequence and the pre-stored actual position sequence, the preliminary predicted position sequence is optimized by a dynamic time warping algorithm to obtain a predicted position sequence; Mapping adjustment is performed according to the predicted position sequence to obtain a target center position set.
[0010] In an optional implementation, optimizing the preliminary predicted position sequence by a dynamic time warping algorithm according to the preliminary predicted position sequence and the pre-stored actual position sequence to obtain the predicted position sequence includes: Calculate the corresponding Euclidean distances according to the preliminary predicted position sequence and the pre-stored actual position sequence to obtain a distance matrix; Creating a cumulative distance matrix according to the distance matrix, and filling the cumulative distance matrix through dynamic programming to obtain an optimal path; Backtracking the optimal path by comparing the cumulative distances of adjacent cells of the cumulative distance matrix to form a matching pair list; The preliminary predicted position sequence is adjusted by an interpolation method according to the matching pair list to obtain a predicted position sequence.
[0011] In an optional implementation, rendering the target category, the target center position set, and the target size information into the original image to obtain a rendered image includes: Calculating the frame corner coordinates according to the target center position set and the target size information; matching a target color and a target shape according to the target category; Draw a target logo on the original image according to the frame corner coordinates, the target color and the target shape by using a drawing function in a pre-stored image processing library to obtain a rendered image; Among them, the border corner coordinates are calculated by the following formula: ;
[0012] in, Indicates the coordinates of the upper left corner of the border. Indicates the coordinates of the upper right corner of the border. Indicates the coordinates of the lower left corner of the border. Indicates the coordinates of the lower right corner of the border. represents the center point coordinates, Indicates the target width in the target size information. Indicates the target height in the target size information.
[0013] In an optional implementation, the rendered image is compared with the rendered image of the previous frame to determine whether the target of the current frame is the same as the target of the previous frame. If the target of the current frame is the same as the target of the previous frame, the pre-stored target motion model is not updated and target detection of the next frame is performed; If the target of the current frame is different from the target of the previous frame, determine whether the pre-stored tracker list contains the target motion model of the current frame. If the pre-stored tracker list does not contain the target motion model of the current frame, the corresponding tracker is invalid and the invalid tracker is deleted. If the pre-stored tracker list contains the target motion model of the current frame, the tracker list is updated and the predicted target position is generated based on the rendered image.
[0014] In a second aspect, the present invention provides a target detection optimization system based on a Center Point model, comprising: Data acquisition module, used to obtain the original picture of the current frame; The target information reasoning module is used to input the original image into the CenterPoint model for reasoning to obtain target category, center offset position and target size information; A center position generation module, used to input the target category, the center offset position, the target size information and the historical target center position set pre-stored in the previous frame into a pre-trained target center point motion model to obtain a target center position set; A picture rendering module, rendering the target category, the target center position set and the target size information into the original picture to obtain a rendered picture; The target tracking module is used to perform target detection according to the rendered image and to track the target according to the target detection result.
[0015] In a third aspect, the present invention also provides an electronic device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements any one of the above-mentioned target detection optimization methods based on the Center Point model.
[0016] In a fourth aspect, the present invention also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the above-mentioned target detection optimization methods based on the Center Point model.
[0017] Compared with the prior art, the present invention has the following beneficial effects: The present invention discloses a target detection optimization method based on Center Point model, including obtaining the original picture of the current frame; inputting the original picture into the CenterPoint model for reasoning to obtain the target category, center offset position and target size information; inputting the target category, the center offset position, the target size information and the historical target center position set stored in the previous frame into the pre-trained target center point motion model to obtain the target center position set; rendering the target category, the target center position set and the target size information into the original picture to obtain a rendered picture; performing target detection according to the rendered picture, and performing target tracking according to the target detection result. The present invention analyzes the picture of the current frame by using the CenterPoint model to obtain the category, position offset and size information of the target. Then, combined with the target position data of the previous frame, the position of the target in the current frame is predicted and updated. This information is marked on the original picture to form a rendered picture with the detection result, so as to realize the recognition and tracking of the target, thereby realizing the accurate detection of fast-moving or obscured targets to improve the accuracy of target tracking. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a flow chart of a target detection optimization method based on a Center Point model provided by a first embodiment of the present invention; Figure 2 It is a schematic diagram of a target detection optimization system based on a Center Point model provided in the second embodiment of the present invention. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] Reference Figure 1 The first embodiment of the present invention provides a target detection optimization method based on a Center Point model, comprising the following steps: S11, obtaining the original image of the current frame; S12, inputting the original image into the CenterPoint model for reasoning to obtain target category, center offset position and target size information; S13, inputting the target category, the center offset position, the target size information, and a historical target center position set pre-stored in a previous frame into a pre-trained target center point motion model to obtain a target center position set; S14, rendering the target category, the target center position set and the target size information into the original image to obtain a rendered image; S15, performing target detection according to the rendered image, and performing target tracking according to the target detection result.
[0021] In step S11, the original picture of the current frame is obtained.
[0022] Specifically, obtaining the original image of the current frame involves receiving real-time data from the sensors equipped on the vehicle, including cameras, LiDAR, and radar, which are used to capture image information of the vehicle's surroundings. The original image obtained is unprocessed image data that directly reflects the real-time status of the environment in which the autonomous driving vehicle is located, providing the necessary visual input for subsequent target detection and tracking. This step ensures that the system can analyze and make decisions based on the latest environmental information, which is crucial for achieving accurate target detection.
[0023] In step S12, the original image is input into the CenterPoint model for reasoning to obtain the target category, center offset position and target size information.
[0024] In a specific implementation, the inputting the original image into the CenterPoint model for reasoning to obtain the target category, center offset position and target size information includes: Input the original image into the ResNet backbone network of the optimized CenterPoint model to extract image features and obtain a feature map; The feature pyramid network is used to fuse the scene features of the feature map to obtain a fused feature map. The fused feature map is input into CenterBox Head for position prediction to obtain the center offset position and target size information; Input the fused feature map into the RetinaNet network for regression to obtain a target prediction box; The center offset position and the target prediction box are input into the KeyPoint Head for type recognition to obtain the target category.
[0025] In a specific implementation, the optimization process of the CenterPoint model includes: Input the original image into the CNN backbone network of the initial CenterPoint model to extract image features and obtain the initial feature map; Input the feature map into CenterBox Head for position prediction to obtain the initial center offset position; Inputting the initial center offset position into the KeyPoint Head to perform target type recognition to obtain an initial target category; Inputting the initial feature map into the MaskNetHead model to obtain a pixel-level segmentation map, wherein the pixel-level segmentation map includes pixel information of the target category and the category region area; Binarizing the pixel-level segmentation map to obtain a binary image; Performing target judgment according to the binary image, the initial target category, the category area and the initial center offset, and counting the initial target category, the category area and the initial center offset into a target list according to the judgment result; The target truth table in the original image is compared with the target list and the positioning loss, classification loss and segmentation loss are calculated. The initial CenterPoint model is iteratively optimized through back propagation according to the calculation results to obtain the CenterPoint model.
[0026] Specifically, the original image is first fed into the ResNet backbone network to extract image features and generate feature maps. As a deep residual learning framework, ResNet solves the gradient vanishing problem in deep networks by introducing skip connections, making training more stable and efficient. In this process, the convolutional layer applies a series of filters to the input original image, each of which is responsible for capturing a specific type of feature, such as edges, textures, or shapes, thereby constructing a multi-level abstract representation.
[0027] Among them, the process of applying a series of filters to the input original picture by the convolution layer is one of the core steps of feature extraction by convolutional neural network (CNN). Each filter is a small, learnable two-dimensional matrix, which is usually small in size (such as 3x3 or 5x5) and has the same depth as the input data. When applied to color images, this means that the filter not only slides in width and height, but also across the three RGB color channels.
[0028] Specifically, suppose there is a size The original image, Represents the width, represents the height, and represents the depth, for example, for a color image, , corresponding to the red, green, and blue channels). Let there be a The filter here represents the spatial dimensions of the filter (i.e., width and height), and the depth of the filter Must match the depth of the input image. The convolution operation performs a dot product between this filter and a local region of the input image and sums the results to produce a single output value. This process can be formally expressed as: ;
[0029] in, represents the original image, represents the filter (also called convolution kernel), is the coordinate position on the output feature map, are relative coordinates within the filter window, represents the depth index of the input image and filter, is the preset bias term.
[0030] The above formula describes how a single filter works with a local area of the input image to calculate an element in the output feature map. In order to obtain a complete feature map, the filter slides over the entire input image with a certain step size and repeats the above calculation. And using zero padding, the convolution operation can be performed without changing the output size.
[0031] Considering that there are multiple such filters in the convolutional layer, each filter can capture different features (such as edges, textures, etc.), then the result of all these filters working together is a multi-channel feature map. filters, the final output feature map will be of size A three-dimensional tensor of is the padding size, is the preset step size.
[0032] For example, if there is a The original image is sized The filter has a step size of 1 and no additional padding. Then according to the above formula, the size of the output feature map will be This is because each filter produces a There are six activation maps of size , and since there are six different filters, there are six layers of such activation maps in total.
[0033] Next, the feature pyramid network (FPN) is used to fuse the obtained feature maps with scene features to generate a fused feature map. FPN enhances the feature expression capabilities at different scales through top-down paths and lateral connections. Specifically, FPN starts from the top-level features, gradually passes information downward, and combines it with high-resolution features from lower levels, which can effectively combine global context information and local detail information, thereby improving detection performance.
[0034] Subsequently, the fused feature maps are input into two different head modules: one is CenterBox Head for position prediction, and the other is RetinaNet for regressing the target prediction box. The task of CenterBox Head is to estimate the offset of the object center relative to the center of the grid cell and the object size, i.e., width and height. This process involves key point detection, where each category has a heat map representing the potential target center position. For each positive sample (i.e., the true object center), its position deviation relative to the grid cell center is calculated, and the actual size of the object is regressed. RetinaNet focuses on the precise positioning of the bounding box, predicts the candidate area through the anchor mechanism, and classifies and regresses it.
[0035] In order to determine the target category, the fused feature map needs to be input into the KeyPoint Head. The KeyPoint Head performs type recognition based on the previously obtained center offset position and target prediction box. The key here is to locate the key points of the object, such as the four corners of a vehicle or the joints of a pedestrian. Through the information of these key points, the system can accurately determine the specific category of the object.
[0036] As for the optimization process of the CenterPoint model, the original image is first input into the CNN backbone network of the initial CenterPoint model to extract the initial feature map. Then, these feature maps will enter the CenterBox Head for preliminary position prediction to generate the initial center offset position; at the same time, the initial feature map will also be sent to the MaskNetHead model to obtain the pixel-level segmentation map. The pixel-level segmentation map contains pixel information about the target category and its coverage area, which is crucial for understanding the spatial layout of the object.
[0037] After that, the pixel-level segmentation map is binarized and converted into a binary image containing only two categories: foreground (target) and background. This step simplifies the subsequent analysis process and makes it easier to distinguish which areas belong to the target of interest. Then, a comprehensive evaluation is performed based on the binary image, initial target category, category area, and initial center offset information to form a target list. This list records the relevant attributes of all detected targets.
[0038] Finally, the real target information (truth table) annotated in the original image is compared with the target list, and the positioning loss, classification loss and segmentation loss are calculated based on this. The positioning loss measures the difference between the predicted box and the real box; the classification loss reflects the degree of match between the predicted category and the actual category; the segmentation loss focuses on the consistency between the pixel-level segmentation result and the real label. Through the back propagation algorithm, the model parameters are adjusted according to these loss functions, and iterates continuously until the optimal solution is reached, and finally the optimization of the CenterPoint model is completed.
[0039] In summary, this series of operations not only improves the robustness and accuracy of the model in complex environments, but also ensures that the model can quickly respond to real-time changing driving scene requirements while maintaining high accuracy. This design idea reflects the innovative application and development trend of modern computer vision technology in the field of autonomous driving.
[0040] In step S13, the target category, the center offset position, the target size information and the historical target center position set pre-stored in the previous frame are input into the pre-trained target center point motion model to obtain the target center position set.
[0041] In a specific implementation, the target category, the center offset position, the target size information, and the historical target center position set stored in the previous frame are input into a pre-trained target center point motion model to obtain the target center position set, including: Use Kalman filter as the basic model to establish the initial motion prediction model; Dynamically adjusting the process noise covariance matrix and the measurement noise covariance matrix according to the pre-stored target speed, pre-stored target acceleration and pre-stored motion position of the previous frame to obtain a motion prediction model; Input the target category, the center offset position, the target size information and the historical target center position set pre-stored in the previous frame into the motion prediction model to perform Kalman filter prediction to generate a preliminary prediction position sequence; According to the preliminary predicted position sequence and the pre-stored actual position sequence, the preliminary predicted position sequence is optimized by a dynamic time warping algorithm to obtain a predicted position sequence; Mapping adjustment is performed according to the predicted position sequence to obtain a target center position set.
[0042] In a specific implementation, the step of optimizing the preliminary predicted position sequence by a dynamic time warping algorithm based on the preliminary predicted position sequence and the pre-stored actual position sequence to obtain the predicted position sequence includes: Calculate the corresponding Euclidean distances according to the preliminary predicted position sequence and the pre-stored actual position sequence to obtain a distance matrix; Creating a cumulative distance matrix according to the distance matrix, and filling the cumulative distance matrix through dynamic programming to obtain an optimal path; Backtracking the optimal path by comparing the cumulative distances of adjacent cells of the cumulative distance matrix to form a matching pair list; The preliminary predicted position sequence is adjusted by an interpolation method according to the matching pair list to obtain a predicted position sequence.
[0043] Specifically, firstly, an initial motion prediction model is established based on the Kalman filter. The model dynamically adjusts the process noise covariance matrix and the measurement noise covariance matrix by combining the historical target's velocity, acceleration, and motion position information, thereby obtaining a more accurate motion prediction model. Specifically, the process noise covariance matrix reflects the estimation of the uncertainty of the internal changes of the system, while the measurement noise covariance matrix measures the uncertainty of the observations. For different motion modes (such as acceleration, uniform speed, or deceleration), selecting an appropriate process noise covariance matrix can improve the tracking stability and convergence speed of the filter.
[0044] In the process of dynamically adjusting the process noise covariance matrix and the measurement noise covariance matrix, the core is to reflect the uncertainty of the changes within the system and the uncertainty of the observations based on the speed, acceleration, and motion position information of the historical target. This adjustment is crucial to improving the performance of the Kalman filter, ensuring that the filter can adapt to the actual behavior of the system and maintain good robustness and accuracy in the face of environmental changes. First, consider the process noise covariance matrix, which reflects the uncertainty and external disturbances in the system model. When new measurement data arrives, the appropriateness of the process noise covariance matrix currently in use can be determined by evaluating the difference between the predicted state and the actual measurement. If it is found that the predicted state error is large, it indicates that the changes within the system are beyond expectations. At this time, the process noise covariance matrix needs to be increased to allow for greater state uncertainty. Conversely, if the predicted state is consistent with the actual measurement, the process noise covariance matrix can be appropriately reduced, thereby reducing the trust in the new measurement data. The specific adjustment formula is as follows: ;
[0045] here, Indicates The process noise covariance matrix after iterations; is the current estimated state covariance matrix; is the state covariance matrix saved after the previous iteration.
[0046] Next is the adjustment of the measurement noise covariance matrix. The measurement noise covariance matrix describes the level of random error introduced in the measurement process. Similarly, by comparing the residuals (i.e., errors) between the predicted state and the actual measurement results, it can be determined whether the existing measurement noise covariance matrix is reasonable. If the residuals deviate significantly from zero, it means that the measurement noise is larger than expected, and the measurement noise covariance matrix should be increased; on the contrary, if the residuals are small, it means that the measurement is more accurate, and the measurement noise covariance matrix can be reduced. The adjustment formula is: ;
[0047] in, Representative The measurement noise covariance matrix after iterations is, is the actual measured value, The predicted measurement value calculated based on the predicted state.
[0048] To achieve the above adjustments, an adaptive mechanism is added to the algorithm. After each update step, the prediction error is checked and the process noise covariance matrix and the measurement noise covariance matrix are modified accordingly. For example, in a specific implementation, assume that a moving target is being tracked and the position of the target is measured using a radar. Since the target may be affected by unknown random disturbances, an adaptive unscented Kalman filter (AUKF) method is used to estimate the position and velocity of the target. AUKF not only performs standard prediction and update operations, but also dynamically updates the process noise covariance matrix and the measurement noise covariance matrix based on the prediction error to ensure that the filter is always in the optimal working state.
[0049] The significance of this adaptive adjustment is that it enables the Kalman filter to automatically adjust parameters without fully understanding the system characteristics, thereby better adapting to changes in actual conditions. Even in nonlinear or non-Gaussian noise environments, it can maintain a high estimation accuracy. In addition, this method also reduces the workload of manual parameter adjustment and improves the automation of the system.
[0050] Next, the target category, center offset position, and target size information mentioned above are input into the optimized motion prediction model together with the historical target center position set stored in the previous frame to perform the prediction stage of the Kalman filter. In this stage, the model predicts the state at the next moment based on the current state estimate and generates a preliminary predicted position sequence. The prediction formula is: ;
[0051] in, is at the time step Predictions of future states, is a pre-stored state transition matrix that describes how the system state evolves over time. is the state estimate for the previous time step, is the pre-stored control input matrix, is the pre-stored external control vector; is the forecast error covariance matrix, is the forecast error covariance matrix of the previous time step, is the process noise covariance matrix.
[0052] Then, in order to further optimize the preliminary predicted position sequence, the Dynamic Time Warping (DTW) algorithm was introduced. DTW is used to compare two time series that may have different lengths and minimize the cumulative distance between them by finding the best matching path between the two sequences. Here, the preliminary predicted position sequence is compared with the pre-stored actual position sequence in order to find the most similar correspondence between the two. The Euclidean distance is calculated to form a distance matrix, and then a cumulative distance matrix is created and filled in this matrix through a dynamic programming method until the optimal path is found. Once the optimal path is determined, the adjacent cells in the cumulative distance matrix can be backtracked to form a series of matching pair lists. These matching pairs define how each point in the preliminary predicted position sequence should be adjusted to better match the actual position sequence. The last step is to adjust the preliminary predicted position sequence using an interpolation method to ensure that the adjusted sequence is not only morphologically close to the actual position sequence, but also maintains continuity and smoothness. This step improves the quality of the predicted position sequence, making the target center position set output in the end more accurate and reliable.
[0053] In summary, by combining the Kalman filter with the DTW algorithm, we can not only effectively predict the target trajectory, but also enhance the model's ability to understand the target behavior in complex environments. This method can provide more accurate target tracking results while ensuring real-time performance, which is of great significance for application scenarios such as autonomous driving.
[0054] In step S14, the target category, the target center position set and the target size information are rendered into the original image to obtain a rendered image.
[0055] In a specific implementation, rendering the target category, the target center position set, and the target size information into the original image to obtain a rendered image includes: Calculate the frame corner coordinates according to the target center position set and the target size information; matching a target color and a target shape according to the target category; Draw a target logo on the original image according to the frame corner coordinates, the target color and the target shape by using a drawing function in a pre-stored image processing library to obtain a rendered image; Among them, the border corner coordinates are calculated by the following formula: ;
[0056] in, Indicates the coordinates of the upper left corner of the border. Indicates the coordinates of the upper right corner of the border. Indicates the coordinates of the lower left corner of the border. Indicates the coordinates of the lower right corner of the border. represents the center point coordinates, Indicates the target width in the target size information. Indicates the target height in the target size information.
[0057] Specifically, we first need to determine the coordinates of the four corner points of the rectangular frame surrounding the target based on the known target center position and target size information. The positions of these corner points can be accurately calculated using the following formula: ;
[0058] in, Indicates the coordinates of the upper left corner of the border. Indicates the coordinates of the upper right corner of the border. Indicates the coordinates of the lower left corner of the border. Indicates the coordinates of the lower right corner of the border. represents the center point coordinates, Indicates the target width in the target size information. Indicates the target height in the target size information.
[0059] Next, select the corresponding color and shape for annotation according to the category to which each target belongs. For example, for pedestrian detection tasks, a red rectangular box can be selected, while for vehicles, a blue elliptical box may be used. This classification labeling not only helps to distinguish different types of objects, but also makes the results more intuitive and easy to understand. The choice of color is based on a pre-defined color mapping table, while the shape depends on the needs of the specific application scenario.
[0060] The final step is to actually draw the target logo. This step relies on drawing functions provided by one or more image processing libraries. For example, calling or Function, passing in the previously calculated four corner point coordinates as parameters, and specifying the fill color and other style attributes.
[0061] In summary, by rendering the target category, center location set and size information into the original image, the detection results are effectively displayed. This method not only helps to quickly understand the image content, but also ensures that key information can be quickly identified and improves decision-making efficiency.
[0062] In step S15, target detection is performed based on the rendered image, and target tracking is performed based on the target detection result.
[0063] In a specific implementation, performing target detection according to the rendered image and optimizing the target motion model according to the target detection result includes: Compare the rendered image with the rendered image of the previous frame to determine whether the target of the current frame is the same as the target of the previous frame. If the target of the current frame is the same as the target of the previous frame, do not update the pre-stored target motion model and perform target detection for the next frame. If the target of the current frame is different from the target of the previous frame, determine whether the pre-stored tracker list contains the target motion model of the current frame. If the pre-stored tracker list does not contain the target motion model of the current frame, the corresponding tracker is invalid and the invalid tracker is deleted. If the pre-stored tracker list contains the target motion model of the current frame, the tracker list is updated and the predicted target position is generated based on the rendered image.
[0064] Specifically, first, the rendered image of the current frame is compared with the previous frame to determine whether the target has changed. This judgment is based on image feature matching technology, such as using a deep learning model to extract the similarity score between the two frames or calculate the degree of change in the target position, size, and appearance features. When it is found that the target of the current frame is consistent with the target of the previous frame, it means that the target has not moved significantly or changed in shape, so there is no need to update the pre-stored target motion model, and directly enter the target detection process of the next frame.
[0065] However, if the target in the current frame is different from the target in the previous frame, it is necessary to further check whether the corresponding target motion model already exists in the pre-stored tracker list. This step ensures that the system can still identify and track the same target even if the target is occluded, suddenly accelerated, or in other complex situations. For those targets that no longer appear in the current frame, that is, the tracker is invalid, these invalid trackers will be removed from the tracker list to maintain the validity and real-time performance of the list.
[0066] If the pre-stored tracker list contains the target motion model of the current frame, the update operation of the tracker list will be triggered. The update process involves several key links: one is state prediction, the second is data association, and the third is state update and trajectory management. Specifically, in the state prediction stage, the Kalman filter is combined with a simplified motion model (such as uniform motion or uniformly accelerated motion) to predict the position of the target at a certain moment in the future based on the historical information of the previous frames.
[0067] Next, in the data association stage, the Hungarian algorithm is used to match the predicted target position with the actual detected target position. The goal is to find one or more pairs of prediction boxes and detection boxes that are most likely to belong to the same physical entity. Once the matching is completed, the state update link can be entered to adjust the state parameters of the Kalman filter while maintaining the target trajectory. At this time, if there is a high-confidence detection box that is not matched and its confidence exceeds the set threshold, a new tracking trajectory is created; conversely, for low-confidence detection boxes or trajectories that have not been successfully matched for a long time, they are marked for deletion.
[0068] Finally, the process of generating the predicted target position not only relies on historical trajectory analysis, but also takes into account factors such as the target's speed and acceleration. By dynamically adjusting the process noise covariance matrix and the measurement noise covariance matrix, the prediction results are ensured to be as close to the actual situation as possible. For example, when the target accelerates or decelerates, the process noise covariance matrix is appropriately increased to allow for greater state uncertainty; while for more stable motion modes, the process noise covariance matrix is reduced to reduce sensitivity to external disturbances. Similarly, the measurement noise covariance matrix is adjusted according to the size of the residual to ensure the reliability of the measurement value.
[0069] In summary, by performing target detection on the rendered image and optimizing the target motion model based on the detection results, effective tracking of moving targets is achieved. This method not only improves the stability and accuracy of the tracking system, but also demonstrates good adaptability when facing complex scenes. In addition, by continuously updating and maintaining the tracker list, it ensures that each target can receive continuous attention and maintain tracking continuity even in the event of occlusion or rapid movement. The entire process achieves a seamless transition from detection to tracking.
[0070] The following describes the working process of the present invention using a relatively common scenario as an example. Figure 1 , a target detection optimization method based on Center Point model, comprising the following steps: When applying the target detection optimization method based on the Center Point model, the self-driving vehicle captures the original picture of the current road scene through the equipped camera. As the vehicle drives, the sensor continuously updates the environmental information. Every time the system receives a new frame of image, it will immediately input it into the pre-trained CenterPoint model for reasoning analysis.
[0071] Next, the system uses the ResNet backbone network to extract image features and fuses scene features at different levels with the help of the feature pyramid network. After processing by the CenterBox Head and RetinaNet networks, the target's location prediction and bounding box are determined, and then the KeyPoint Head further identifies the specific category of the target. These steps work together to enable the system to accurately understand each element in the image.
[0072] For continuous video streams, the system not only relies on the information of the current frame, but also combines the historical data of the previous frame to estimate the movement trend of the target. The motion prediction model established by the Kalman filter dynamically adjusts the parameters to ensure accurate tracking of the target position even in fast movement or in the presence of occlusion. The dynamic time warping algorithm in this process helps optimize the predicted path and ensure the consistency and coherence of the trajectory.
[0073] After completing the above processing, the system renders the identified target information back to the original image to display the detection results in an intuitive form. Each target is assigned a specific color and shape to clearly distinguish different object types. Finally, a new round of target detection and tracking is performed based on the rendered image to maintain the system's real-time responsiveness.
[0074] Throughout the process, the method significantly enhances the ability to understand target behavior in complex traffic environments, especially for challenging fast-moving objects or partial occlusions. This method ensures that autonomous vehicles can perceive the surrounding environment more reliably, improves driving safety, and also promotes the overall performance of intelligent transportation systems.
[0075] Reference Figure 2 The second embodiment of the present invention provides a target detection optimization system based on a Center Point model, comprising: Data acquisition module, used to obtain the original picture of the current frame; The target information reasoning module is used to input the original image into the CenterPoint model for reasoning to obtain target category, center offset position and target size information; A center position generation module, used to input the target category, the center offset position, the target size information and the historical target center position set pre-stored in the previous frame into a pre-trained target center point motion model to obtain a target center position set; A picture rendering module, rendering the target category, the target center position set and the target size information into the original picture to obtain a rendered picture; The target tracking module is used to perform target detection according to the rendered image and to track the target according to the target detection result.
[0076] It should be noted that the target detection optimization device based on the Center Point model provided in an embodiment of the present invention is used to execute all the process steps of the target detection optimization method based on the Center Point model in the above embodiment. The working principles and beneficial effects of the two correspond one to one, and thus will not be repeated here.
[0077] An embodiment of the present invention further provides an electronic device. The electronic device includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a target detection optimization program based on a Center Point model. When the processor executes the computer program, the steps in the above-mentioned target detection optimization method embodiments based on a Center Point model are implemented, such as Figure 1 Alternatively, when the processor executes the computer program, the functions of the modules / units in the above-mentioned device embodiments are realized, such as the target detection optimization module based on the Center Point model.
[0078] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, which are used to describe the execution process of the computer program in the electronic device.
[0079] The electronic device may be a computing device such as a desktop computer, a notebook, a PDA, and a smart tablet. The electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art will appreciate that the above components are merely examples of electronic devices and do not constitute a limitation on the electronic device. The electronic device may include more or fewer components than the above components, or may combine certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.
[0080] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the electronic device, and uses various interfaces and lines to connect various parts of the entire electronic device.
[0081] The memory can be used to store the computer program and / or module, and the processor realizes various functions of the electronic device by running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0082] Wherein, if the module / unit integrated in the electronic device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0083] It should be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the accompanying drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art may understand and implement it without paying any creative effort.
[0084] The specific embodiments described above further illustrate the purpose, technical solutions and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. It is particularly pointed out that for those skilled in the art, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention should be included in the scope of protection of the present invention.
Claims
1. A target detection optimization method based on Center Point model, characterized in that: Executed by the controller, including: Get the original picture of the current frame; Input the original image into the CenterPoint model for reasoning to obtain target category, center offset position and target size information; Inputting the target category, the center offset position, the target size information and a historical target center position set pre-stored in a previous frame into a pre-trained target center point motion model to obtain a target center position set; Rendering the target category, the target center position set and the target size information into the original image to obtain a rendered image; Target detection is performed based on the rendered image, and target tracking is performed based on the target detection result.
2. The target detection optimization method based on Center Point model according to claim 1, characterized in that: The original image is input into the CenterPoint model for reasoning to obtain target category, center offset position and target size information, including: Input the original image into the ResNet backbone network of the optimized CenterPoint model to extract image features and obtain a feature map; Using a feature pyramid network to perform scene feature fusion on the feature map to obtain a fused feature map; The fused feature map is input into CenterBox Head for position prediction to obtain the center offset position and target size information; Input the fused feature map into the RetinaNet network for regression to obtain a target prediction box; The center offset position and the target prediction box are input into the KeyPoint Head for type recognition to obtain the target category.
3. The target detection optimization method based on Center Point model according to claim 2, characterized in that: The optimization process of the CenterPoint model includes: Input the original image into the CNN backbone network of the initial CenterPoint model to extract image features and obtain the initial feature map; Input the feature map into CenterBox Head for position prediction to obtain the initial center offset position; Inputting the initial center offset position into the KeyPoint Head to perform target type recognition to obtain an initial target category; Inputting the initial feature map into the MaskNetHead model to obtain a pixel-level segmentation map, wherein the pixel-level segmentation map includes pixel information of the target category and the category region area; Binarizing the pixel-level segmentation map to obtain a binary image; Performing target judgment according to the binary image, the initial target category, the category area and the initial center offset, and counting the initial target category, the category area and the initial center offset into a target list according to the judgment result; The target truth table in the original image is compared with the target list and the positioning loss, classification loss and segmentation loss are calculated. The initial CenterPoint model is iteratively optimized through back propagation according to the calculation results to obtain the CenterPoint model.
4. The target detection optimization method based on Center Point model according to claim 1, characterized in that: The target category, the center offset position, the target size information and the historical target center position set stored in the previous frame are input into the pre-trained target center point motion model to obtain the target center position set, including: Use Kalman filter as the basic model to establish the initial motion prediction model; Dynamically adjusting the process noise covariance matrix and the measurement noise covariance matrix according to the pre-stored target speed, pre-stored target acceleration and pre-stored motion position of the previous frame to obtain a motion prediction model; Input the target category, the center offset position, the target size information and the historical target center position set pre-stored in the previous frame into the motion prediction model to perform Kalman filter prediction to generate a preliminary prediction position sequence; According to the preliminary predicted position sequence and the pre-stored actual position sequence, the preliminary predicted position sequence is optimized by a dynamic time warping algorithm to obtain a predicted position sequence; Mapping adjustment is performed according to the predicted position sequence to obtain a target center position set.
5. The target detection optimization method based on Center Point model according to claim 4 is characterized in that: The step of optimizing the preliminary predicted position sequence by a dynamic time warping algorithm according to the preliminary predicted position sequence and the pre-stored actual position sequence to obtain a predicted position sequence includes: Calculate the corresponding Euclidean distances according to the preliminary predicted position sequence and the pre-stored actual position sequence to obtain a distance matrix; Creating a cumulative distance matrix according to the distance matrix, and filling the cumulative distance matrix through dynamic programming to obtain an optimal path; Backtracking the optimal path by comparing the cumulative distances of adjacent cells of the cumulative distance matrix to form a matching pair list; The preliminary predicted position sequence is adjusted by an interpolation method according to the matching pair list to obtain a predicted position sequence.
6. The target detection optimization method based on Center Point model according to claim 1, characterized in that: The step of rendering the target category, the target center position set, and the target size information into the original image to obtain a rendered image includes: Calculate the frame corner coordinates according to the target center position set and the target size information; matching a target color and a target shape according to the target category; Draw a target logo on the original image according to the frame corner coordinates, the target color and the target shape by using a drawing function in a pre-stored image processing library to obtain a rendered image; Among them, the border corner coordinates are calculated by the following formula: ;in, Indicates the coordinates of the upper left corner of the border. Indicates the coordinates of the upper right corner of the border. Indicates the coordinates of the lower left corner of the border. Indicates the coordinates of the lower right corner of the border. represents the center point coordinates, Indicates the target width in the target size information. Indicates the target height in the target size information.
7. The target detection optimization method based on Center Point model according to claim 1, characterized in that: The performing target detection according to the rendered image and optimizing the target motion model according to the target detection result includes: Compare the rendered image with the rendered image of the previous frame to determine whether the target of the current frame is the same as the target of the previous frame. If the target of the current frame is the same as the target of the previous frame, do not update the pre-stored target motion model and perform target detection for the next frame. If the target of the current frame is different from the target of the previous frame, determine whether the pre-stored tracker list contains the target motion model of the current frame. If the pre-stored tracker list does not contain the target motion model of the current frame, the corresponding tracker is invalid and the invalid tracker is deleted. If the pre-stored tracker list contains the target motion model of the current frame, the tracker list is updated and the predicted target position is generated based on the rendered image.
8. A target detection optimization system based on Center Point model, characterized in that: include: Data acquisition module, used to obtain the original picture of the current frame; The target information reasoning module is used to input the original image into the CenterPoint model for reasoning to obtain target category, center offset position and target size information; A center position generation module, used to input the target category, the center offset position, the target size information and the historical target center position set pre-stored in the previous frame into a pre-trained target center point motion model to obtain a target center position set; A picture rendering module, rendering the target category, the target center position set and the target size information into the original picture to obtain a rendered picture; The target tracking module is used to perform target detection according to the rendered image, and to track the target according to the target detection result.
9. An electronic device, characterized in that: It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, it implements the target detection optimization method based on the Center Point model as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the target detection optimization method based on the Center Point model as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Three-dimensional target detection method based on point cloud time sequence information fusion
CN112418084A
Detection result real-time visualization method based on gradient magnetic field
CN112578462A
Space target multi-source data parameterized simulation and MinCenter Net fusion detection method
CN113096058A
Method, device and equipment for detecting and tracking dynamic target of unmanned excavator
CN116740146A
Methods and systems for intelligent collection and analysis of vehicle data
US20190025813A1
Cited By
Target tracking method, electronic equipment and computer readable storage medium
CN120655682A