A vehicle 3D target detection method, system, device and medium

CN120852737BActive Publication Date: 2026-09-29NANJING JITU NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510799549.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2026-09-29
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

[0002]目前,单目视觉3D目标检测算法通常分为两类,第一种是基于热力图的方式,热力图反映了网络预测结果的置信度分布,优势在于预测精度较高,但是多个热力图的引入以及复杂的后处理操作带来了巨大的计算开销,无法满足实时性的推理要求

Benefits of technology

[0009]通过采集待检测区域的RGB图像数据并进行预处理,利用CenterNet网络实现3D目标检测,具有如下几点效果:第一、预处理步骤能够将不同来源、格式各异的RGB图像统一为标准化的目标图像格式,确保输入数据的一致性和规范性,为后续的检测过程提供稳定可靠的基础,从而提高检测的准确性和稳定性。第二、CenterNet网络的使用充分发挥了其在目标检测领域的优势,主干网络能够高效地对输入图像进行特征提取,捕捉到图像中丰富的语义信息和空间特征,为后续的预测提供有力支持。预测头的设计则能够同时预测目标中心点热力图、8维中心点向量以及中心点偏移量,这种多维度的预测方式使得检测结果更加全面和精确。通过对目标中心点热力图进行偏移量调整,能够进一步优化中心点的定位精度,确保检测到的目标中心点位置更加准确,进而提高整个3D目标检测的准确性。第三、本发明能够检测机动车、非机动车和行人等多种类型的目标,具有良好的通用性和适应性,能够满足复杂场景下的多样化检测需求。整合后的中心点热力图和8维中心点向量输出,为后续的分析和应用提供了丰富的信息,便于进一步的处理和决策制定,例如在交通监控、自动驾驶等领域,可以基于这些检测结果实现对交通流量的分析、车辆和行人的行为预测等功能,提升系统的智能化水平和应用价值。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852737B_ABST
    Figure CN120852737B_ABST
Patent Text Reader

Abstract

The application discloses a kind of vehicle 3D target detection method, system, equipment and medium, it is related to vehicle target detection technical field, method includes: the at least one RGB image data of any detection area is collected, and all RGB images are preprocessed, and the target RGB image of uniform format is obtained;The target RGB image data is sequentially input into pre-training model, and the 3D target detection result corresponding to each target RGB image data is obtained, and the 3D target detection result includes: the center point coordinate of any target and the 8-dimensional center point vector of any target, and the target is: motor vehicle, non-motor vehicle or pedestrian.The application can ensure that the position of the detected target center point is more accurate, and then improve the accuracy of the whole 3D target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle target detection technology, and in particular to a method, system, device and medium for 3D vehicle target detection. Background Technology

[0002] Currently, monocular vision 3D object detection algorithms are generally divided into two categories. The first is based on heatmaps, which reflect the confidence distribution of the network's prediction results. Its advantage lies in high prediction accuracy, but the introduction of multiple heatmaps and complex post-processing operations brings huge computational overhead, failing to meet real-time inference requirements. The second category is based on regression, which offers faster inference speeds compared to the first method. However, this type of algorithm is typically anchor-based, and the setting of anchor size and proportion has a significant impact on the final prediction accuracy, resulting in weak generalization ability. Summary of the Invention

[0003] The technical problem to be solved by this invention is to address the shortcomings of existing technologies, and specifically provides a method, system, device, and medium for 3D vehicle target detection, as detailed below:

[0004] 1) In a first aspect, the present invention provides a method for 3D target detection of vehicles, the specific technical solution of which is as follows:

[0005] Acquire at least one RGB image data of any region to be detected, and preprocess all RGB images to obtain a target RGB image in a uniform format;

[0006] The target RGB image data are sequentially input into the pre-trained model to obtain the 3D target detection result corresponding to each target RGB image data. The 3D target detection result includes: the center point coordinates of any target and the 8-dimensional center point vector of any target. The target is: motor vehicle, non-motor vehicle or pedestrian.

[0007] The pre-trained model is the CenterNet network, which includes a backbone network and a prediction head. The backbone network is used to extract features from the input image, and the prediction head is used to predict the target center point heatmap, 8-dimensional center point vector, and center point offset corresponding to the input image based on the feature extraction results. The target center point heatmap is adjusted based on the center point offset to generate a center point heatmap. The center point heatmap and the 8-dimensional center point vector are then integrated and output.

[0008] The beneficial effects of the vehicle 3D target detection method provided by this invention are as follows:

[0009] By collecting and preprocessing RGB image data of the target area, 3D object detection is achieved using the CenterNet network, achieving the following effects: First, the preprocessing step unifies RGB images from different sources and with varying formats into a standardized target image format, ensuring the consistency and standardization of input data and providing a stable and reliable foundation for subsequent detection processes, thereby improving the accuracy and stability of detection. Second, the use of the CenterNet network fully leverages its advantages in the field of object detection. The backbone network can efficiently extract features from the input image, capturing rich semantic information and spatial features, providing strong support for subsequent prediction. The prediction head design can simultaneously predict the target center point heatmap, the 8-dimensional center point vector, and the center point offset. This multi-dimensional prediction method makes the detection results more comprehensive and accurate. By adjusting the offset of the target center point heatmap, the positioning accuracy of the center point can be further optimized, ensuring that the detected target center point position is more accurate, thereby improving the overall accuracy of 3D object detection. Third, this invention can detect various types of targets such as motor vehicles, non-motor vehicles, and pedestrians, possessing good versatility and adaptability, and can meet diverse detection needs in complex scenarios. The integrated center point heatmap and 8-dimensional center point vector output provide rich information for subsequent analysis and application, facilitating further processing and decision-making. For example, in fields such as traffic monitoring and autonomous driving, these detection results can be used to analyze traffic flow and predict the behavior of vehicles and pedestrians, thereby improving the system's intelligence level and application value.

[0010] Based on the above solution, the present invention can be further improved as follows.

[0011] Furthermore, the process of obtaining the training set for the pre-trained model is as follows:

[0012] The target training area is captured by an image acquisition device to obtain at least n target images. A cube annotation tool is used to annotate any target image with a 3D bounding box to obtain the input image corresponding to each target image. All input images are integrated to obtain the training set.

[0013] Furthermore, the process of annotating any target image with a 3D bounding box using the cube annotation tool is as follows:

[0014] Determine the initial 3D bounding box of any target in any target image. The initial 3D bounding box is composed of the 8 vertices corresponding to two non-overlapping rectangles.

[0015] The position coordinates of any one of the eight vertices are adjusted using a preset strategy, and the 3D bounding box of the target is determined based on the adjusted vertex.

[0016] Furthermore, the loss functions of the pre-trained model include: center point heatmap loss function, vector loss function, and offset loss function.

[0017] 2) Secondly, the present invention also provides a 3D target detection system, the specific technical solution of which is as follows:

[0018] The acquisition module is used to: acquire at least one RGB image data of any region to be detected, and preprocess all RGB images to obtain a target RGB image in a uniform format;

[0019] The detection module is used to: sequentially input the target RGB image data into the pre-trained model to obtain the 3D target detection result corresponding to each target RGB image data. The 3D target detection result includes: the center point coordinates of any target and the 8-dimensional center point vector of any target. The target is: motor vehicle, non-motor vehicle or pedestrian.

[0020] The pre-trained model is the CenterNet network, which includes a backbone network and a prediction head. The backbone network is used to extract features from the input image, and the prediction head is used to predict the target center point heatmap, 8-dimensional center point vector, and center point offset corresponding to the input image based on the feature extraction results. The target center point heatmap is adjusted based on the center point offset to generate a center point heatmap. The center point heatmap and the 8-dimensional center point vector are then integrated and output.

[0021] Based on the above solution, the present invention can be further improved as follows.

[0022] Furthermore, the process of obtaining the training set for the pre-trained model is as follows:

[0023] The target training area is captured by an image acquisition device to obtain at least n target images. A cube annotation tool is used to annotate any target image with a 3D bounding box to obtain the input image corresponding to each target image. All input images are integrated to obtain the training set.

[0024] Furthermore, the process of annotating any target image with a 3D bounding box using the cube annotation tool is as follows:

[0025] Determine the initial 3D bounding box of any target in any target image. The initial 3D bounding box is composed of the 8 vertices corresponding to two non-overlapping rectangles.

[0026] The position coordinates of any one of the eight vertices are adjusted using a preset strategy, and the 3D bounding box of the target is determined based on the adjusted vertex.

[0027] Furthermore, the loss functions of the pre-trained model include: center point heatmap loss function, vector loss function, and offset loss function.

[0028] 3) In a third aspect, the present invention also provides an electronic device, the electronic device including a processor coupled to a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to enable the electronic device to perform any of the methods described above.

[0029] 4) In a fourth aspect, the present invention also provides a computer-readable storage medium storing at least one computer program, which is loaded and executed by a processor to enable a computer to perform any of the above methods.

[0030] It should be noted that the beneficial effects of the technical solutions of the second to fourth aspects of the present invention and their corresponding possible implementations can be found in the above description of the technical effects of the first aspect and its corresponding possible implementations, and will not be repeated here. Attached Figure Description

[0031] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0032] Figure 1 This is a flowchart illustrating a vehicle 3D target detection method according to an embodiment of the present invention;

[0033] Figure 2 This is a schematic diagram of the model output result of a vehicle 3D target detection method according to an embodiment of the present invention;

[0034] Figure 3 This is a structural framework diagram of an electronic device according to the present invention. Detailed Implementation

[0035] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0036] like Figure 1 As shown, a vehicle 3D target detection method according to an embodiment of the present invention includes the following steps:

[0037] S1, acquire at least one RGB image data of any region to be detected, and preprocess all RGB images to obtain a target RGB image in a uniform format;

[0038] S2, input the target RGB image data sequentially into the pre-trained model to obtain the 3D target detection result corresponding to each target RGB image data. The 3D target detection result includes: the center point coordinates of any target and the 8-dimensional center point vector of any target. The target is: motor vehicle, non-motor vehicle or pedestrian.

[0039] The pre-trained model is the CenterNet network, which includes a backbone network and a prediction head. The backbone network is used to extract features from the input image, and the prediction head is used to predict the target center point heatmap, 8-dimensional center point vector, and center point offset corresponding to the input image based on the feature extraction results. The target center point heatmap is adjusted based on the center point offset to generate a center point heatmap. The center point heatmap and the 8-dimensional center point vector are then integrated and output.

[0040] The beneficial effects of the vehicle 3D target detection method provided by this invention are as follows:

[0041] By collecting and preprocessing RGB image data of the target area, 3D object detection is achieved using the CenterNet network, achieving the following effects: First, the preprocessing step unifies RGB images from different sources and with varying formats into a standardized target image format, ensuring the consistency and standardization of input data and providing a stable and reliable foundation for subsequent detection processes, thereby improving the accuracy and stability of detection. Second, the use of the CenterNet network fully leverages its advantages in the field of object detection. The backbone network can efficiently extract features from the input image, capturing rich semantic information and spatial features, providing strong support for subsequent prediction. The prediction head design can simultaneously predict the target center point heatmap, the 8-dimensional center point vector, and the center point offset. This multi-dimensional prediction method makes the detection results more comprehensive and accurate. By adjusting the offset of the target center point heatmap, the positioning accuracy of the center point can be further optimized, ensuring that the detected target center point position is more accurate, thereby improving the overall accuracy of 3D object detection. Third, this invention can detect various types of targets such as motor vehicles, non-motor vehicles, and pedestrians, possessing good versatility and adaptability, and can meet diverse detection needs in complex scenarios. The integrated center point heatmap and 8-dimensional center point vector output provide rich information for subsequent analysis and application, facilitating further processing and decision-making. For example, in fields such as traffic monitoring and autonomous driving, these detection results can be used to analyze traffic flow and predict the behavior of vehicles and pedestrians, thereby improving the system's intelligence level and application value.

[0042] In S1, the area to be detected refers to the road area that the data collection device can collect, and the road area is open to vehicles and pedestrians. Depending on the location where the data collection device is set up, the location of the area to be detected will also be different, including but not limited to: intersections, highway intersections, etc.

[0043] In another embodiment of this solution, the preprocessing of all RGB images can be as follows:

[0044] Use interpolation algorithms (such as bilinear interpolation, nearest neighbor interpolation, etc.) to scale the images to ensure that the width and height of all input images are consistent with the model requirements.

[0045] It should be noted that, in order to ensure that the subsequent pre-trained model can more clearly identify the features in the input image, during the image scaling process, it is necessary to ensure that the pixels of the scaled first image are the same as those of the original input image, or that the pixels of the first image are within the preset range of the input image's pixels. The preset range is adjusted according to actual needs. If the area to be detected is a road section or intersection with high traffic volume, the pixel requirement is high to avoid missing detailed features, so the preset range is set to be smaller. In addition, if the acquisition device that acquires RGB image data has a high pixel count, the preset range can be adjusted to be larger accordingly. The specific adjustment data is not limited here.

[0046] Since this solution targets more than one detection area, before scaling, the unique ID of the acquisition device corresponding to the RGB image data is determined, and the corresponding historical scaling ratio is determined based on the unique ID. The RGB image data corresponding to the unique ID is then uniformly scaled based on the historical scaling ratio.

[0047] In another embodiment of this solution, traffic flow for any area to be detected is statistically analyzed in real time, and the historical scaling ratio corresponding to the unique ID is adjusted based on the statistical results. The specific process is as follows:

[0048] The target time period type corresponding to the current moment is determined (time period types include: peak time, night time, daily time, etc.), and the first historical scaling ratio corresponding to the target time period type is retrieved from the historical database. Based on the difference between the average value per second of traffic flow during daily time and the average value per second during peak time, three difference ranges are constructed. The difference data corresponding to the first range, the second range, and the third range gradually increase. The difference between the number of targets in the 3D target detection results corresponding to the previous moment and the number of targets in the 3D target detection results corresponding to the current moment is calculated to obtain the calculation result. When the calculation result falls into the first range or is less than the minimum value of the first range, the historical scaling ratio corresponding to the unique ID is maintained for processing. When the calculation result falls into the second range, the first historical scaling ratio is used as the historical scaling ratio corresponding to the unique ID for processing. When the calculation result falls into the third range, the difference calculation is continuously performed within a fixed period, and the calculation result is associated with the time. The number of calculation results falling into different ranges and the corresponding ratios are determined within the fixed period (with the previous moment as the end time of the fixed period). Based on the weights corresponding to the different ranges set in advance, the historical scaling ratio corresponding to the current moment is determined.

[0049] The process of determining the historical scaling ratio corresponding to the current moment is as follows:

[0050] The historical scaling ratio corresponding to the current moment is determined by the sum of the first result, the second result, and the third result.

[0051] The first result is the product of the proportion falling into the first range and the weight corresponding to the first range and the historical scaling ratio corresponding to the unique ID;

[0052] The second result is: the product of the proportion falling into the second range and the weight corresponding to the second range and the first historical scaling ratio;

[0053] The third result is the product of the proportion falling into the third range, the weight corresponding to the third range, and the sum of each second historical scaling ratio, where the second historical scaling ratio is the corresponding historical scaling ratio determined by the calculation results based on falling into the third range, which is found in the historical database.

[0054] Furthermore, when the calculation result falls within the third range, and the time period from the start of collecting at least one RGB image data of any region to be detected to the end of the previous time is smaller than the fixed period, then the first historical scaling ratio is used as the initial historical scaling ratio for processing.

[0055] It should be noted that there are no specific limitations on how the endpoint values ​​corresponding to the first, second, and third ranges are determined.

[0056] In another embodiment of this scheme, the CenterNet network consists of two parts: a backbone network and a prediction head. The backbone network uses DLA34 for feature extraction from the input image. The model input has a height and width of 960×544, with a downsampling factor R=4. After feature extraction by the backbone network, the feature map has a height and width of 240×136. The prediction head consists of three convolutional modules, each composed of a 3×3 convolutional kernel, a ReLU activation function, and a 1×1 convolutional kernel, used to predict the target center point heatmap, the 8-dimensional center point vector, and the center point offset, respectively. The center point heatmap prediction head branch has C 1×1 convolutional kernels, where C represents the number of categories in the 3D object detection dataset; the center point vector prediction head branch has 8 1×1 convolutional kernels, corresponding to the 8 dimensions of the vector; and the center point offset prediction head branch has 2 1×1 convolutional kernels, used to predict the offset of the center point in the x-axis and y-axis directions during downsampling.

[0057] Furthermore, the process of obtaining the training set for the pre-trained model is as follows:

[0058] The target training area is captured by an image acquisition device to obtain at least n target images. A cube annotation tool is used to annotate any target image with a 3D bounding box to obtain the input image corresponding to each target image. All input images are integrated to obtain the training set.

[0059] Furthermore, the process of annotating any target image with a 3D bounding box using the cube annotation tool is as follows:

[0060] Determine the initial 3D bounding box of any target in any target image. The initial 3D bounding box is composed of the 8 vertices corresponding to two non-overlapping rectangles.

[0061] The position coordinates of any one of the eight vertices are adjusted using a preset strategy, and the 3D bounding box of the target is determined based on the adjusted vertex.

[0062] In another embodiment of this scheme, the Cuboid cube annotation tool is used to annotate the 3D bounding boxes. Each 3D bounding box consists of two rectangles, one in front and one behind. During annotation, the size and relative position of the rectangles can be adjusted according to the target size to fit the target edge and highlight the perspective relationship of the target in the 2D image. Simultaneously, the x-axis coordinates of the top and bottom corners on each side of each rectangle are the same, ensuring that opposite sides of the rectangle are always parallel and equal. This approach ensures that the 3D detection box fits the target while limiting certain degrees of freedom, guaranteeing consistent annotation across different targets. Each RGB image in the 3D object detection dataset corresponds to one input image. The input image is a text file with a .txt extension, recording the category index of the target in the corresponding image and the absolute coordinates of the eight corner points of the 3D bounding box in the image, line by line.

[0063] Furthermore, the loss functions of the pre-trained model include: center point heatmap loss function, vector loss function, and offset loss function.

[0064] In another embodiment of this scheme, the intersection of the lines connecting the three sets of opposite center points of the 3D bounding box is denoted as the center point of the 3D bounding box (i.e., the target center point), and the 8-dimensional center point vector is denoted as V. c =(l,r,t,b) o ,f,b a (θ, γ), where l represents the distance between the center point of the 3D bounding box and the left midpoint, r represents the distance between the center point of the 3D bounding box and the right midpoint, t represents the distance between the center point of the 3D bounding box and the top midpoint, and b represents the distance between the center point of the 3D bounding box and the top midpoint. o f represents the distance between the center point of the 3D bounding box and the bottom midpoint, f represents the distance between the center point of the 3D bounding box and the front midpoint, and b represents the distance between the center point of the 3D bounding box and the front midpoint. a Let θ represent the distance between the center point of the 3D bounding box and the subsequent midpoint, γ represent the angle between f and the positive x-axis, and γ represent the angle between r and the positive x-axis. All of these parameters can be calculated using labeled data from a 3D object detection dataset during the training phase. Therefore, the loss function L of the center-point-based 3D object detection model is... 3Ddet Including the center point heatmap loss function L K Vector loss function L Vc and offset loss function L off The specific calculation formula is as follows:

[0065] L 3Ddet =L K +λ Vc L Vc +λ off L off

[0066] λ Vc and λ off The hyperparameters of the loss function are set to 0.1 and 1 respectively; LK For the center point heatmap loss, for each real target center point p in the input image, after downsampling, an equivalent real target center point is obtained on the predicted feature map by rounding down, denoted as . Where P is the coordinate of the center point of the real target in the input image, denoted as P = (x, y), R is the downsampling factor of the input image, and then a Gaussian kernel is applied. Distribute all real target center points onto the heatmap Above, here σ p The target scale adaptive deviation coefficient. This represents the x-axis coordinate of the equivalent true target center point on the predicted feature map. This represents the y-axis coordinate of the equivalent true target center point on the predicted feature map, and is a heatmap of the network prediction. The loss L between the heatmap and the actual target center point K This can be represented by focal loss:

[0067]

[0068] In the above formula, α and β are hyperparameters of focal loss, and N is the number of target center points in the image.

[0069] L Vc The center point vector loss is calculated using the Smooth L1 loss formula, as follows:

[0070]

[0071] L off The loss is for center point offset, where Δb represents the difference between the predicted center point vector and the true center point vector. To reduce the target center point offset error introduced during downsampling, the model needs to predict the offset for each target center point. Offset prediction can improve the detection accuracy of small target locations. The center point offset loss during the training phase can be represented by L1 loss:

[0072]

[0073] in, Where N is the actual offset, and N is the number of target center points in the image. The model represents the equivalent ground truth center point on the feature map. The predicted offset.

[0074] Example 1: RGB image data is collected using vehicle-mounted cameras or roadside cameras. Based on practical experience, the image data should include different scenes and perspectives to improve the model's detection capability and generalization performance during the training phase.

[0075] A 3D object detection dataset was created, and the Cuboid tool of the CVAT (Computer Vision Annotation Tool) data annotation platform was used to annotate the objects with 3D bounding boxes. The eight corner points of the 3D bounding box are formed by the corner points of the preceding and following rectangles. The input image is stored in the following format: Where cls_id is the target category index. and These represent the absolute coordinates of the four corner points of the front and back rectangles of the 3D bounding box in the image, respectively, with i = 1 representing the top left corner point of the rectangle.

[0076] During the data loading phase of training, the coordinates of the target center point are calculated using the input image, denoted as (C). x C y ), and the coordinates of the center points of each face of the 3D bounding box. The distance from the target center point to the front center point of the 3D bounding box (F) x ,F y Taking the distance f as an example, the calculation formula is as follows:

[0077]

[0078] f = [(F x -C x ) 2 +(F y -C y ) 2 ] 1 / 2

[0079] Build the improved CenterNet network, and train the CenterNet model using the images collected in the above steps and the input image until the loss function converges.

[0080] After training, the network model will predict the category, center point coordinates, and center point vector of all targets in the image. Post-processing converts the model's predictions into 3D bounding box corner coordinates. The calculation methods for the midpoint coordinates of the front and back rectangles of the 3D bounding box are as follows:

[0081] F x ,F y =C x +cosθ·f,C y +sinθ·f

[0082] B x B y =C x +cos(π-θ)·b a C y -sin(π-θ)·b a

[0083] Then, using the midpoint coordinates of the two rectangles and the center point vector predicted by the model, the coordinates of the eight corner points of the 3D bounding box are calculated. Taking the front rectangle as an example, the calculation method is as follows:

[0084]

[0085]

[0086] After obtaining the coordinates of the 8 corner points, 3D bounding boxes are drawn in sequence to visualize the detection results. The detection results are as follows: Figure 2 As shown.

[0087] In the above embodiments, although the steps are numbered S1, S2, etc., they are only specific embodiments given by the present invention. Those skilled in the art can adjust the execution order of S1, S2, etc. according to the actual situation, which is also within the protection scope of the present invention. It can be understood that in some embodiments, some or all of the above embodiments may be included.

[0088] The present invention also provides a 3D target detection system, the specific technical solution of which is as follows:

[0089] The acquisition module is used to: acquire at least one RGB image data of any region to be detected, and preprocess all RGB images to obtain a target RGB image in a uniform format;

[0090] The detection module is used to: sequentially input the target RGB image data into the pre-trained model to obtain the 3D target detection result corresponding to each target RGB image data. The 3D target detection result includes: the center point coordinates of any target and the 8-dimensional center point vector of any target. The target is: motor vehicle, non-motor vehicle or pedestrian.

[0091] The pre-trained model is the CenterNet network, which includes a backbone network and a prediction head. The backbone network is used to extract features from the input image, and the prediction head is used to predict the target center point heatmap, 8-dimensional center point vector, and center point offset corresponding to the input image based on the feature extraction results. The target center point heatmap is adjusted based on the center point offset to generate a center point heatmap. The center point heatmap and the 8-dimensional center point vector are then integrated and output.

[0092] Based on the above solution, the present invention can be further improved as follows.

[0093] Furthermore, the process of obtaining the training set for the pre-trained model is as follows:

[0094] The target training area is captured by an image acquisition device to obtain at least n target images. A cube annotation tool is used to annotate any target image with a 3D bounding box to obtain the input image corresponding to each target image. All input images are integrated to obtain the training set.

[0095] Furthermore, the process of annotating any target image with a 3D bounding box using the cube annotation tool is as follows:

[0096] Determine the initial 3D bounding box of any target in any target image. The initial 3D bounding box is composed of the 8 vertices corresponding to two non-overlapping rectangles.

[0097] The position coordinates of any one of the eight vertices are adjusted using a preset strategy, and the 3D bounding box of the target is determined based on the adjusted vertex.

[0098] Furthermore, the loss functions of the pre-trained model include: center point heatmap loss function, vector loss function, and offset loss function.

[0099] It should be noted that the beneficial effects of the 3D target detection system provided in the above embodiments are the same as those of the vehicle 3D target detection method described above, and will not be repeated here. Furthermore, the system provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the system can be divided into different functional modules according to the actual situation to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process is detailed in the method embodiments, and will not be repeated here.

[0100] like Figure 3 As shown, an electronic device 300 according to an embodiment of the present invention includes a processor 320 coupled to a memory 310. The memory 310 stores at least one computer program 330, which is loaded and executed by the processor 320 to enable the electronic device 300 to implement any of the above-mentioned methods. Specifically:

[0101] The electronic device 300 can vary considerably due to differences in configuration or performance. It may include one or more processors 320 (Central Processing Units, CPUs) and one or more memories 310. The one or more memories 310 store at least one computer program 330, which is loaded and executed by the one or more processors 320 to enable the electronic device 300 to implement the vehicle 3D target detection method provided in the above embodiments. Of course, the electronic device 300 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The electronic device 300 may also include other components for implementing device functions, which will not be elaborated upon here.

[0102] An embodiment of the present invention provides a computer-readable storage medium storing at least one computer program, which is loaded and executed by a processor to enable a computer to implement any of the above-described methods.

[0103] Alternatively, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, a floppy disk, and an optical data storage device, etc.

[0104] In an exemplary embodiment, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform any of the methods described above.

[0105] It should be noted that the terms "first" and "second" in the specification and claims of this application are used to distinguish similar objects and represent a limitation on a specific order or sequence. Where appropriate, the order of use for similar objects can be interchanged so that the embodiments of this application described herein can be implemented in an order other than that shown or described.

[0106] Those skilled in the art will recognize that this invention can be implemented as a system, method, or computer program product. Therefore, this disclosure can be specifically implemented in the following forms: it can be entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, this invention can also be implemented as a computer program product in one or more computer-readable media containing computer-readable program code.

[0107] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.

[0108] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for 3D vehicle target detection, characterized in that, include: Acquire at least one RGB image data of any region to be detected, and preprocess all RGB image data to obtain target RGB image data in a uniform format; The target RGB image data are sequentially input into the pre-trained model to obtain the 3D target detection result corresponding to each target RGB image data. The 3D target detection result includes: the center point coordinates of any target and the 8-dimensional center point vector of any target. The target is: a motor vehicle, a non-motor vehicle, or a pedestrian. The pre-trained model is a CenterNet network, which includes a backbone network and a prediction head. The backbone network is used to extract features from the input image. The prediction head is used to predict the target center point heatmap, 8-dimensional center point vector, and center point offset corresponding to the input image based on the feature extraction results. The target center point heatmap is adjusted based on the center point offset to generate a center point heatmap. The center point heatmap and the 8-dimensional center point vector are integrated and output. The CenterNet network consists of two parts: a backbone network and a prediction head. The backbone network uses DLA34 for feature extraction from the input image. The input image has a height and width of 960×544, and the downsampling ratio R=4. After feature extraction by the backbone network, the feature map has a height and width of 240×136. The prediction head consists of three convolutional modules. Each convolutional module consists of a 3×3 convolutional kernel, a ReLU activation function, and a 1×1 convolutional kernel. These are used to predict the target center point heatmap, the 8-dimensional center point vector, and the center point offset, respectively. The center point heatmap prediction head branch has C 1×1 convolutional kernels, where C represents the number of categories in the 3D object detection dataset. The center point vector prediction head branch has 8 1×1 convolutional kernels, corresponding to the 8 dimensions of the vector. The center point offset prediction head branch has 2 1×1 convolutional kernels, used to predict the offset of the center point in the x-axis and y-axis directions during downsampling. The intersection of the lines connecting the three sets of opposite center points of the 3D bounding box is denoted as the center point of the 3D bounding box, i.e., the target center point. The 8-dimensional center point vector is denoted as... , Indicates the distance between the center point of the 3D bounding box and the left midpoint. Indicates the distance between the center point of the 3D bounding box and the right midpoint. Indicates the distance between the center point of the 3D bounding box and the top midpoint. Indicates the distance between the center point of the 3D bounding box and the midpoint below it. Indicates the distance between the center point of the 3D bounding box and the front midpoint. This represents the distance between the center point of the 3D bounding box and the center point behind it. express The angle between the x-axis and the positive x-axis. express The angle with the positive x-axis, and the loss function of the pre-trained model based on the center point. Including the center point heatmap loss function Vector loss function and offset loss function The specific calculation formula is as follows: and The hyperparameters of the loss function are set to 0.1 and 1 respectively; For the center point heatmap loss, for each real target center point p in the input image, after downsampling, an equivalent real target center point is obtained on the predicted feature map by rounding down, denoted as . The coordinates of the true target center point in the input image are denoted as p=(x,y), and R is the downsampling factor of the input image. Then, a Gaussian kernel is applied. Distribute all real target center points onto the heatmap Up, here The target scale adaptive deviation coefficient. This represents the x-axis coordinate of the equivalent true target center point on the predicted feature map. This represents the y-axis coordinate of the equivalent true target center point on the predicted feature map, and is a heatmap of the network prediction. Loss between the heatmap and the actual target center point This can be represented by focal loss: In the above formula and is the hyperparameter of focal loss, and N is the number of target center points in the image; The center point vector loss is calculated using the Smooth L1 loss formula, as follows: For the center point offset loss, This represents the difference between the predicted center point vector and the true center point vector. To reduce the target center point offset error introduced during downsampling, the model needs to predict the offset for each target center point. The prediction of offsets can improve the detection accuracy of small-sized targets. The center point offset loss during the training phase can be represented by the L1 loss. in, Where N is the actual offset, and N is the number of target center points in the image. The model represents the equivalent ground truth center point on the feature map. The predicted offset.

2. The vehicle 3D target detection method according to claim 1, characterized in that, The process of obtaining the training set for the pre-trained model is as follows: The target training region is captured by an image acquisition device to obtain at least n target images. A cube annotation tool is used to annotate any target image with a 3D bounding box to obtain the input image corresponding to each target image. All input images are integrated to obtain the training set.

3. The vehicle 3D target detection method according to claim 2, characterized in that, The process of annotating any target image with a 3D bounding box using the cube annotation tool is as follows: Determine the initial 3D bounding box of any target in any target image. The initial 3D bounding box is composed of 8 vertices corresponding to two non-overlapping rectangles. The position coordinates of any one of the eight vertices are adjusted using a preset strategy, and the 3D bounding box of the target is determined based on the adjusted vertices.

4. A 3D target detection system, employing the vehicle 3D target detection method as described in claim 1, characterized in that, include: The acquisition module is used to: acquire at least one RGB image data of any region to be detected, and preprocess all RGB image data to obtain a target RGB image in a uniform format; The detection module is used to: sequentially input the target RGB image data into the pre-trained model to obtain the 3D target detection result corresponding to each target RGB image data. The 3D target detection result includes: the center point coordinates of any target and the 8-dimensional center point vector of any target. The target is: a motor vehicle, a non-motor vehicle, or a pedestrian. The pre-trained model is a CenterNet network, which includes a backbone network and a prediction head. The backbone network is used to extract features from the input image. The prediction head is used to predict the target center point heatmap, 8-dimensional center point vector, and center point offset corresponding to the input image based on the feature extraction results. The target center point heatmap is adjusted based on the center point offset to generate a center point heatmap. The center point heatmap and the 8-dimensional center point vector are then integrated and output.

5. A 3D target detection system according to claim 4, characterized in that, The process of obtaining the training set for the pre-trained model is as follows: The target training region is captured by an image acquisition device to obtain at least n target images. A cube annotation tool is used to annotate any target image with a 3D bounding box to obtain the input image corresponding to each target image. All input images are integrated to obtain the training set.

6. A 3D target detection system according to claim 5, characterized in that, The process of annotating any target image with a 3D bounding box using the cube annotation tool is as follows: Determine the initial 3D bounding box of any target in any target image. The initial 3D bounding box is composed of 8 vertices corresponding to two non-overlapping rectangles. The position coordinates of any one of the eight vertices are adjusted using a preset strategy, and the 3D bounding box of the target is determined based on the adjusted vertices.

7. An electronic device, characterized in that, The electronic device includes a processor coupled to a memory storing at least one computer program, which is loaded and executed by the processor to enable the electronic device to perform the method as described in any one of claims 1 to 3.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to enable the computer to perform the method as claimed in any one of claims 1 to 3.