Traffic data set generation method and device based on multiple view angles and corresponding equipment

By extracting features from multi-view image sets and fusing 3D and 2D detection models, combined with automatic pre-annotation and manual fine-tuning, the problem of complex and costly data collection and annotation in existing 3D detection models is solved, and efficient and flexible 3D dataset generation is achieved.

CN121921589APending Publication Date: 2026-04-24ZHIDAO NETWORK TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHIDAO NETWORK TECH (BEIJING) CO LTD
Filing Date
2025-12-01
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

The data collection and annotation process of existing 3D detection models is complex and costly. Reliance on point cloud data results in slow annotation speed and low flexibility. It also prevents direct modification of images and increases the time required for model iteration.

Method used

Feature extraction from multi-view image sets is performed, and feature fusion is combined with 3D and 2D detection models to generate 3D pre-annotated bounding boxes. Efficient 3D annotation results are generated through automatic pre-annotation and manual fine-tuning.

Benefits of technology

It effectively reduces the cost of generating 3D datasets, improves annotation speed and flexibility, simplifies the annotation process, and enhances the efficiency of dataset generation for intelligent transportation systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921589A_ABST
    Figure CN121921589A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a multi-view-angle-based traffic data set generation method and device and corresponding equipment, and the method comprises the steps: obtaining image data of a plurality of cameras, and forming a multi-view-angle image set; performing feature extraction on the multi-view image set to obtain shared features of the multi-view image set; the 3D detection model performs 3D detection on the shared features to obtain a 3D detection result; the 2D detection model performs 2D detection on the shared features to obtain a 2D detection frame set; performing feature fusion on the 3D detection result and the 2D detection frame set to obtain a 3D pre-labeling frame; performing fine adjustment on the 3D pre-labeling frame to obtain a 3D labeling result; mapping the 3D labeling result into a 2D image of the multi-view image set to form a labeled traffic data image; the method has the beneficial effect of improving data set generation efficiency and is suitable for the field of automatic driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of autonomous driving, specifically to a method, apparatus, and corresponding equipment for generating traffic datasets based on multiple perspectives. Background Technology

[0002] In the continuous exploration and advancement of autonomous driving technology, perception capability is the core foundation of the entire system. It is like the "eyes" and "brain" of autonomous vehicles, responsible for collecting and analyzing information about the surrounding environment to provide the basis for vehicle decision-making and action. In the complex perception system, 3D detection technology occupies a pivotal position.

[0003] 3D detection technology uses optical imaging devices, LiDAR, 3D sensors, and other equipment, combined with cameras and pure vision methods, to identify objects on the road, such as vehicles, pedestrians, and traffic signs. It can accurately measure their position, shape, size, and motion, providing autonomous vehicles with richer and more accurate perception information. This is crucial for autonomous vehicles to make safe and accurate decisions in complex and ever-changing traffic environments.

[0004] In existing technologies, the construction and optimization of 3D detection models cannot be separated from high-quality data support. Compared with traditional 2D image data, the collection, annotation and processing of 3D data are more complex and costly.

[0005] On the one hand, 3D data collection requires long-term field testing in various real road environments to obtain sufficiently diverse and representative data samples. This not only requires professional testing equipment and personnel, but also needs to take into account the influence of various factors such as weather, lighting, and traffic conditions, which greatly increases the difficulty and cost of data collection.

[0006] On the other hand, because 3D data has higher dimensions and complexity, 3D data annotators need to combine complex point cloud data and image data to accurately annotate the three-dimensional position, shape and size of objects, which undoubtedly requires a lot of time and manpower.

[0007] In addition, during the debugging of 3D detection models, there is often no point cloud as a reference (especially for 3D detection models based on pure vision algorithms), which makes it impossible to correct problematic data in a timely manner, thereby increasing the time required for model iteration.

[0008] Therefore, a method and related equipment that can effectively improve the efficiency of 3D dataset generation in intelligent transportation systems are particularly important. Summary of the Invention

[0009] To address one of the aforementioned technical deficiencies, this application provides a method, apparatus, and corresponding equipment for generating traffic datasets based on multiple perspectives.

[0010] The first aspect of this application provides a method for generating traffic datasets based on multiple perspectives, comprising the following steps:

[0011] Acquire image data from multiple cameras to form a multi-view image set;

[0012] Feature extraction is performed on a multi-view image set to obtain the shared features of the multi-view image set;

[0013] The 3D detection model performs 3D detection on shared features to obtain 3D detection results;

[0014] The 2D detection model performs 2D detection on shared features to obtain a set of 2D detection boxes;

[0015] Feature fusion is performed on the 3D detection results and the 2D detection box set to obtain 3D pre-annotated boxes;

[0016] Fine-tune the 3D pre-annotation box to obtain the 3D annotation result;

[0017] The 3D annotation results are mapped onto the 2D images of the multi-view image set to form an annotated traffic data image.

[0018] In an optional embodiment of this application, the step of extracting features from the multi-view image set to obtain shared features of the multi-view image set includes:

[0019] Construct a shared feature extraction module based on a deep convolutional network;

[0020] The shared feature extraction module extracts and abstracts image features from a multi-view image set to obtain unified shared features.

[0021] In an optional embodiment of this application, the 3D detection model performs 3D detection on shared features, and the calculation expression for the 3D detection result is as follows:

[0022]

[0023] In equation (1), Represents a set of 3D detection boxes. Indicates shared features;

[0024] This indicates that shared features will be input into the 3D detection model for calculation;

[0025] X k ,Y k Z kl represents the 3D coordinates of the center point of the 3D detection box for category k. k ,w k ,h k θ represents the three-dimensional size of category k. k c represents the orientation angle of category k. k This represents the confidence level of category k within the 3D detection box, and m represents the total number of categories.

[0026] In an optional embodiment of this application, the 2D detection model performs 2D detection on shared features, and the calculation expression for the 2D detection box set is as follows:

[0027]

[0028] In formula (2): Represents a set of 2D detection boxes. This indicates that shared features will be input into the 2D detection model for calculation;

[0029] x j ,y j w represents the coordinates of the center point of the 2D detection box for category j. j ,h j Let c represent the width and height of the 2D bounding box, respectively. j represents the confidence level of category j in the 2D detection box, and m represents the total number of categories.

[0030] In an optional embodiment of this application, the step of feature fusion of the 3D detection results and the 2D detection box set to obtain 3D pre-annotated boxes includes:

[0031] For each bounding box in the 2D bounding box set, the depth is calculated to obtain a set of depth values;

[0032] A confidence-based weighted fusion strategy is used to fuse 3D detection results and depth value sets to obtain fused depth values ​​and confidence scores.

[0033] Based on the fused depth value and confidence level, a 3D pre-annotated bounding box is formed.

[0034] In an optional embodiment of this application, the step of calculating the depth of each detection box in the 2D detection box set to obtain a depth value set includes:

[0035] Obtain the intrinsic parameter matrix of a single camera and the extrinsic parameter matrix among multiple cameras;

[0036] Based on the intrinsic parameter matrix of a single camera and the extrinsic parameter matrix between multiple cameras, the planar coordinates of the target image are projected onto the local coordinates.

[0037] The depth values ​​of the target image are calculated based on its local coordinates to obtain a set of depth values;

[0038] The target image includes each detection box in the 2D detection box set.

[0039] In an optional embodiment of this application, the confidence-based weighted fusion strategy fuses the 3D detection results and the depth value set to obtain the fused depth value and confidence score, including:

[0040] Weight coefficients are constructed based on the confidence scores of the categories in the 3D detection bounding boxes and the category confidence scores in the 2D detection bounding boxes;

[0041] The fused depth value and confidence level are calculated based on the weighting coefficients, 3D detection results, and depth value set.

[0042] A second aspect of this application provides a multi-view traffic dataset generation apparatus, comprising:

[0043] The image acquisition module is used to acquire image data from multiple cameras to form a multi-view image set;

[0044] The shared feature extraction module is used to extract features from a collection of images from multiple perspectives to obtain the shared features of the collection of images from multiple perspectives.

[0045] A 3D detection model is used to perform 3D detection on shared features and obtain 3D detection results.

[0046] A 2D detection model is used to perform 2D detection on shared features to obtain a set of 2D detection boxes;

[0047] The feature fusion module is used to fuse the features of the 3D detection results and the 2D detection box set to obtain 3D pre-annotated boxes;

[0048] The adjustment module is used to fine-tune the 3D pre-annotation box to obtain the 3D annotation result;

[0049] The mapping module is used to map the 3D annotation results onto the 2D images of the multi-view image set to form an annotated traffic data image.

[0050] A third aspect of this application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the above methods.

[0051] A fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of the method as described in any of the above.

[0052] The method, apparatus, and corresponding equipment for generating traffic datasets based on multiple perspectives provided in this application embodiment are used to acquire image data from multiple cameras set up on the roadside to form a multi-perspective image set. The shared features of the multi-perspective image set are detected and fused using a 3D detection model and a 2D detection model to obtain 3D pre-annotated bounding boxes. Then, the 3D pre-annotated bounding boxes are fine-tuned, and the fine-tuned 3D annotation results are mapped onto the 2D images of the multi-perspective image set to form an annotated traffic data image.

[0053] In this application, the 3D annotation result generated by combining a multi-view image set generated from multiple cameras with automatic pre-annotation and manual fine-tuning can effectively solve the problems of high cost caused by the heavy reliance on point clouds in the data annotation of the prior art, as well as the slow annotation speed, inability to directly modify the image, low flexibility and inconvenience of manual image verification. It can effectively improve the generation efficiency of 3D datasets for intelligent transportation systems and has strong practicality.

[0054] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by way of what is pointed out in the written description, claims, and drawings. Attached Figure Description

[0055] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0056] Figure 1 This is a schematic diagram of point cloud annotation in existing technology;

[0057] Figure 2 A flowchart illustrating a multi-view traffic dataset generation method provided in one embodiment of this application;

[0058] Figure 3 This is a flowchart illustrating step S20 of a multi-view traffic dataset generation method provided in an embodiment of this application.

[0059] Figure 4 This is a flowchart illustrating step S50 in a multi-view traffic dataset generation method provided in one embodiment of this application.

[0060] Figure 5 A schematic diagram of the generated 3D pre-annotated bounding box;

[0061] Figure 6 This is a schematic diagram of the local coordinate system;

[0062] Figure 7 This is a schematic diagram of the operation interface of step S60 in the traffic dataset generation method based on multiple perspectives provided in an embodiment of this application;

[0063] Figure 8 A schematic diagram of the structure of a traffic dataset generation device based on multiple perspectives provided in one embodiment of this application;

[0064] In the picture:

[0065] 10 is the image acquisition module, 20 is the shared feature extraction module, 30 is the 3D detection model, 40 is the 2D detection model, 50 is the feature fusion module, 60 is the adjustment module, and 70 is the mapping module;

[0066] 101 represents the camera, and 102 represents a collection of images from multiple perspectives. Detailed Implementation

[0067] The solutions in this application embodiment can be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0068] To make the technical solutions and advantages of the embodiments of this application clearer, the exemplary embodiments of this application will be described in further detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not an exhaustive list of all embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.

[0069] In the process of developing this application, the inventors discovered that most existing 3D annotation schemes are based on laser point clouds, annotating 3D bounding boxes from the point cloud perspective; as shown in the attached figure. Figure 1 In the point cloud annotation diagram shown, the annotator finds the target to be annotated in the point cloud, then annotates it, and then checks it against the image. In addition, existing semi-automatic annotation methods use point cloud-based algorithms. During annotation, the algorithm is first used to pre-annotate the point cloud image, and then the annotation is corrected manually.

[0070] Among the above manual or semi-automatic annotation methods:

[0071] Data labeling heavily relies on point clouds, and point cloud data must be collected during data acquisition (if it is for road test perception model training, LiDAR needs to be installed at traffic intersections); this results in high bandwidth consumption, high memory consumption, and high cost.

[0072] Furthermore, the data transmission frequency of LiDAR is generally 10 Hz, while the transmission frequency of images is greater than 10 Hz, which makes time alignment difficult. Especially in roadside perception, cameras and radars are not triggered by hardware, and the time delay is often greater than 100 ms to 200 ms, resulting in a gap between the data labeled by point cloud and the actual image display, and data errors.

[0073] Furthermore, relying on point cloud annotation data still requires manual verification on the image, which is slow. If data problems are found during use, they cannot be directly modified based on the image, resulting in low flexibility.

[0074] To address the aforementioned issues, this application provides a method for generating traffic datasets based on multiple perspectives, such as... Figure 2 As shown: The method includes the following steps:

[0075] S10: Acquire image data from multiple cameras to form a multi-view image set;

[0076] S20, extract features from the multi-view image set to obtain the shared features of the multi-view image set;

[0077] S30, the 3D detection model performs 3D detection on shared features to obtain 3D detection results;

[0078] S40, the 2D detection model performs 2D detection on shared features to obtain a set of 2D detection boxes;

[0079] S50: Perform feature fusion on the 3D detection results and the 2D detection box set to obtain 3D pre-annotated boxes;

[0080] S60, fine-tune the 3D pre-annotation box to obtain the 3D annotation result;

[0081] S70 maps the 3D annotation results onto the 2D image of the multi-view image set to form an annotated traffic data image.

[0082] In this embodiment, the multiple cameras can be configured in various combinations according to the detection requirements, such as multiple bullet cameras or a combination of bullet camera and fisheye camera.

[0083] In this embodiment, image data from multiple cameras set up on the roadside are acquired to form a multi-view image set. The shared features of the multi-view image set are detected and fused using a 3D detection model and a 2D detection model to obtain a 3D pre-annotated bounding box. Then, the 3D pre-annotated bounding box is fine-tuned, and the fine-tuned 3D annotation result is mapped onto the 2D image of the multi-view image set to form an annotated traffic data image.

[0084] In this application, the 3D annotation result generated by combining a multi-view image set generated from multiple cameras with automatic pre-annotation and manual fine-tuning can effectively solve the problems of high cost caused by the heavy reliance on point clouds in the data annotation of the prior art, as well as the slow annotation speed, inability to directly modify the image, low flexibility and inconvenience of manual image verification. It can effectively improve the generation efficiency of 3D datasets for intelligent transportation systems and has strong practicality.

[0085] like Figure 3 As shown: S20, the step of extracting features from the multi-view image set to obtain the shared features of the multi-view image set includes:

[0086] S201, Construct a shared feature extraction module based on a deep convolutional network;

[0087] S202, through the shared feature extraction module, extracts and abstracts image features from the multi-view image set to obtain unified shared features.

[0088] Specifically, the expression for the multi-view image set can be:

[0089] I = {I1, I2, ..., I} n}; where I represents a set of multi-view images, I n This represents the image data acquired by the nth camera.

[0090] After inputting a multi-view image set I into a shared feature extraction module based on a deep convolutional network, the module extracts and abstracts features from the multi-view image set I through convolutional layers, pooling layers, and other structures, ultimately outputting the shared features. This process can be represented as:

[0091]

[0092] Furthermore, the shared features This refers to the representation in the same feature space obtained after images from different camera perspectives (front view, side view, panoramic view, etc.) are extracted by the same shared network Backbone; in the embodiments of this application, images captured by the same camera are extracted by the same shared network Backbone and mapped to the same dimension and the same scale of feature map.

[0093] The feature map is represented as [B, C, H, W], where B, C, H, and W are core abbreviations describing data dimensions, corresponding to Batch, Channel, Height, and Width, respectively.

[0094] For example, in front-view camera images and side-view camera images:

[0095] After the front-view camera image is input into the Backbone, the feature F_front∈R^(C×H×W) is obtained;

[0096] After inputting the side-view camera image into the Backbone, the feature F_side∈R^(C×H×W) is obtained;

[0097] Due to feature sharing, the two features F_front and F_side are consistent in terms of the number of channels C and the scale H×W, and can be directly concatenated, weighted and other calculations.

[0098] In this embodiment, the Backbone refers to the core convolutional network part responsible for extracting basic image features in complex tasks such as object detection and image segmentation. It serves as the "feature foundation" for subsequent tasks. Its core function is to transform the original pixel information into a feature map with semantic meaning, providing support for subsequent tasks. The shared Backbone design enables the system to fully utilize the complementary information between multiple camera perspectives while reducing the number of parameters, extracting common and different features from images from different angles, and providing a comprehensive and representative feature foundation for subsequent 2D and 3D detection.

[0099] In this embodiment, the 3D detection model and the 2D detection model do not need to train separate backbones; they can achieve the corresponding detection tasks by sharing features.

[0100] Specifically, the 3D detection model can be a BEV Transformer or other spatial projection modules.

[0101] Furthermore, in S30, the 3D detection model performs 3D detection on the shared features, and the calculation expression for the 3D detection result is as follows:

[0102]

[0103] In equation (1), Represents a set of 3D detection boxes. Indicates shared features;

[0104] This indicates that shared features will be input into the 3D detection model for calculation;

[0105] X k ,Y k Z k l represents the 3D coordinates of the center point of the 3D detection box for category k. k ,w k ,h k θ represents the three-dimensional size of category k.k c represents the orientation angle of category k. k This represents the confidence level of category k within the 3D detection box, and m represents the total number of categories.

[0106] Specifically, the category refers to the category of each element image in the target image (in this embodiment, it can be any image data in a multi-view image set), such as: motor vehicles (cars, trucks, buses, new energy vehicles, etc.), non-motor vehicles (electric vehicles, bicycles, tricycles, etc.), pedestrians (ordinary pedestrians, cyclists, children, the elderly), etc.

[0107] In this embodiment, the 3D detection model can project features from different perspectives and align them to a unified coordinate system (e.g., a bird's-eye view plane), so that information from all cameras can be fused into a global representation. The shared features provide the 3D detection model with comprehensive information from multi-view images, enabling the 3D detection model to more accurately locate targets from a bird's-eye view.

[0108] Further, in step S40, the 2D detection model performs 2D detection on the shared features, and the calculation expression for the 2D detection box set is as follows:

[0109]

[0110] In formula (2): Represents a set of 2D detection boxes. This indicates that shared features will be input into the 2D detection model for calculation;

[0111] x j ,y j w represents the coordinates of the center point of the 2D detection box for category j. j ,h j Let c represent the width and height of the 2D bounding box, respectively. j represents the confidence level of category j in the 2D detection box, and m represents the total number of categories.

[0112] In this embodiment, after integrating the shared features, the output results (2D detection box set) for each camera viewpoint demonstrate the target detection capability after multi-view image feature fusion.

[0113] like Figure 4 , Figure 5 As shown: S50 involves feature fusion of the 3D detection results and the 2D detection box set to obtain 3D pre-annotated boxes; including:

[0114] S501, Perform depth calculation on each detection box in the 2D detection box set to obtain a set of depth values;

[0115] S502, a confidence-based weighted fusion strategy, fuses 3D detection results and depth value sets to obtain fused depth values ​​and confidence scores;

[0116] S503 generates 3D pre-annotated bounding boxes based on the fused depth values ​​and confidence levels.

[0117] In this embodiment of the application, for each detection box (target image) in the set of 2D detection boxes output by the 2D detection model, its depth value can be calculated by inverse perspective transformation (IPM).

[0118] Specifically, step S501 involves calculating the depth of each detection box in the 2D detection box set to obtain a set of depth values; including:

[0119] S5011, obtain the intrinsic parameter matrix of a single camera and the extrinsic parameter matrix between multiple cameras;

[0120] S5012 projects the planar coordinates of the target image onto the local coordinates based on the intrinsic parameter matrix of a single camera and the extrinsic parameter matrix between multiple cameras.

[0121] S5013, Calculate the depth value of the target image based on the local coordinates of the target image to obtain a set of depth values;

[0122] The target image includes each detection box in the 2D detection box set.

[0123] In this embodiment of the application, the intrinsic parameter matrix of the camera and the extrinsic parameter matrix between multiple cameras can be determined using the Zhang Zhengyou calibration method; wherein, the extrinsic parameter matrix between multiple cameras refers to the transformation relationship between each camera coordinate system and the local coordinate system; the local coordinate system is defined by humans, and the origin of the vehicle end is generally the center of the rear axle of the vehicle, and the origin of the road end is generally the center of the intersection.

[0124] like Figure 6 As shown, the definition of the local coordinate system is as follows:

[0125] Origin: Center of the crossroads; x-axis: South; y-axis: East; z-axis: Up (consistent with a right-handed coordinate system).

[0126] The camera coordinate system is defined as follows:

[0127] The optical center of each camera is the origin of its coordinate system, with the x-axis pointing to the right of the camera, the y-axis pointing downwards, and the z-axis pointing in front of the lens (conforming to the camera coordinate system convention).

[0128] The transformation formula from the camera coordinate system to the local coordinate system is:

[0129]

[0130] in, This represents the extrinsic parameters of camera i relative to the local coordinate system;

[0131] X local It is a 3D point in the local coordinate system, X cam,i It is a 3D point in the camera coordinate system of camera i.

[0132] Furthermore, the core of multi-camera extrinsic calibration is to determine the rotation matrix of each camera relative to the local coordinate system using a calibration board or known 3D points in the scene. Translation vector The main steps of calibration include:

[0133] For single-camera calibration, the extrinsic parameters of each camera relative to the local coordinate system are calculated using a calibration board (such as a checkerboard) or known 3D points in the field. The positions of the corner points on the calibration board in the local coordinate system are known (x is south, y is east, and z is altitude). In this embodiment, the extrinsic parameters can be solved based on the 2D-3D correspondence using the PnP (Perspective-n-Point) algorithm.

[0134] Multi-camera geometric constraints: If multiple cameras have overlapping fields of view, the extrinsic parameters of all cameras can be optimized by using the 3D points observed together; if there are no overlapping fields of view, external equipment (such as lidar) or calibration boards placed multiple times at different locations are required to provide a global reference.

[0135] Global optimization uses observation data from all cameras to construct a global optimization problem, minimizing reprojection error and improving extrinsic parameter accuracy.

[0136] In this embodiment, assuming the camera intrinsic parameter matrix K and extrinsic parameter matrix are [R|t], both obtained through calibration, the projection relationship between the image point (u, v) and the ground plane Z = 0 is as follows:

[0137]

[0138] Where H is the homography matrix; the image point (u, v) is the coordinate of the 3D point on the 2D image plane;

[0139] For any 2D bounding box, the midpoint of the bottom edge (u c ,v c The corresponding ground coordinates (X, Y) represent the spatial location of the target, and the depth value is expressed as:

[0140]

[0141] In this embodiment of the application, step S502, based on a confidence-weighted fusion strategy, fuses the 3D detection results and the depth value set to obtain the fused depth value and confidence, including:

[0142] S5021, construct weight coefficients based on the confidence scores of the categories in the 3D detection boxes and the confidence scores of the categories in the 2D detection boxes;

[0143] S5022 calculates the fused depth value and confidence level based on the weighting coefficients, 3D detection results, and depth value set.

[0144] Specifically, the expression for the weighting coefficient is:

[0145]

[0146] Where, α k This represents the weight coefficient of category k. This represents the confidence level of category k within the 3D detection box. This represents the confidence level of category k within the 2D detection box.

[0147] The expressions for the fused depth value and confidence level are as follows:

[0148]

[0149] in, and These represent the depth and confidence level of category k after fusion, respectively. This represents the depth value of category k within the 3D detection bounding box. This represents the depth value of category k within the 3D detection box.

[0150] In this embodiment, the intrinsic parameters and extrinsic parameters of multiple cameras are obtained, and 3D pre-annotation boxes are generated by combining 2D and 3D detection models. Then, the 3D pre-annotation boxes can be fine-tuned by using an interactive interface to obtain 3D annotation results. This achieves the goal of directly modifying the 3D pre-annotation boxes in the image, thereby quickly generating high-precision 3D data.

[0151] like Figure 7 As shown: In this embodiment of the application, the 3D pre-annotated bounding box can be fine-tuned by setting a threshold. For example, if the deviation between the 3D detection box and the actual vehicle exceeds 30%, manual correction is required. Fine-tuning methods may include:

[0152] The design modification function allows users to move the center point of the 3D detection frame up, down, left, and right using the buttons, and increase / decrease the direction angle and length / width / height.

[0153] When a target image is missed, you can click on the missed location with the mouse to create a 3D box and use the modification function to modify it until the accuracy requirement is met.

[0154] In this embodiment, step S70, mapping the 3D annotation result to a 2D image of a multi-view image set, essentially involves drawing the 3D bounding box from the 3D annotation result onto the 2D image. The core of this process is projecting the eight corner points of the 3D bounding box from the three-dimensional coordinate system onto the two-dimensional coordinates of the camera image plane. The specific calculation process may include:

[0155] S701, Let the coordinates of the corner points of the 3D bounding box in the world coordinate system (or local coordinate system) be: X=(X,Y,Z);

[0156] Converted to homogeneous coordinates: Where w represents the coordinate extension; Z represents the perpendicular distance (depth) from the point to the camera's optical center;

[0157] Camera extrinsic parameters are determined by the rotation matrix. Translation vector The resulting extrinsic parameter matrix is:

[0158]

[0159] The camera's intrinsic parameter matrix is:

[0160]

[0161] Among them, f x f y c represents the focal length of the camera along the x-axis and y-axis of the image, respectively. x c y These represent the pixel coordinates of the intersection point of the camera's optical principal axis and the image plane (ideally, the center of the image).

[0162] S702 transforms each corner point of the 3D bounding box to the camera coordinate system;

[0163] Specifically, taking one of the corner points X world For example, the calculation expression for transforming it to the camera coordinate system is:

[0164] X cam =RX world +t;

[0165] Written in homogeneous coordinates:

[0166] Among them, X cam =(X c ,Y c Z c )T ;X c ,Y c Z c These represent the 3D points in the final camera coordinate system;

[0167] S703 projects 3D points from the camera coordinate system onto the image plane (pixel coordinates);

[0168] Based on the pinhole camera model, x img =K·X cam ;

[0169] Written in homogeneous coordinates:

[0170]

[0171] Where (u′,v′,w) represents the projected 2D homogeneous coordinates, which need to be further processed into pixel coordinates; the pixel coordinates (u,v) are obtained through homogeneous normalization:

[0172]

[0173] It should be understood that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order constraint on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the diagram may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0174] like Figure 8 As shown: One embodiment of this application provides a traffic dataset generation apparatus based on multiple perspectives, including:

[0175] Image acquisition module 10 is used to acquire image data through multiple cameras 101 to form a multi-view image set 102;

[0176] The shared feature extraction module 20 is used to extract features from the multi-view image set 102 to obtain the shared features of the multi-view image set;

[0177] 3D detection model 30 is used to perform 3D detection on shared features and obtain 3D detection results;

[0178] 2D detection model 40 is used to perform 2D detection on shared features to obtain a set of 2D detection boxes;

[0179] The feature fusion module 50 is used to perform feature fusion on the 3D detection results and the 2D detection box set to obtain 3D pre-annotated boxes;

[0180] The adjustment module 60 is used to fine-tune the 3D pre-annotation box to obtain the 3D annotation result;

[0181] The mapping module 70 is used to map the 3D annotation results onto the 2D images of the multi-view image set to form an annotated traffic data image.

[0182] Unlike traditional point cloud-based annotation methods, this embodiment uses cameras to acquire roadside image data and generates 3D annotation results through a combination of automatic pre-annotation and manual fine-tuning, ultimately producing an annotated traffic data image. This not only improves annotation generation efficiency but also facilitates modification of annotation results, making it easier to add, delete, modify, and query data, greatly simplifying the annotation process and improving data generation efficiency.

[0183] Specific limitations regarding the aforementioned multi-view traffic dataset generation device can be found in the limitations of the multi-view traffic dataset generation method described above, and will not be repeated here. Each module in the aforementioned multi-view traffic dataset generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0184] In one embodiment, a computer device is provided, comprising a processor, memory, a network interface, and a database connected via a system bus. The processor provides computational and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements the above-described method for generating traffic datasets based on multiple perspectives. This includes a memory and a processor; the memory stores a computer program, and the processor executes the computer program to implement any step in the above-described method for generating traffic datasets based on multiple perspectives.

[0185] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, can perform any of the steps in the above-described method for generating traffic datasets based on multiple perspectives.

[0186] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0187] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0188] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0189] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0190] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0191] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for generating traffic datasets based on multiple perspectives, characterized in that: Includes the following steps: Acquire image data from multiple cameras to form a multi-view image set; Feature extraction is performed on a multi-view image set to obtain the shared features of the multi-view image set; The 3D detection model performs 3D detection on shared features to obtain 3D detection results; The 2D detection model performs 2D detection on shared features to obtain a set of 2D detection boxes; Feature fusion is performed on the 3D detection results and the 2D detection box set to obtain 3D pre-annotated boxes; Fine-tune the 3D pre-annotation box to obtain the 3D annotation result; The 3D annotation results are mapped onto the 2D images of the multi-view image set to form an annotated traffic data image.

2. The method for generating traffic datasets based on multiple perspectives according to claim 1, characterized in that, The step of extracting features from a set of multi-view images to obtain shared features of the set includes: Construct a shared feature extraction module based on a deep convolutional network; The shared feature extraction module extracts and abstracts image features from a multi-view image set to obtain unified shared features.

3. The method for generating traffic datasets based on multiple perspectives according to claim 2, characterized in that, The 3D detection model performs 3D detection on shared features, and the calculation expression for the 3D detection result is as follows: In equation (1), Represents a set of 3D detection boxes. Indicates shared features; This indicates that shared features will be input into the 3D detection model for calculation; X k ,Y k Z k l represents the 3D coordinates of the center point of the 3D detection box for category k. k ,w k ,h k θ represents the three-dimensional size of category k. k c represents the orientation angle of category k. k This represents the confidence level of category k within the 3D detection box, and m represents the total number of categories.

4. The method for generating traffic datasets based on multiple perspectives according to claim 3, characterized in that, The 2D detection model performs 2D detection on shared features, and the calculation expression for the 2D detection box set is as follows: In formula (2): Represents a set of 2D detection boxes. This indicates that shared features will be input into the 2D detection model for calculation; x j ,y j w represents the coordinates of the center point of the 2D detection box for category j. j ,h j Let c represent the width and height of the 2D bounding box, respectively. j represents the confidence level of category j in the 2D detection box, and m represents the total number of categories.

5. The method for generating traffic datasets based on multiple perspectives according to claim 4, characterized in that, The step of fusing features between the 3D detection results and the 2D detection box set to obtain 3D pre-annotated boxes includes: For each bounding box in the 2D bounding box set, the depth is calculated to obtain a set of depth values; A confidence-based weighted fusion strategy is used to fuse 3D detection results and depth value sets to obtain fused depth values ​​and confidence scores. Based on the fused depth value and confidence level, a 3D pre-annotated bounding box is formed.

6. The method for generating traffic datasets based on multiple perspectives according to claim 5, characterized in that, The process of calculating the depth of each detection box in the 2D detection box set to obtain a set of depth values ​​includes: Obtain the intrinsic parameter matrix of a single camera and the extrinsic parameter matrix among multiple cameras; Based on the intrinsic parameter matrix of a single camera and the extrinsic parameter matrix between multiple cameras, the planar coordinates of the target image are projected onto the local coordinates. The depth values ​​of the target image are calculated based on its local coordinates to obtain a set of depth values; The target image includes each detection box in the 2D detection box set.

7. The method for generating traffic datasets based on multiple perspectives according to claim 6, characterized in that, The confidence-based weighted fusion strategy fuses the 3D detection results and the depth value set to obtain the fused depth value and confidence score, including: Weight coefficients are constructed based on the confidence scores of the categories in the 3D detection bounding boxes and the category confidence scores in the 2D detection bounding boxes; The fused depth value and confidence level are calculated based on the weighting coefficients, 3D detection results, and depth value set.

8. A traffic dataset generation device based on multiple perspectives, characterized in that, include: The image acquisition module is used to acquire image data from multiple cameras to form a multi-view image set. The shared feature extraction module is used to extract features from a collection of images from multiple perspectives to obtain the shared features of the collection of images from multiple perspectives. A 3D detection model is used to perform 3D detection on shared features and obtain 3D detection results. A 2D detection model is used to perform 2D detection on shared features to obtain a set of 2D detection boxes; The feature fusion module is used to fuse the features of the 3D detection results and the 2D detection box set to obtain 3D pre-annotated boxes; The adjustment module is used to fine-tune the 3D pre-annotation box to obtain the 3D annotation result; The mapping module is used to map the 3D annotation results onto the 2D images of the multi-view image set to form an annotated traffic data image.

9. A computer device, comprising: The method includes a memory and a processor, the memory storing a computer program, characterized in that the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.