Road object matching method, device, storage medium and electronic device
By acquiring and matching the 2D images and 3D point cloud images of the target road scene, combining historical matching information and target position conditions, the problem of inaccurate matching of road objects in the prior art is solved, improving matching accuracy and reducing computing costs.
Patent Information
- Application Number
- CN202411663304.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-20
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-11-20
AI Technical Summary
In the prior art, the matching of 2D information and 3D information of road objects is inaccurate, resulting in low positioning accuracy and high calculation cost.
By obtaining the 2D image and 3D point cloud image of the target road scene, the positional relationship of the candidate object is identified, and the object matching relationship is determined based on the historical matching information and the target position conditions.
Improves the matching accuracy of road objects, reduces calculation costs, and enhances robustness to initial values.
Smart Images

Figure CN119169323B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving, and more specifically, to a method, device, storage medium and electronic device for matching road objects. Background Art
[0002] In the field of large-scale visual positioning, direct 2D-3D matching technology has been proposed to improve positioning accuracy. This technology aims to achieve accurate positioning by directly matching pixels in 2D images with points in 3D models. Existing technologies usually start from initial pose estimation based on optimization methods and match through iterative methods, but these methods usually require a large number of optimization steps, have high computational costs, and are sensitive to initial values. In addition, current deep learning models mainly rely on fully supervised learning strategies, which have high requirements for labeled data. In practical applications, obtaining a large amount of accurately labeled data is expensive and time-consuming.
[0003] In addition, the existing technology for 3D information detection is susceptible to the loss of key visual cues (e.g., glare or blur, invisible lanes in the dark, or severe shadows), which can cause detection failure in some cases; that is, there is a technical problem in the existing technology of inaccurate matching of 2D information and 3D information of road objects. Summary of the invention
[0004] The embodiments of the present application provide a road object matching method, device, storage medium and electronic device to at least solve the technical problem of inaccurate matching of road objects in the related art.
[0005] According to one aspect of an embodiment of the present application, a road object matching method is provided, comprising: obtaining a first road object set and a second road object set that match a target road scene, wherein the first road object set is obtained by performing image recognition on a target road image corresponding to the target road scene, and the second road object set is obtained by performing point cloud recognition on a point cloud image corresponding to the target road scene; obtaining an object position relationship between a plurality of candidate objects in the first road object set and the second road object set; determining an object matching relationship between the first candidate object and the second candidate object when the object position relationship between the first candidate object of the first road object set and the second candidate object of the second road object set meets a target position condition, wherein the first candidate object and the second candidate object having the object matching relationship are used to indicate the same target road object in the target road scene; and determining a second candidate object that meets the historical matching condition from the second road object set according to historical matching information of the first candidate object when the object position relationship between the first candidate object of the first road object set and any candidate object in the second road object set does not meet the target position condition.
[0006] According to another aspect of an embodiment of the present application, a road object matching device is also provided, including: a first acquisition unit, which acquires a first road object set and a second road object set that match a target road scene, wherein the first road object set is obtained by performing image recognition on a target road image corresponding to the target road scene, and the second road object set is obtained by performing point cloud recognition on a point cloud image corresponding to the target road scene; a second acquisition unit, which acquires an object position relationship between the first road object set and a plurality of candidate objects in the second road object set; and a first determination unit, which determines, when the object position relationship between a first candidate object in the first road object set and a second candidate object in the second road object set satisfies a target position condition. , determining an object matching relationship between a first candidate object and a second candidate object, wherein the first candidate object and the second candidate object having the object matching relationship are used to indicate the same target road object in the target road scene; a second determining unit, when an object position relationship between a first candidate object in the first road object set and any candidate object in the second road object set does not meet the target position condition, determines a second candidate object that meets the historical matching condition from the second road object set according to historical matching information of the first candidate object, wherein the historical matching information includes an object matching result corresponding to at least one historical road scene, and a collection timestamp corresponding to the historical road scene is earlier than a collection timestamp of the target road scene.
[0007] As an optional solution, the above-mentioned road object matching device also includes: a third determination unit, which is used to determine the object matching relationship between the first candidate object and the second candidate object when the area of the overlapping area between the first annotation box corresponding to the first candidate object and the second annotation box of the second candidate object is greater than or equal to the area threshold; determine the object matching relationship between the first candidate object and the second candidate object when the annotation box distance between the first annotation box corresponding to the first candidate object and the second annotation box of the second candidate object is less than or equal to the distance threshold; wherein the first annotation box is used to indicate a first image area of the first candidate object in the target road image, and the second annotation box is used to indicate a second image area of the second candidate object in the target road image, and the second image area is a projection result of point cloud information based on the matching of the second candidate object on the target road image.
[0008] As an optional solution, the above-mentioned third determination unit includes: a third acquisition module, used to acquire a point cloud image collected from the target road scene; perform point cloud recognition on the point cloud image to obtain a second reference road object set, wherein the second reference road object set includes multiple reference objects; based on the point cloud information matched by each of the second reference road object sets, determine the reference annotation boxes obtained by projecting the multiple reference objects in the target road image respectively; according to the occlusion relationship of the multiple reference annotation boxes in the target road image, screen out multiple unobstructed reference annotation boxes from the multiple reference annotation boxes; and determine the second road object set according to the candidate objects indicated by each of the multiple unobstructed reference annotation boxes.
[0009] As an optional solution, the third determination unit includes: a fourth acquisition module, which is used to acquire a second reference annotation frame subset matching the first annotation frame from a second reference annotation frame set matching the second road object set when the overlapping areas between the first annotation frame and the respective candidate annotation frames of any candidate object in the second road object set are all smaller than the area threshold, wherein the area of the overlapping areas between the second reference annotation frame in the second reference annotation frame subset and the first annotation frame is greater than 0 and smaller than the area threshold; determine a target annotation frame from the second reference annotation frame subset, wherein the area of the overlapping area between the first annotation frame and the target annotation frame is greater than the area of the overlapping areas corresponding to the other annotation frames in the second reference annotation frame subset; and determine that the third candidate object indicated by the target annotation frame has an object matching relationship with the first road object when the annotation frame distance between the target annotation frame and the first annotation frame meets the target distance condition.
[0010] As an optional solution, the fourth acquisition module includes: a fifth acquisition module, used to obtain the target annotation box and the annotation box distance between multiple first reference annotation boxes in the first reference annotation box set, wherein the first reference annotation box is used to indicate the image area of the candidate object in the first road object set in the target road image, and the multiple first reference annotation boxes include the first annotation box; when the annotation box distance between the target annotation box and the first annotation box is smaller than the annotation box distance corresponding to each of the other annotation boxes in the first reference annotation box set, determine that the annotation box distance between the target annotation box and the first annotation box meets the target distance condition.
[0011] As an optional scheme, the above-mentioned second determination unit also includes a sixth acquisition module, which is used to obtain historical matching information, wherein the historical matching information includes object matching results of the first candidate object in multiple historical road scenes, the multiple historical road scenes and the target road scene are road scenes corresponding to multiple information collection operations respectively, the multiple information collection operations are information collection operations corresponding to multiple continuous timestamps respectively, the information collection operations include image data collection operations and point cloud data collection operations, and the object matching results are used to indicate a second reference object that matches the first candidate object in the current road scene; according to the object matching results corresponding to each of the multiple historical road scenes, a second candidate object that meets the historical matching conditions is determined from at least one second reference object.
[0012] As an optional scheme, the above-mentioned sixth acquisition module includes: a fourth determination module, used to determine the second reference object with the largest number of matches as the second candidate object that meets the historical matching conditions; obtain the historical matching information of the second reference object with the largest number of matches; and determine the second reference object as the second candidate object that meets the historical matching conditions based on the historical matching information of the second reference object.
[0013] As an optional scheme, the above-mentioned fourth determination module is also used to determine that the second reference object does not meet the historical matching conditions when the historical matching information of the second reference object indicates that the candidate object with the most matches with the second reference object is not the first candidate object; determine the object acquisition order according to the number of matches corresponding to at least one second reference object; acquire a second reference object from at least one second reference object in sequence as the current reference object according to the object acquisition order; and determine the current reference object as the second candidate object that meets the historical matching conditions when the historical matching information of the current reference object indicates that the first candidate object is the candidate object with the most matches with the current reference object.
[0014] As an optional scheme, the above-mentioned fourth determination module is also used to obtain a second annotation box corresponding to a second candidate object that meets the historical matching conditions, and a second verification annotation box corresponding to a second verification candidate object that matches the first candidate object determined according to the key frame matching information; when the area of the overlapping area between the second verification annotation box and the second annotation box is greater than or equal to the area threshold, determine that the matching result is correct.
[0015] As an optional scheme, the above-mentioned fourth determination module is also used to determine the second candidate object that matches the first candidate object the least number of times in the matching results as the second candidate object to be corrected that meets the verification conditions, wherein the matching results include multiple first candidate objects and second candidate objects with a matching relationship; obtain the key frame that matches the second candidate object to be corrected, and the key frame matching information corresponding to the key frame; and modify the second candidate object to be corrected that matches the first candidate object to be corrected into the second correction object according to the key frame matching information.
[0016] According to another aspect of the embodiment of the present application, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above road object matching method.
[0017] According to another aspect of an embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the road object matching method through the computer program.
[0018] In the above-mentioned embodiment of the present application, a first road object set and a second road object set matching the target road scene are obtained, wherein the first road object set is obtained by performing image recognition on a target road image corresponding to the target road scene, and the second road object set is obtained by performing point cloud recognition on a point cloud image corresponding to the target road scene; an object position relationship between a plurality of candidate objects in the first road object set and the second road object set is obtained; when the object position relationship between a first candidate object in the first road object set and a second candidate object in the second road object set satisfies a target position condition, an object matching relationship between the first candidate object and the second candidate object is determined, wherein the first candidate object and the second candidate object having the object matching relationship are used to indicate the same target road object in the target road scene; when the object position relationship between the first candidate object in the first road object set and any candidate object in the second road object set does not satisfy the target position condition, a second candidate object satisfying the historical matching condition is determined from the second road object set according to historical matching information of the first candidate object, wherein the historical matching information includes an object matching result corresponding to each of at least one historical road scene, and the acquisition timestamp corresponding to the historical road scene is earlier than the acquisition timestamp of the target road scene.
[0019] Through the above-mentioned implementation mode of the present application, a first road object set and a second road object set matching the target road scene are obtained; then, the object position relationship between the first road object set and multiple candidate objects in the second road object set is obtained; it is achieved that when the object position relationship between the first candidate object in the first road object set and the second candidate object in the second road object set meets the target position condition, the object matching relationship between the first candidate object and the second candidate object is determined, wherein the first candidate object and the second candidate object having the object matching relationship are used to indicate the same target road object in the target road scene; when the object position relationship between the first candidate object in the first road object set and any candidate object in the second road object set does not meet the target position condition, the second candidate object that meets the historical matching condition is determined from the second road object set based on the historical matching information of the first candidate object, thereby solving the technical problem of inaccurate matching of 2D information and 3D information of road objects in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0021] Figure 1is a schematic diagram of an application environment of an optional road object matching method according to an embodiment of the present application;
[0022] Figure 2 is a flow chart of an optional road object matching method according to an embodiment of the present application;
[0023] Figure 3 is a schematic diagram of an optional road object matching method according to an embodiment of the present application;
[0024] Figure 4 is a schematic diagram of another optional road object matching method according to an embodiment of the present application;
[0025] Figure 5 is a schematic diagram of another optional road object matching method according to an embodiment of the present application;
[0026] Figure 6 is a schematic diagram of another optional road object matching method according to an embodiment of the present application;
[0027] Figure 7 is a schematic diagram of another optional road object matching method according to an embodiment of the present application;
[0028] Figure 8 is a flowchart of another optional road object matching method according to an embodiment of the present application;
[0029] Fig. 9 is a schematic diagram of a road object matching device according to an embodiment of the present application;
[0030] Fig.10 It is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0031] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0033] According to one aspect of an embodiment of the present application, a road object matching method is provided. Optionally, the road object matching method can be but is not limited to being applied to: Figure 1 In the hardware environment shown. Optionally, the road object matching method provided by the present application can be applied to a vehicle terminal. Figure 1 A side view of a vehicle terminal 101 is shown, which can travel on a travel surface 113. The vehicle terminal 101 includes an onboard navigation system 103, a memory 102 storing a digitized road map 104, a space monitoring system 117, a vehicle controller 109, a GPS (Global Positioning System) sensor 110, an HMI (Human / Machine Interface) device 111, and also includes an autonomous controller 112 and a telematics controller 114.
[0034] In one embodiment, the space monitoring system 117 includes one or more space sensors and systems, which are used to monitor the visible area 105 in front of the vehicle terminal 101. The space monitoring system 117 also includes a space monitoring controller 118; the space sensors used to monitor the visible area 105 include a laser radar sensor 106, a radar sensor 107, a camera 108, etc. The space monitoring controller 118 can be used to generate data related to the visible area 105 based on the data input from the space sensor. The space monitoring controller 118 can determine the linear range, relative speed and trajectory of the vehicle terminal 101 based on the input from the space sensor, for example, determine the current speed of the vehicle and the relative speed compared to the front vehicle. The space sensor of the vehicle terminal space monitoring system 117 may include an object positioning sensing device, and the object positioning sensing device may include a range sensor, which can be used to locate the front object such as the front vehicle object.
[0035] The camera 108 is advantageously mounted and positioned on the vehicle terminal 101 in a position that allows capturing an image of the visible area 105, wherein at least a portion of the visible area 105 includes a portion of the travel surface 113 in front of the vehicle terminal 101 and including the trajectory of the vehicle terminal 101. The visible area 105 may also include the surrounding environment. Other cameras may also be used, for example, including a second camera disposed on the rear or side portion of the vehicle terminal 101 to monitor the rear of the vehicle terminal 101 and one of the right or left sides of the vehicle terminal 101.
[0036] The autonomous controller 112 is configured to implement autonomous driving or advanced driver assistance system (ADAS) vehicle terminal functionality. Such functionality may include a vehicle terminal onboard control system capable of providing a certain level of driving automation. Driving automation may include a range of dynamic driving and vehicle terminal operations. Driving automation may include a certain level of automatic control or intervention involving a single vehicle terminal function (e.g., steering, acceleration, and / or braking). For example, the above-mentioned autonomous controller can be used to match scenes of road objects by performing the following steps:
[0037] S102, acquiring a first road object set and a second road object set matching the target road scene, wherein the first road object set is obtained by performing image recognition on a target road image corresponding to the target road scene, and the second road object set is obtained by performing point cloud recognition on a point cloud image corresponding to the target road scene;
[0038] S104, obtaining an object position relationship between a plurality of candidate objects in the first road object set and the second road object set;
[0039] S106, when the object position relationship between the first candidate object in the first road object set and the second candidate object in the second road object set meets the target position condition, determining an object matching relationship between the first candidate object and the second candidate object, wherein the first candidate object and the second candidate object having the object matching relationship are used to indicate the same target road object in the target road scene;
[0040] S108. When the object position relationship between the first candidate object in the first road object set and any candidate object in the second road object set does not satisfy the target position condition, a second candidate object satisfying the historical matching condition is determined from the second road object set based on historical matching information of the first candidate object, wherein the historical matching information includes an object matching result corresponding to at least one historical road scene, and a collection timestamp corresponding to the historical road scene is earlier than a collection timestamp of the target road scene.
[0041] The HMI device 111 provides human-machine interaction for the purpose of guiding the operation of the infotainment system, GPS (Global Positioning System) sensor 110, onboard navigation system 103 and the like, and includes a controller. The HMI device 111 monitors operator requests and provides the operator with status, service and maintenance information of the vehicle terminal system. The HMI device 111 communicates with multiple operator interface devices and / or controls the operation of multiple operator interface devices. The HMI device 111 can also communicate with one or more devices that monitor biometric data associated with the vehicle terminal operator. For simplicity of description, the HMI device 111 is depicted as a single device, but in the embodiments of the system described herein, it can be configured as multiple controllers and associated sensing devices.
[0042] Operator controls may be included in the passenger compartment of the vehicle terminal 101, and may include, by way of non-limiting example, a steering wheel, an accelerator pedal, a brake pedal, and an operator input device, which is an element of the HMI device 111. The operator controls enable a vehicle terminal operator to interact with the operating vehicle terminal 101 and direct the operation of the vehicle terminal 101 to provide passenger transportation.
[0043] The onboard navigation system 103 uses the digitized road map 104 for the purpose of providing navigation support and information to the vehicle terminal operator. The autonomous controller 112 uses the digitized road map 104 for the purpose of controlling the autonomous vehicle terminal operation or ADAS vehicle terminal functions.
[0044] The vehicle terminal 101 may include a telematics controller 114, which includes a wireless telematics communication system capable of communicating outside the vehicle terminal (including communicating with a communication network 115 having wireless and wired communication capabilities). The wireless telematics communication system includes an off-board server 116 capable of short-range wireless communication with a mobile terminal.
[0045] As an optional implementation, Figure 2 As shown, the road object matching method can be executed by an electronic device, and the specific steps include:
[0046] S202, acquiring a first road object set and a second road object set matching the target road scene, wherein the first road object set is obtained by performing image recognition on a target road image corresponding to the target road scene, and the second road object set is obtained by performing point cloud recognition on a point cloud image corresponding to the target road scene;
[0047] S204, obtaining an object position relationship between a plurality of candidate objects in the first road object set and the second road object set;
[0048] S206, determining an object matching relationship between the first candidate object in the first road object set and the second candidate object in the second road object set when the object position relationship between the first candidate object and the second candidate object meets the target position condition, wherein the first candidate object and the second candidate object having the object matching relationship are used to indicate the same target road object in the target road scene;
[0049] S208. When the object position relationship between the first candidate object in the first road object set and any candidate object in the second road object set does not satisfy the target position condition, a second candidate object satisfying the historical matching condition is determined from the second road object set based on historical matching information of the first candidate object, wherein the historical matching information includes an object matching result corresponding to at least one historical road scene, and a collection timestamp corresponding to the historical road scene is earlier than a collection timestamp of the target road scene.
[0050] In S202 of the above embodiment, a first road object set and a second road object set matching the target road scene are obtained, wherein the first road object set is obtained by performing image recognition on a target road image corresponding to the target road scene, and the second road object set is obtained by performing point cloud recognition on a point cloud image corresponding to the target road scene;
[0051] As an optional implementation, for example, the target road scene is a crossroad scene; the first road object set can be obtained by performing image recognition on a target road image corresponding to the crossroad scene, for example, obtaining a 2D true value information set of multiple road objects (such as cars, buses, trains, etc.) at the crossroad;
[0052] Optionally, the onboard camera of the second vehicle in the first lane captures the road scene image of the current intersection, and further performs feature extraction to convert the image into a high-dimensional feature representation. Then, an object detection algorithm (such as FasterR-CNN, YOLO, etc.) is used to detect and locate the object on the feature map, determine whether each candidate area contains the object, and accurately locate the object position. These algorithms can identify and locate the vehicle in the image and output the 2D bounding box of the vehicle.
[0053] The second road object set may be obtained by performing point cloud recognition on a point cloud image corresponding to the target road scene, for example, by acquiring 3D point cloud information of multiple road objects (such as cars, buses, trains, etc.) at the intersection through, but not limited to, laser scanning, and then obtaining a 3D true value information set;
[0054] In S204 in the above embodiment, the object position relationship between the plurality of candidate objects in the first road object set and the second road object set is obtained.
[0055] Optionally, for example, 3D information corresponding to the red car in the second road object set is obtained, and the 3D information is projected onto the 2D image. The object position relationship between the first road object set and the multiple candidate objects in the second road object set can be determined by projection, that is, the matching relationship between the 2D true value information and the 3D true value information is preliminarily obtained. The coordinates of the 3D point can be converted from the world coordinate system to the camera coordinate system by, but not limited to, determining the internal parameters (such as focal length, principal point) and external parameters (such as position and rotation relative to the world coordinate system) of the camera, and then applying the internal parameters of the camera to project the 3D coordinates onto the 2D plane; the above position relationship includes, but is not limited to: unobstructed, partially obstructed, fully obstructed, same position, different positions, etc.
[0056] In S206 in the above embodiment, when the object position relationship between the first candidate object in the first road object set and the second candidate object in the second road object set satisfies the target position condition, the object matching relationship between the first candidate object and the second candidate object is determined, wherein the first candidate object and the second candidate object having the object matching relationship are used to indicate the same target road object in the target road scene.
[0057] It can be understood that the target position condition may be that the first candidate object of the first road object set and the second candidate object of the second road object set are at the same position and the corresponding 2D information and 3D information respectively;
[0058] Preliminarily, candidate objects are selected from the first road object set and the second road object set respectively according to the vehicle type, for example, cars are selected as candidate objects, and then the 2D information of all cars in the first road object set and the 3D information of all cars in the second road object set are obtained. The above object matching relationship is used to indicate whether the 2D information and the 3D information are matched successfully. If the match is successful, the first candidate object and the second candidate object are used to indicate the same target road object in the target road scene.
[0059] Optionally, to determine whether the object position relationship between the first candidate object in the first road object set and the second candidate object in the second road object set meets the target position condition, it is possible to calculate whether the numerical value of the bounding_box_overlap of the 2D true value box and the 3D projection box is greater than a set threshold. Specifically, the IoU of the 2D true value box and the 3D projection box is calculated. The IoU is an indicator for measuring the degree of overlap of two boxes on a 2D image. When the IoU is greater than a certain threshold, it should be noted that the above threshold can be dynamically determined according to the complexity of the actual scene, and it is directly determined that the 2D information and the 3D information correspond to the same road object, and the matching is successful. This is only an example.
[0060] The corresponding 3D information can also be reconstructed based on the 2D information in the first road object set and further matched with the 3D information in the second road object set, which is not specifically limited here.
[0061] In S208 of the above embodiment, when the object position relationship between the first candidate object of the first road object set and any candidate object in the second road object set does not meet the target position condition, a second candidate object that meets the historical matching condition is determined from the second road object set based on the historical matching information of the first candidate object, wherein the historical matching information includes an object matching result corresponding to at least one historical road scene, and the acquisition timestamp corresponding to the historical road scene is earlier than the acquisition timestamp of the target road scene.
[0062] It can be understood that, when the 2D truth box and the 3D projection box do not meet the target position condition (for example, no overlap), a second candidate object that meets the historical matching condition is further determined from the second road object set based on the historical matching information of the first candidate object. For example, the above-mentioned first candidate object is the first car in the left lane, and the historical matching information matching the car is obtained. For example, the historical matching information indicates that the 3D information corresponding to the vehicle ID 3 has the largest number of successful matches with the car, then the 3D information corresponding to the vehicle ID 3 in the second road object set is determined to be the second candidate object that meets the historical matching condition.
[0063] Through the above-mentioned implementation mode of the present application, a first road object set and a second road object set matching the target road scene are obtained; then, the object position relationship between the first road object set and multiple candidate objects in the second road object set is obtained; it is achieved that when the object position relationship between the first candidate object in the first road object set and the second candidate object in the second road object set meets the target position condition, the object matching relationship between the first candidate object and the second candidate object is determined, wherein the first candidate object and the second candidate object having the object matching relationship are used to indicate the same target road object in the target road scene; when the object position relationship between the first candidate object in the first road object set and any candidate object in the second road object set does not meet the target position condition, the second candidate object that meets the historical matching condition is determined from the second road object set based on the historical matching information of the first candidate object, thereby solving the technical problem of inaccurate matching of 2D information and 3D information of road objects in the prior art.
[0064] In an optional implementation, when the object position relationship between the first candidate object in the first road object set and the second candidate object in the second road object set meets the target position condition, determining the object matching relationship between the first candidate object and the second candidate object includes at least one of the following:
[0065] Method 1: determining an object matching relationship between the first candidate object and the second candidate object when the area of an overlapping region between a first annotation box corresponding to the first candidate object and a second annotation box of the second candidate object is greater than or equal to an area threshold;
[0066] Method 2: determining an object matching relationship between the first candidate object and the second candidate object when the distance between the first marked box corresponding to the first candidate object and the second marked box of the second candidate object is less than or equal to the distance threshold;
[0067] Among them, the first annotation box is used to indicate the first image area of the first candidate object in the target road image, and the second annotation box is used to indicate the second image area of the second candidate object in the target road image, and the second image area is the projection result of the point cloud information matched based on the second candidate object in the target road image.
[0068] In the first method, when the area of the overlapping region between the first annotation box corresponding to the first candidate object and the second annotation box of the second candidate object is greater than or equal to the area threshold, determining the object matching relationship between the first candidate object and the second candidate object;
[0069] Matching can be performed based on IoU, by calculating whether the value of the bounding_box_overlap (area of the overlapping area) of the 2D true value box and the 3D projection box is greater than the set threshold. If the IoU of the 2D true value box and the 3D projection box of the same category is greater than the set threshold, the 2D true value box and the 3D projection box will be associated to determine that the object matching relationship between the first candidate object and the second candidate object is successful.
[0070] The area threshold may be manually set based on experience, or may be determined based on the product of the area of the annotation box and a fixed ratio.
[0071] Optionally, in the above-mentioned second method, when the distance between the first annotation box corresponding to the first candidate object and the second annotation box of the second candidate object is less than or equal to the distance threshold, the object matching relationship between the first candidate object and the second candidate object is determined;
[0072] The above-mentioned annotation box distance can be but is not limited to the distance between the center point of the first annotation box and the center point of the second annotation box. When the distance between the center point of the 2D true value box and the center point of the 3D projection box is less than a certain threshold, the object matching relationship between the first candidate object and the second candidate object is determined to be a successful match.
[0073] In another optional implementation, the object matching relationship between the first candidate object and the second candidate object is jointly determined in combination with the first method and the second method. For example, the area of the 2D true value frame is the smallest, and the 2D true value frame matching the first candidate object and the 3D projection frame matching the second candidate object are obtained. The overlapping area between the 2D true value frame and the 3D projection frame is greater than 50% of the area of the 2D true value frame and less than 90% of the 2D true value frame (i.e., partial overlap). It is further determined whether the distance between the center point of the 2D true value frame and the center point of the 3D projection frame is less than 1 mm. If the distance between the center point of the 2D true value frame and the center point of the 3D projection frame is less than 1 mm, the object matching relationship between the first candidate object and the second candidate object is determined to be a successful match.
[0074] It should be noted that, to determine whether the distance between the center point of the 2D true value box and the center point of the 3D projection box is less than a certain threshold, it is also possible to obtain the distance between the center point of the 2D true value box matching the first candidate object and the center point of the 3D projection box matching all the second candidate objects, and determine that the distance between the center point of the first candidate object and the center point of the target second candidate object is the smallest of all distances. In this case, it can be determined that the object matching relationship between the first candidate object and the target second candidate object is a successful match, and no specific restrictions are made here.
[0075] Through the above-mentioned implementation mode of the present application, judging the matching road objects by the overlapping conditions and the center point distance conditions can significantly improve the matching accuracy, efficiency and robustness, while improving the consideration of spatial relationships and optimizing the entire matching process.
[0076] In an optional implementation, before obtaining the object position relationship between the first road object set and the plurality of candidate objects in the second road object set, the method includes:
[0077] S1, obtaining a point cloud image collected from a target road scene;
[0078] S2, performing point cloud recognition on the point cloud image to obtain a second reference road object set, wherein the second reference road object set includes a plurality of reference objects;
[0079] S3, determining reference annotation boxes respectively projected by the plurality of reference objects in the target road image based on the point cloud information respectively matched by the second reference road object set;
[0080] S4, selecting a plurality of unoccluded reference annotation frames from the plurality of reference annotation frames according to an occlusion relationship between the plurality of reference annotation frames in the target road image;
[0081] S5: Determine a second road object set according to the candidate objects indicated by each of the multiple reference annotation boxes that are not blocked.
[0082] It can be understood that in the above steps S1-S3, a point cloud image obtained by collecting the target road scene is obtained; point cloud recognition is performed on the point cloud image to obtain a second reference road object set, wherein the second reference road object set includes multiple reference objects, such as 3D information of multiple cars; based on the point cloud information matched by each of the second reference road object sets, the reference annotation frames obtained by projecting the multiple reference objects in the target road image are determined, that is, the corresponding multiple 3D projection frames are projected according to the 3D information corresponding to each car; for example, based on the camera intrinsic parameters and the joint calibration parameters of the camera and the lidar, the 3D target truth frame is projected to the 2D image. This step projects the 3D frame to the 2D image.
[0083] In the above steps S4-S5, multiple unobstructed reference annotation frames are screened out from the multiple reference annotation frames according to the occlusion relationship between the multiple reference annotation frames in the target road image; and the second road object set is determined according to the candidate objects indicated by each of the multiple unobstructed reference annotation frames.
[0084] The above process is described below in a specific implementation manner:
[0085] The 2D box projected from the 3D target truth box onto the 2D image is screened. Due to the characteristics of point cloud, the cloud-based truth model can output targets that are completely invisible in the 2D image based on point cloud data. Here, those truth targets that are completely invisible in the 2D image need to be eliminated.
[0086] Specifically, this scheme sets three discrimination conditions. According to the fact that the target truth value is larger when it is near and smaller when it is far away, and although the 3D truth value frame is projected, the depth information of the 3D truth value frame is still retained, the size, distance, and relative relationship of the center pixel coordinates of the 2D projection frame can be used to eliminate the true value target 3D projection frames that are completely invisible from the 2D image perspective.
[0087] like Figure 3 As shown, taking the front view image as an example, the dotted box represents the front view image, the solid line truth box on the left is the 3D truth box corresponding to the vehicle object 302 projected onto the 2D image, and the dotted line truth box on the left is the 3D truth box corresponding to the vehicle object 304 projected onto the 2D image. The solid line truth box on the left surrounds the dotted line truth box on the left, and the vehicle object 304 is farther away from the vehicle 308 longitudinally, so the dotted line truth box on the left can be eliminated because this object cannot be seen under the front view camera.
[0088] Through the above-mentioned implementation mode recorded in the present application, the possibility of false matching is reduced by eliminating the truth value boxes projected onto the 2D image from the 3D truth value boxes that are not visible from the camera perspective. Since these occluded targets are not visible in the 2D image, if they are included in the matching process, the matching algorithm may mistakenly match these invisible targets with visible targets, thereby reducing the accuracy of the matching.
[0089] In an optional implementation, when the object position relationship between the first candidate object in the first road object set and the second candidate object in the second road object set meets the target position condition, determining the object matching relationship between the first candidate object and the second candidate object further includes:
[0090] S1, when the overlapping areas between the first annotation box and the candidate annotation boxes of any candidate object in the second road object set are all smaller than the area threshold, a second reference annotation box subset matching the first annotation box is obtained from the second reference annotation box set matching the second road object set, wherein the overlapping area between the second reference annotation box in the second reference annotation box subset and the first annotation box is greater than 0 and smaller than the area threshold;
[0091] S2, determining a target annotation frame from the second reference annotation frame subset, wherein the area of an overlapping region between the first annotation frame and the target annotation frame is greater than the area of an overlapping region corresponding to other annotation frames in the second reference annotation frame subset;
[0092] S3: When the distance between the target annotation box and the first annotation box meets the target distance condition, determine that the third candidate object indicated by the target annotation box has an object matching relationship with the first road object.
[0093] It can be understood that, in the above step S1, when the overlapping areas between the first annotation frame and the candidate annotation frames of any candidate object in the second road object set are all smaller than the area threshold, a second reference annotation frame subset matching the first annotation frame is obtained from the second reference annotation frame set matching the second road object set, wherein the area of the overlapping areas between the second reference annotation frames in the second reference annotation frame subset and the first annotation frame is greater than 0 and smaller than the area threshold;
[0094] For example, the first annotation box is the 2D true value box corresponding to the first car on the first lane on the left, and the overlapping area between the candidate annotation box (3D projection box) of any candidate object in the second road object set is smaller than the area threshold (for example, equal to the area size of the 2D box of the car), then a second reference annotation box subset matching the first annotation box is obtained. It should be noted that the area of the overlapping area between the second reference annotation box in the second reference annotation box subset and the first annotation box is greater than 0 and smaller than the area threshold. For example, the overlapping area between the second reference annotation box in the second reference annotation box subset and the first annotation box is 50%-80% of the area of the 2D box of the car.
[0095] In steps S2-S3, a target annotation box is determined from the second reference annotation box subset, and the area of the overlapping region between the first annotation box and the target annotation box is greater than the area of the overlapping regions corresponding to other annotation boxes in the second reference annotation box subset; when the annotation box distance between the target annotation box and the first annotation box meets the target distance condition, it is determined that the third candidate object indicated by the target annotation box has an object matching relationship with the first road object.
[0096] As an optional implementation, the IoU values in S1 that are not 0 with the 2D true value box are arranged from large to small, and the second reference annotation box subset with the largest overlapping area with the first annotation box is determined as the target annotation box, and the Euclidean distance of the center point pixel coordinates of the 2D true value box (first annotation box) and all second reference annotation box subsets (3D projection boxes) are calculated and sorted, and further matched in combination with the IoU sorting result and the center point Euclidean distance sorting result;
[0097] For example, the 3D information corresponding to the target annotation box with the largest overlapping area with the first annotation box in the second reference annotation box subset is then determined to determine whether the distance between the center point of the first annotation box and the center point of the reference annotation box with index 3 is less than 1 mm. If the target distance condition is met, it is determined that the first candidate object corresponding to the first annotation box and the second candidate object corresponding to the reference annotation box with index 3 are successfully matched, that is, they belong to the same road object.
[0098] Through the above implementation of the present application, when the overlapping areas between the first annotation frame and the respective candidate annotation frames of any candidate object in the second road object set are all smaller than the area threshold, a second reference annotation frame subset matching the first annotation frame is obtained from the second reference annotation frame set matching the second road object set; and then a target annotation frame is determined from the second reference annotation frame subset, so that when the annotation frame distance between the target annotation frame and the first annotation frame satisfies the target distance condition, it is determined that the third candidate object indicated by the target annotation frame has an object matching relationship with the first road object, thereby realizing the dual constraint judgment of the matching condition, and thus improving the accuracy of the matching result.
[0099] In an optional implementation, after determining the target annotation frame from the second reference annotation frame subset, the method further includes:
[0100] S1, obtaining a target annotation box and an annotation box distance between a plurality of first reference annotation boxes in a first reference annotation box set, wherein the first reference annotation box is used to indicate an image area of a candidate object in a first road object set in a target road image, and the plurality of first reference annotation boxes include the first annotation box;
[0101] S2: When the annotation box distances between the target annotation box and the first annotation box are smaller than the annotation box distances corresponding to other annotation boxes in the first reference annotation box set, determine the annotation box distance between the target annotation box and the first annotation box to meet the target distance condition.
[0102] As an optional implementation, for example, the distance between the center point of the target annotation box with an index of 3 and the center points of multiple first reference annotation boxes including the first annotation box is obtained, and a double determination of the two types of sorting results is performed. For example, the index of the 2D true value box (first annotation box) with an index of 1 and the 3D true value projection box (target annotation box) with the largest IoU value is 3;
[0103] Moreover, the 2D truth frame indexed as 1 is the closest to the Euclidean distance of the center pixel coordinates of the 3D truth projection frame indexed as 3, which is smaller than the distance between the center point of the target annotation frame and the center points of other first reference annotation frames, i.e., it satisfies the dual judgment of maximum overlap and minimum distance between center points. The 2D truth frame indexed as 1 and the 3D truth projection frame indexed as 3 are associated and matched, thereby improving the matching accuracy of road objects.
[0104] In an optional implementation, determining a second candidate object that meets the historical matching condition from the second road object set according to the historical matching information of the first candidate object includes:
[0105] S1, obtaining historical matching information, wherein the historical matching information includes object matching results of a first candidate object in a plurality of historical road scenes, the plurality of historical road scenes and the target road scene are road scenes corresponding to a plurality of information collection operations respectively, the plurality of information collection operations are information collection operations corresponding to a plurality of continuous timestamps respectively, the information collection operations include image data collection operations and point cloud data collection operations, and the object matching result is used to indicate a second reference object that matches the first candidate object in a current road scene;
[0106] S2: Determine, according to the object matching results corresponding to each of the plurality of historical road scenes, a second candidate object that meets the historical matching condition from at least one second reference object.
[0107] It should be noted that in step S1, the historical matching information includes the object matching results of the first candidate object in multiple historical road scenes, the multiple historical road scenes and the target road scene are road scenes corresponding to multiple information collection operations respectively, the multiple information collection operations are information collection operations corresponding to multiple continuous timestamps respectively, the information collection operations include image data collection operations and point cloud data collection operations, and the object matching results are used to indicate the second reference object that matches the first candidate object in the current road scene.
[0108] Optionally, the historical matching information includes multiple corresponding 2D acquisition information and 3D acquisition information of the first candidate object at multiple consecutive timestamps of the intersection scene, including under different weather conditions, and matching results of the 2D acquisition information and 3D acquisition information of the first candidate object.
[0109] Then, in step S2, according to the object matching results corresponding to the plurality of historical road scenes, a second candidate object satisfying the historical matching condition is determined from at least one second reference object.
[0110] For example, if the second candidate object that matches the first candidate object the most times from the historical matching information is the car target with index 3, then the information of the second candidate object that matches the first candidate object is determined to be the 3D information corresponding to the car target with index 3 according to the historical matching information.
[0111] By acquiring historical matching information and then determining a second candidate object that meets the historical matching conditions from at least one second reference object based on the object matching results corresponding to each of a plurality of historical road scenes, the optimal matching result between the 2D information and the 3D information is guaranteed to the greatest extent, thereby improving the accuracy of matching the road object information.
[0112] In an optional implementation, according to the object matching results corresponding to each of the plurality of historical road scenes, determining a second candidate object satisfying the historical matching condition from at least one second reference object includes one of the following:
[0113] Method 1: determine the second reference object with the most matching times as the second candidate object that meets the historical matching condition;
[0114] Method 2: Obtain historical matching information of a second reference object having the largest number of matches; and determine, based on the historical matching information of the second reference object, that the second reference object is a second candidate object that meets the historical matching condition.
[0115] As an optional implementation, for example, the Car target with 2D true value ID 5 has matched with the Car target with 3D projection ID 3 5 times, matched with the Car target with 3D projection ID 6 4 times, and matched with the Car target with 3D projection ID 7 once in 10 consecutive historical frames. We sort the historical matching results in descending order. From the descending results, it can be seen that the Car target with 2D true value ID 5 has matched with the Car target with 3D projection ID 3 the most times in the historical matching stage, so the Car target with 3D projection ID 3 is determined to be the second candidate object that meets the historical matching conditions.
[0116] As another optional implementation, for example, the Car target with a 2D true value ID of 5 matches the Car target with a 3D projection ID of 3 the most times in 10 consecutive historical frames, and further obtains the historical matching information of the Car target with a 3D projection ID of 3 (the second reference object), and searches for the 2D true value ID that matches the Car target with a 3D projection ID of 3 the most times in the historical matching results. If the 2D true value ID is 5, the second reference object is determined to be the second candidate object that meets the historical matching conditions. This is only an example and is not specifically limited.
[0117] In an optional implementation manner, after acquiring the historical matching information of the second reference object with the largest number of matches, the method further includes:
[0118] S1, when the historical matching information of the second reference object indicates that the candidate object with the most matches with the second reference object is not the first candidate object, determining that the second reference object does not meet the historical matching condition;
[0119] S2, determining an object acquisition order according to the number of matches corresponding to each of the at least one second reference objects;
[0120] S3, according to the object acquisition order, sequentially acquiring a second reference object from at least one second reference object as the current reference object;
[0121] S4. When the historical matching information of the current reference object indicates that the first candidate object is the candidate object with the largest number of matches with the current reference object, determine the current reference object as the second candidate object that meets the historical matching condition.
[0122] The above process is described below in a specific implementation manner:
[0123] First, assume that the Car target with 2D true value ID 5 in the nth frame has no 3D projection frame target matching it after two rounds of matching strategies: overlapping area matching and projection frame center point distance matching. For example, the IoU between the Car target with 2D true value ID 5 and all 3D projection frames is 0. In S1-S4, first count whether the Car target with 2D true value ID 5 has always existed in the past 10 consecutive frames. If it has always existed, it is the most ideal situation. If it has not always existed, only count the continuous scenes where the Car target with 2D true value ID 5 in the past has appeared continuously to the current frame, such as the n-5th frame to the nth frame.
[0124] Assume that the Car target with 2D true value frame ID 5 has always appeared in 10 consecutive historical frames, and count the matching status of the Car target with 2D true value frame ID 5 in 10 consecutive historical frames. Figure 4 As shown, the Car target with 2D truth frame ID 5 has matched the Car target with 3D truth projection frame ID 3 5 times, matched the Car target with 3D truth projection frame ID 6 4 times, and matched the Car target with 3D truth projection frame ID 7 once in 10 consecutive historical frames. The historical matching results are arranged in descending order.
[0125] From the descending results, we can see that in the historical matching stage, the Car target with 2D truth frame ID 5 matches the Car target with 3D truth projection frame ID 3 the most times. This is the process of counting the matching information of the 3D projection frame based on the 2D truth target frame. It is also necessary to count the matching information of the 2D truth target frame based on the 3D projection frame.
[0126] Further based on the result that the Car target with 2D truth box ID 5 matches the Car target with 3D projection ID 3 the most times in the above historical matching stage, the matching information of the Car target with 3D truth projection box ID 3 in the historical matching stage is further counted.
[0127] Assuming that the Car target with the ID of 3 in the 3D true value projection frame is also matched the most times with the Car target with the ID of 5 in the 2D true value frame in the historical matching information, the Car target with the ID of 5 in the 2D true value frame is matched with the Car target with the ID of 3 in the 3D true value projection frame in the nth frame;
[0128] If the Car target with 3D truth projection frame ID 3 is not matched the most times with the Car target with 2D truth frame ID 5 in the historical matching information, further based on the statistical results of the above step, the historical matching information statistics of the Car target with 3D truth projection frame ID 6, which has the second most matches with the Car target with 2D truth ID 5, are performed, and the second candidate object is determined in this way, thereby improving the accuracy of the matching results.
[0129] In an optional implementation, after determining the object matching relationship between the first candidate object and the second candidate object, the method further includes:
[0130] S1, acquiring a second annotation box corresponding to a second candidate object that meets a historical matching condition, and a second verification annotation box corresponding to a second verification candidate object that matches the first candidate object and is determined according to key frame matching information;
[0131] S2: When the area of the overlapping region between the second verification annotation box and the second annotation box is greater than or equal to the area threshold, determine that the matching result is correct.
[0132] The above steps S1-S2 are described below in a specific implementation manner:
[0133] We search for the nearest key frame based on the current nth frame. The key frame is manually extracted and sent for manual annotation. The number of key frames in a continuous frame segment is fixed. The number of key frames and the key frame extraction rules are set in advance. If the nearest key frame cannot be found, no key frame correction is performed. Here, it is assumed that based on the nth frame, the nearest key frame m is found, and the key frame matching information corresponding to key frame m is obtained.
[0134] For example, the second candidate object that meets the historical matching conditions and matches the first candidate object is the Car target with a 2D truth value ID of 5 in the current nth frame. The 2d truth value information (second verification candidate object) that matches the first candidate object is determined based on the joint annotation matching information of the 2D continuous frame truth value and the 3D continuous frame truth value in the key frame m (for example, the key frame matching information). It is determined whether the overlapping area between the 3d truth value box of the second verification candidate object in the key frame matching information and the 2d truth box corresponding to the Car target with a 2D truth value ID of 5 is greater than the threshold condition. If the threshold condition is met, the matching result can be verified to be correct.
[0135] It should be noted that the above key frame matching information may be, but is not limited to, the matching relationship between the first candidate object and the second candidate object manually marked by a person skilled in the art selecting key frames in various scenarios.
[0136] Through the above implementation, it can be confirmed whether the Car target with 2D true value ID 5 in the nth frame of the current continuous frame segment also exists in the key frame m; if so, Figure 5 As shown, for example, according to the IoU overlap value, it is determined that the Car target with a 2D true value box ID of 5 corresponds to the Car target with a 2D true value box ID of 14 in the key frame joint annotation result;
[0137] Then the matching result is verified, that is, the Car target with 3D true value projection frame ID 15 matched with the Car target with 2D true value frame ID 14 determined according to the key frame matching information (key frame joint annotation result) in key frame m is calculated, and whether the IoU with the Car target with 3D true value projection frame ID 3 in the nth frame in the matching result is greater than the threshold. If it is greater than the threshold condition, it is considered that the Car target with 2D true value frame ID 14 matched by the key frame m joint annotation and the Car target with 3D true value projection frame ID 15 matched therewith are completely corresponding to the Car target with 2D true value frame ID 5 and the Car target with 3D true value projection frame ID 3 matched therewith in the nth frame, thereby confirming that the result based on the third round of matching for the nth frame is correct, that is, the matching result of the nth frame with the Car target with 2D true value frame ID 5 and the Car target with 3D true value projection frame ID 3 is correct.
[0138] Through the above implementation, a second annotation box corresponding to the second candidate object that meets the historical matching condition is obtained, and a second verification annotation box corresponding to the second verification candidate object that matches the first candidate object determined according to the key frame matching information is obtained; when the area of the overlapping area between the second verification annotation box and the second annotation box is greater than or equal to the area threshold, the matching result is determined to be correct. The historical matching result is verified to ensure the accuracy of the historical matching result, thereby improving the accuracy of the entire matching process.
[0139] In an optional implementation, after determining the object matching relationship between the first candidate object and the second candidate object, the method further includes:
[0140] S1, determining the second candidate object that has the least number of matches with the first candidate object in the matching results as the second candidate object to be corrected that meets the verification condition, wherein the matching results include a plurality of first candidate objects and second candidate objects that have a matching relationship;
[0141] S2, acquiring a key frame matching the second candidate object to be corrected, and key frame matching information corresponding to the key frame;
[0142] S3: modify the second candidate object to be corrected that matches the first candidate object into a second corrected object according to the key frame matching information.
[0143] The above steps S1-S3 are described below in a specific implementation manner:
[0144] like Figure 6 As shown, in the n-10th frame of the left matching result, the Car target with a 2D true value ID of 5 matches the 3D true value projection frame ID of 7, and this match only occurs in one frame, which is the least number of times. Therefore, the second candidate object with a 3D true value projection frame ID of 7 is determined as the second candidate object to be corrected;
[0145] Further optimization is performed with the help of key frames. Assume that the key frame closest to the n-10th frame is key frame m. For example, the Car target with 2D truth frame ID 5 in the nth frame corresponds to the Car target with 2D truth frame ID 14 in the manual joint matching annotation (key frame matching information) of the key frame. The Car target with 3D truth projection frame ID 15 in key frame m that matches the Car target with 2D truth frame ID 14 is calculated based on the Car target with 3D truth projection frame ID 15. The IoU of all 3D projection frames in the n-10th frame is calculated, and then the matching is sorted from large to small. It is found that among all 3D projection frames in the n-10th frame, the Car target with ID 6 has the largest overlap with the Car target with 3D truth projection frame ID 15 in key frame m. Based on this, the 3D truth projection frame ID of the n-10th frame that matches the Car target with 2D truth frame ID 5 is modified to 6, such as Figure 6 The matching results are shown on the right.
[0146] Through the above implementation, the second candidate object that matches the first candidate object the least times in the matching results is determined as the second candidate object to be corrected that meets the verification condition; the key frame that matches the second candidate object to be corrected and the key frame matching information corresponding to the key frame are obtained; the second candidate object to be corrected that matches the first candidate object is modified to the second correction object according to the key frame matching information. The abnormal matching information in the matching results is corrected by the annotation results corresponding to the key frame, thereby improving the accuracy of the road object matching results.
[0147] The following is a complete implementation method to illustrate this solution:
[0148] Different from the current method of obtaining true values for perception model algorithm evaluation, the cloud-based truth value large model trained in this application supports the output of 2D continuous frame truth value results and 3D continuous frame truth value results. In order to solve the problem of serious offset in the position and size of the 3D truth value box of the 3D continuous frame truth value, and the possible inability to output the truth value due to the sparse point cloud at distant targets, after the 2D and 3D continuous frame truth values are generated, three rounds of matching and one round of key frame manual annotation auxiliary optimization algorithms are designed to perform joint annotation of 2D and 3D continuous frame truth values.
[0149] Joint matching annotation can also effectively solve the problem that the vehicle-side algorithm model has large errors in measuring the distance to distant targets in pure 3D space, resulting in the inability to associate with the true value during evaluation. This is because through joint matching annotation, this step can be associated in a 2D frame.
[0150] The unified perspective of joint truth matching is the 2D image perspective. First, based on the camera intrinsic parameters and the joint calibration parameters of the camera and lidar, the 3D target truth frame is projected to the 2D image. This step projects the 3D frame to the 2D image. Due to the maximum circumscribed rectangle, the projected 2D frame will expand, which is normal.
[0151] Secondly, the 2D frame projected onto the 2D image from the 3D target truth frame needs to be screened. Due to the characteristics of point clouds, the cloud-based truth model can output targets that are completely invisible in 2D images based on point cloud data. In this step, the truth targets that are completely invisible in 2D images need to be eliminated, because the framework of the entire joint matching annotation process is set within the 2D image range. Three discrimination conditions are set, based on the target truth value being larger when near and smaller when far, and the fact that the depth information of the 3D truth frame is still retained despite the projection operation on the 3D truth frame. The size of the 2D projection frame, the distance, and the relative relationship of the center pixel coordinates of the frame can be used to eliminate the 3D projection frames of the truth targets that are completely invisible from the 2D image perspective.
[0152] like Figure 3As shown, taking the front view image as an example, the dotted box represents the front view image, the solid line truth box on the left is the 3D truth box corresponding to the vehicle object 302 projected onto the 2D image, and the dotted line truth box on the left is the 3D truth box corresponding to the vehicle object 304 projected onto the 2D image. The solid line truth box on the left surrounds the dotted line truth box on the left, and the vehicle object 304 is farther away from the vehicle 308 longitudinally, so the dotted line truth box on the left can be eliminated because this object cannot be seen under the front view camera.
[0153] By eliminating the 3D true value boxes that are not visible from the camera's perspective and projecting them onto the 2D image, we enter the first round of matching.
[0154] In the first round of matching, the target category is first divided into 2D truth boxes and 3D projection boxes, and the matching process is framed in the same category of 2D truth boxes and 3D projection boxes. The first round of matching is mainly based on IoU matching, and direct matching is performed by calculating whether the numerical value of the bounding_box_overlap of the 2D truth box and the 3D projection box is greater than the set threshold. If the IoU of the 2D truth box and the 3D projection box of the same category is greater than the set threshold, the 2D truth box and the 3D projection box will be associated. Because we retain the information of the original 3D truth box of the 3D projection box, we can give the 2D truth box 3D information, such as horizontal and vertical depth and horizontal and vertical speed information, and because it is a continuous frame truth, the target tracking ID of the 2D truth box and the 3D projection frame can also be associated. In this round of matching, targets that are close to the vehicle and not too coupled with other targets will be jointly matched. However, because this round of matching is completely based on the IoU value, only when the IoU result is greater than the set threshold will it be associated, which will cause many missed matches. Therefore, the first round of matching is a coarse match.
[0155] like Figure 7 As shown, box 704 represents the 3D true value projection box, and box 702 represents the 2D true value box. Since the first round of matching will only associate 2D and 3D true value targets whose IoU values are greater than the selected threshold, there will be many missed matching targets in one round of matching results, because the 3D true value projection box of the farther target is not only expanded but also offset, and there will be no IoU information, so it cannot be matched.
[0156] In order to solve the problem of complex scenes with many true targets or targets occluding or overlapping each other, a two-round matching method is proposed.
[0157] The second round of matching is based on the matching process after eliminating the results of the first round of matching. This round of matching will only associate the targets whose IoU values of the 2D true value box and the 3D projection box are not 0. This step will sort the IoU values of the 2D true value box from large to small, calculate the Euclidean distance of the pixel coordinates of the center point of the 2D true value box and the 3D projection box, and sort them. Matching is performed based on the IoU sorting results and the center point Euclidean distance sorting results.
[0158] Double judgment is performed on the two types of sorting results. For example, the 2D truth box with index 1 has the largest IoU value with its 3D truth projection box with index 3, and the 2D truth box with the closest Euclidean distance to the center pixel coordinates of the 3D truth projection box with index 3 has index 1. That is, double judgment is performed, and the 2D truth box with index 1 and the 3D truth projection box with index 3 are associated and matched.
[0159] The second round of matching is still to solve the matching of targets whose IoU values between the 2D true value box and the 3D projection box are not 0, and the third round of matching is mainly to solve the matching of targets whose IoU values are 0.
[0160] First, based on the elimination of the first and second round matching results, the 2D true value boxes and 3D target boxes with an IoU value of 0 are screened out, and the matching results of each 2D true value box in the historical 10 consecutive frames (if less than 10 frames, from the start frame to the current frame) are calculated for analysis, and the joint annotation information of the key frames is combined to assist in optimization. The specific matching ideas are as follows:
[0161] First, assume that the Car target with 2D true value ID 5 in the nth frame has no 3D projection frame target matching it after the first and second rounds of matching. That is, the IoU between the Car target with 2D true value ID 5 and all 3D projection frames is 0. At the beginning of the third round of matching, we first count whether the Car target with 2D true value ID 5 has always existed in the past 10 consecutive frames. If it has always existed, it is the most ideal situation. If it has not always existed, we only count the continuous scenes where the Car target with 2D true value ID 5 in the past has appeared continuously to the current frame, such as the n-5th frame to the nth frame.
[0162] Now assume that in the 10 consecutive historical frames, the Car target with 2D true value ID 5 has always appeared, and count the matching of the Car target with 2D true value ID 5 in the 10 consecutive historical frames. Figure 4As shown in the figure, the Car target with 2D truth value ID 5 has matched with the Car target with 3D projection ID 3 5 times, 4 times, and 1 time in 10 consecutive historical frames. We sort the historical matching results in descending order. From the descending results, we can see that the Car target with 2D truth value ID 5 matches the Car target with 3D projection ID 3 the most times in the historical matching stage. This is the process of counting the matching information of the 3D projection frame based on the 2D truth value target frame. We also need to count the matching information of the 2D truth value target frame based on the 3D projection frame.
[0163] For example, based on the result that the Car target with 2D true value ID 5 matches the Car target with 3D projection ID 3 the most times in the above historical matching stage, we will count the matching information of the Car target with 3D projection ID 3 in the historical matching stage. Assuming that the Car target with 3D projection ID 3 also matches the Car target with 2D true value ID 5 the most times in the historical matching information, we will match the Car target with 2D true value ID 5 with the Car target with 3D projection ID 3 in the nth frame; if the Car target with 3D projection ID 3 does not match the Car target with 2D true value ID 5 the most times in the historical matching information, then we will count the historical matching information of the Car target with 3D projection ID 6, which has the second most matches with the Car target with 2D true value ID 5, based on the statistical results of the above step, and so on.
[0164] After the above two steps, assuming that the Car target with 2D true value ID 5 has been matched with the Car target with 3D projection ID 3, the matching sequence is corrected and optimized based on the manual joint annotation results of the key frames. The specific ideas are as follows:
[0165] First, we look for the nearest key frame based on the current nth frame. The key frame is manually extracted and sent for manual annotation. The number of key frames in a continuous frame segment is fixed. The number of key frames and the key frame extraction rules are set in advance. If the nearest key frame cannot be found, no key frame correction is performed. Suppose that the nearest key frame m is found based on the nth frame. Based on the joint annotation information of the key frame, we can know the joint annotation matching information of the 2D continuous frame truth value and the 3D continuous frame truth value of the key frame (such as Figure 5 As shown in the joint annotation result of the key frames on the right, based on the joint annotation information of the key frames, the IoU is calculated for the Car target with the 2D truth value ID of 5 in the current nth frame and the 2D truth box annotated by the joint matching of the key frames. The purpose is: first, to confirm whether the Car target with the 2D truth value ID of 5 in the nth frame of the current continuous frame segment also exists in the key frame m; second, if it exists, Figure 5, the Car target with 2D truth value ID 5 in the nth frame corresponds to the Car target with 2D truth value ID 14 marked by the manual joint matching of the key frame, and then calculates whether the IoU of the Car target with 3D truth value ID 15 matched with the Car target with 2D truth value ID 14 in the key frame m and the Car target with 3D projection ID 3 in the nth frame is greater than the threshold. If it is greater than the threshold condition, it is considered that the Car target with 2D truth value ID 14 and the Car target with 3D truth value ID 15 matched by the joint annotation of the key frame m are completely corresponding to the Car target with 2D truth value ID 5 and the Car target with 3D projection ID 3 matched by the nth frame, thereby confirming that the result based on the third round of matching is correct for the nth frame.
[0166] Then, yes Figure 6 In the n-10th frame, the Car target with 2D truth ID 5 matches the 3D projection truth ID 7, and this match only appears in one frame, so the key frame is used for optimization. Assume that the key frame closest to the n-10th frame is also key frame m. This assumption is valid in most cases because there are few key frames and the span is large.
[0167] Based on the previous step, it is known that the Car target with 2D truth value ID 5 in the nth frame corresponds to the Car target with 2D truth value ID 14 marked by the manual joint matching of the key frame. The Car target with 3D truth value ID 15 in the key frame m matches the Car target with 2D truth value ID 14. Based on the Car target with 3D truth value ID 15, the IoU of all 3D projection frames in the n-10th frame and it is calculated, and then sorted from large to small for matching. It is concluded that in all 3D projection frames of the n-10th frame, the Car target with ID 6 corresponds to the Car target with 3D truth value ID 15 in the key frame m. Based on this, the 3D projection truth value ID of the Car target matching the 2D truth value ID 5 in the n-10th frame is modified to 6. The modified result is as follows Figure 6 Shown on the right.
[0168] Flowchart Figure 8 The above process is explained as follows:
[0169] On the one hand, S802 is executed to obtain the area of the 2D truth box; S804, the 3D truth box point cloud xyz determines the distance before and after; S806, the 2D truth box border coordinates determine the relative position;
[0170] Execute S808 according to S802-S806, and comprehensively judge the inclusion relationship to determine the occlusion;
[0171] On the other hand, S810 is executed to obtain the true value of the 2D continuous frame;
[0172] S812, obtaining 3D continuous frame true values;
[0173] S814, projecting the 3D target truth frame onto the 2D image; executing S816 in conjunction with S808 to filter out objects that are not completely blocked;
[0174] On the other hand, S818, manual labeling of key frames; combining S810 and S816 to perform S820, three rounds of matching;
[0175] S822, detection result of the perception algorithm model to be evaluated; and executing S824 in combination with step S820 to associate the detection result.
[0176] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0177] According to another aspect of the embodiments of the present application, a road object matching device for implementing the road object matching method is also provided. Fig. 9 As shown, the device comprises:
[0178] A first acquisition unit 902 acquires a first road object set and a second road object set that match the target road scene, wherein the first road object set is obtained by performing image recognition on a target road image corresponding to the target road scene, and the second road object set is obtained by performing point cloud recognition on a point cloud image corresponding to the target road scene;
[0179] A second acquiring unit 904 is configured to acquire an object position relationship between a plurality of candidate objects in the first road object set and the second road object set;
[0180] A first determining unit 906 determines an object matching relationship between the first candidate object in the first road object set and the second candidate object in the second road object set if the object position relationship between the first candidate object in the first road object set and the second candidate object in the second road object set meets the target position condition, wherein the first candidate object and the second candidate object having the object matching relationship are used to indicate the same target road object in the target road scene;
[0181] The second determination unit 908 determines, from the second road object set, a second candidate object that meets the historical matching condition, based on historical matching information of the first candidate object, when the object position relationship between the first candidate object in the first road object set and any candidate object in the second road object set does not meet the target position condition, wherein the historical matching information includes an object matching result corresponding to at least one historical road scene, and the collection timestamp corresponding to the historical road scene is earlier than the collection timestamp of the target road scene.
[0182] As an optional solution, the above-mentioned road object matching device also includes: a third determination unit, which is used to determine the object matching relationship between the first candidate object and the second candidate object when the area of the overlapping area between the first annotation box corresponding to the first candidate object and the second annotation box of the second candidate object is greater than or equal to the area threshold; determine the object matching relationship between the first candidate object and the second candidate object when the annotation box distance between the first annotation box corresponding to the first candidate object and the second annotation box of the second candidate object is less than or equal to the distance threshold; wherein the first annotation box is used to indicate a first image area of the first candidate object in the target road image, and the second annotation box is used to indicate a second image area of the second candidate object in the target road image, and the second image area is a projection result of point cloud information based on the matching of the second candidate object on the target road image.
[0183] As an optional solution, the above-mentioned third determination unit includes: a third acquisition module, used to acquire a point cloud image collected from the target road scene; perform point cloud recognition on the point cloud image to obtain a second reference road object set, wherein the second reference road object set includes multiple reference objects; based on the point cloud information matched by each of the second reference road object sets, determine the reference annotation boxes obtained by projecting the multiple reference objects in the target road image respectively; according to the occlusion relationship of the multiple reference annotation boxes in the target road image, screen out multiple unobstructed reference annotation boxes from the multiple reference annotation boxes; and determine the second road object set according to the candidate objects indicated by each of the multiple unobstructed reference annotation boxes.
[0184] As an optional solution, the third determination unit includes: a fourth acquisition module, which is used to acquire a second reference annotation frame subset matching the first annotation frame from a second reference annotation frame set matching the second road object set when the overlapping areas between the first annotation frame and the respective candidate annotation frames of any candidate object in the second road object set are all smaller than the area threshold, wherein the area of the overlapping areas between the second reference annotation frame in the second reference annotation frame subset and the first annotation frame is greater than 0 and smaller than the area threshold; determine a target annotation frame from the second reference annotation frame subset, wherein the area of the overlapping area between the first annotation frame and the target annotation frame is greater than the area of the overlapping areas corresponding to the other annotation frames in the second reference annotation frame subset; and determine that the third candidate object indicated by the target annotation frame has an object matching relationship with the first road object when the annotation frame distance between the target annotation frame and the first annotation frame meets the target distance condition.
[0185] As an optional solution, the fourth acquisition module includes: a fifth acquisition module, used to obtain the target annotation box and the annotation box distance between multiple first reference annotation boxes in the first reference annotation box set, wherein the first reference annotation box is used to indicate the image area of the candidate object in the first road object set in the target road image, and the multiple first reference annotation boxes include the first annotation box; when the annotation box distance between the target annotation box and the first annotation box is smaller than the annotation box distance corresponding to each of the other annotation boxes in the first reference annotation box set, determine that the annotation box distance between the target annotation box and the first annotation box meets the target distance condition.
[0186] As an optional scheme, the above-mentioned second determination unit also includes a sixth acquisition module, which is used to obtain historical matching information, wherein the historical matching information includes object matching results of the first candidate object in multiple historical road scenes, the multiple historical road scenes and the target road scene are road scenes corresponding to multiple information collection operations respectively, the multiple information collection operations are information collection operations corresponding to multiple continuous timestamps respectively, the information collection operations include image data collection operations and point cloud data collection operations, and the object matching results are used to indicate a second reference object that matches the first candidate object in the current road scene; according to the object matching results corresponding to each of the multiple historical road scenes, a second candidate object that meets the historical matching conditions is determined from at least one second reference object.
[0187] As an optional scheme, the above-mentioned sixth acquisition module includes: a fourth determination module, used to determine the second reference object with the largest number of matches as the second candidate object that meets the historical matching conditions; obtain the historical matching information of the second reference object with the largest number of matches; and determine the second reference object as the second candidate object that meets the historical matching conditions based on the historical matching information of the second reference object.
[0188] As an optional scheme, the above-mentioned fourth determination module is also used to determine that the second reference object does not meet the historical matching conditions when the historical matching information of the second reference object indicates that the candidate object with the most matches with the second reference object is not the first candidate object; determine the object acquisition order according to the number of matches corresponding to at least one second reference object; acquire a second reference object from at least one second reference object in sequence as the current reference object according to the object acquisition order; and determine the current reference object as the second candidate object that meets the historical matching conditions when the historical matching information of the current reference object indicates that the first candidate object is the candidate object with the most matches with the current reference object.
[0189] As an optional scheme, the above-mentioned fourth determination module is also used to obtain a second annotation box corresponding to a second candidate object that meets the historical matching conditions, and a second verification annotation box corresponding to a second verification candidate object that matches the first candidate object determined according to the key frame matching information; when the area of the overlapping area between the second verification annotation box and the second annotation box is greater than or equal to the area threshold, determine that the matching result is correct.
[0190] As an optional scheme, the above-mentioned fourth determination module is also used to determine the second candidate object that matches the first candidate object the least number of times in the matching results as the second candidate object to be corrected that meets the verification conditions, wherein the matching results include multiple first candidate objects and second candidate objects with a matching relationship; obtain the key frame that matches the second candidate object to be corrected, and the key frame matching information corresponding to the key frame; and modify the second candidate object to be corrected that matches the first candidate object to be corrected into the second correction object according to the key frame matching information.
[0191] For specific embodiments, reference may be made to the examples shown in the above-mentioned road object matching method, which will not be described in detail in this example.
[0192] Among them, the memory 1002 can be used to store software programs and modules, such as program instructions / modules corresponding to the road object matching method and device in the embodiment of the present invention. The processor 1004 executes various functional applications and data processing by running the software programs and modules stored in the memory 1002, that is, realizing the above-mentioned road object matching method. The memory 1002 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1002 may further include a memory remotely located relative to the processor 1004, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 1002 can be specifically, but not limited to, used to store file information such as target logical files. As an example, if Fig.10 As shown, the memory 1002 may include, but is not limited to, the first acquisition unit 902, the second acquisition unit 904, the first determination unit 906, and the second determination unit 908 in the road object matching device. In addition, it may also include, but is not limited to, other module units in the road object matching device, which will not be described in detail in this example.
[0193] Optionally, the transmission device 1006 is used to receive or send data via a network. Specific examples of the network may include a wired network and a wireless network. In one example, the transmission device 1006 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers via a network cable so as to communicate with the Internet or a local area network. In one example, the transmission device 1006 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0194] In addition, the electronic device mentioned above further includes: a display 1008, and a connection bus 1010, which is used to connect various module components in the electronic device mentioned above.
[0195] According to one aspect of the present application, a computer program product is provided, the computer program product comprising a computer program / instruction, the computer program / instruction comprising a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium. When the computer program is executed by a central processing unit, various functions provided by the embodiments of the present application are executed.
[0196] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0197] It should be noted that the computer system of the electronic device is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0198] In particular, according to an embodiment of the present application, the process described in each method flow chart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer readable medium, and the computer program contains a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium. When the computer program is executed by a central processing unit, various functions defined in the system of the present application are executed.
[0199] According to one aspect of the present application, a computer-readable storage medium is provided, and a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above-mentioned various optional implementations.
[0200] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0201] S1, obtaining a first road object set and a second road object set matching a target road scene, wherein the first road object set is obtained by performing image recognition on a target road image corresponding to the target road scene, and the second road object set is obtained by performing point cloud recognition on a point cloud image corresponding to the target road scene;
[0202] S2, obtaining an object position relationship between a plurality of candidate objects in the first road object set and the second road object set;
[0203] S3, when the object position relationship between the first candidate object in the first road object set and the second candidate object in the second road object set meets the target position condition, determining an object matching relationship between the first candidate object and the second candidate object, wherein the first candidate object and the second candidate object having the object matching relationship are used to indicate the same target road object in the target road scene;
[0204] S4. When the object position relationship between the first candidate object in the first road object set and any candidate object in the second road object set does not satisfy the target position condition, a second candidate object satisfying the historical matching condition is determined from the second road object set based on historical matching information of the first candidate object, wherein the historical matching information includes an object matching result corresponding to at least one historical road scene, and a collection timestamp corresponding to the historical road scene is earlier than a collection timestamp of the target road scene.
[0205] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing hardware related to the electronic device through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0206] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0207] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling one or more computer devices (which can be personal computers, servers or network devices, etc.) to execute all or part of the steps of the methods of each embodiment of the present application.
[0208] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0209] In the several embodiments provided in the present application, it should be understood that the disclosed user equipment can be implemented in other ways. Among them, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0210] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0211] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0212] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A road object matching method, characterized in that: include: Acquire a first road object set and a second road object set that match a target road scene, wherein the first road object set is obtained by performing image recognition on a target road image corresponding to the target road scene, and the second road object set is obtained by performing point cloud recognition on a point cloud image corresponding to the target road scene; Acquire a point cloud image acquired from the target road scene; perform point cloud recognition on the point cloud image to obtain a second reference road object set, wherein the second reference road object set includes a plurality of reference objects; determine reference annotation frames obtained by projecting the plurality of reference objects in the target road image based on the point cloud information matched by each of the second reference road object sets; select a plurality of unobstructed reference annotation frames from the plurality of reference annotation frames based on an occlusion relationship of the plurality of reference annotation frames in the target road image; determine the second road object set based on candidate objects indicated by each of the plurality of unobstructed reference annotation frames; Acquire an object position relationship between a plurality of candidate objects in the first road object set and the second road object set; In the case where the object position relationship between a first candidate object in the first road object set and a second candidate object in the second road object set meets a target position condition, determining the object matching relationship between the first candidate object and the second candidate object comprises at least one of the following: determining the object matching relationship between the first candidate object and the second candidate object when the area of an overlapping area between a first annotation box corresponding to the first candidate object and a second annotation box of the second candidate object is greater than or equal to an area threshold; determining the object matching relationship between the first candidate object and the second candidate object when the annotation box distance between the first annotation box corresponding to the first candidate object and the second annotation box of the second candidate object is less than or equal to a distance threshold; wherein the first annotation box is used to indicate a first image area of the first candidate object in the target road image, the second annotation box is used to indicate a second image area of the second candidate object in the target road image, the second image area is a projection result of point cloud information based on the matching of the second candidate object on the target road image, and the first candidate object and the second candidate object having the object matching relationship are used to indicate the same target road object in the target road scene; When the object position relationship between the first candidate object in the first road object set and any one of the candidate objects in the second road object set does not satisfy the target position condition, the second candidate object that satisfies the historical matching condition is determined from the second road object set based on historical matching information of the first candidate object, wherein the historical matching information includes object matching results corresponding to at least one historical road scene, and the acquisition timestamp corresponding to the historical road scene is earlier than the acquisition timestamp of the target road scene.
2. The method according to claim 1, characterized in that: The determining of the object matching relationship between the first candidate object in the first road object set and the second candidate object in the second road object set when the object position relationship between the first candidate object and the second candidate object in the second road object set meets the target position condition further includes: When the overlapping areas between the first annotation box and the candidate annotation boxes of any candidate object in the second road object set are both smaller than the area threshold, a second reference annotation box subset matching the first annotation box is acquired from a second reference annotation box set matching the second road object set, wherein the overlapping areas between the second reference annotation boxes in the second reference annotation box subset and the first annotation box are greater than 0 and smaller than the area threshold; Determine a target annotation frame from the second reference annotation frame subset, wherein the area of an overlap region between the first annotation frame and the target annotation frame is greater than the area of an overlap region corresponding to other annotation frames in the second reference annotation frame subset; When the annotation box distance between the target annotation box and the first annotation box meets a target distance condition, it is determined that a third candidate object indicated by the target annotation box has the object matching relationship with the first road object.
3. The method according to claim 2, characterized in that After determining the target annotation frame from the second reference annotation frame subset, the method further includes: Acquire a distance between the target annotation frame and a plurality of first reference annotation frames in a first reference annotation frame set, wherein the first reference annotation frame is used to indicate an image region of the candidate object in the first road object set in the target road image, and the plurality of first reference annotation frames include the first annotation frame; When the annotation box distances between the target annotation box and the first annotation box are all smaller than the annotation box distances corresponding to other annotation boxes in the first reference annotation box set, the annotation box distance between the target annotation box and the first annotation box is determined to satisfy the target distance condition.
4. The method according to claim 1, characterized in that The determining, based on the historical matching information of the first candidate object, the second candidate object that meets the historical matching condition from the second road object set includes: Acquire the historical matching information, wherein the historical matching information includes object matching results of the first candidate object in a plurality of historical road scenes, the plurality of historical road scenes and the target road scene are road scenes corresponding to a plurality of information collection operations respectively, the plurality of information collection operations are information collection operations corresponding to a plurality of continuous timestamps respectively, the information collection operations include image data collection operations and point cloud data collection operations, and the object matching result is used to indicate a second reference object that matches the first candidate object in the current road scene; According to the object matching results corresponding to each of the plurality of historical road scenes, a second candidate object satisfying the historical matching condition is determined from at least one second reference object.
5. The method according to claim 4, characterized in that The determining, based on the object matching results corresponding to each of the plurality of historical road scenes, the second candidate object satisfying the historical matching condition from at least one of the second reference objects comprises one of the following: Determine the second reference object having the largest number of matches as the second candidate object satisfying the historical matching condition; The historical matching information of the second reference object with the largest number of matches is obtained; and according to the historical matching information of the second reference object, the second reference object is determined to be the second candidate object that meets the historical matching condition.
6. The method according to claim 5, characterized in that After acquiring the historical matching information of the second reference object with the largest number of matches, the method further includes: If the historical matching information of the second reference object indicates that the candidate object with the largest number of matches with the second reference object is not the first candidate object, determining that the second reference object does not satisfy the historical matching condition; Determining an object acquisition order according to the number of matches corresponding to at least one of the second reference objects; According to the object acquisition order, sequentially acquiring one of the second reference objects from at least one of the second reference objects as a current reference object; When the historical matching information of the current reference object indicates that the first candidate object is the candidate object with the largest number of matches with the current reference object, the current reference object is determined as the second candidate object that satisfies the historical matching condition.
7. The method according to claim 5, characterized in that After determining the object matching relationship between the first candidate object and the second candidate object, the method further includes: Acquire a second annotation box corresponding to the second candidate object that meets the historical matching condition, and a second verification annotation box corresponding to the second verification candidate object that matches the first candidate object and is determined according to the key frame matching information; When the area of the overlapping region between the second verification annotation box and the second annotation box is greater than or equal to the area threshold, it is determined that the matching result is correct.
8. The method according to claim 5, characterized in that After determining the object matching relationship between the first candidate object and the second candidate object, the method further includes: Determine the second candidate object that has the least number of matches with the first candidate object in the matching results as the second candidate object to be corrected that meets the verification condition, wherein the matching results include a plurality of the first candidate objects and the second candidate objects that have a matching relationship; Acquire a key frame matching the second candidate object to be corrected, and key frame matching information corresponding to the key frame; The second candidate object to be corrected that matches the first candidate object is modified into a second corrected object according to the key frame matching information.
9. A road object matching device, characterized in that: include: A first acquisition unit is configured to acquire a first road object set and a second road object set that match a target road scene, wherein the first road object set is obtained by performing image recognition on a target road image corresponding to the target road scene, and the second road object set is obtained by performing point cloud recognition on a point cloud image corresponding to the target road scene; A second acquisition unit acquires a point cloud image acquired from the target road scene; performs point cloud recognition on the point cloud image to obtain a second reference road object set, wherein the second reference road object set includes a plurality of reference objects; based on the point cloud information matched by each of the second reference road object sets, determines reference annotation frames obtained by projecting the plurality of reference objects in the target road image; based on the occlusion relationship of the plurality of reference annotation frames in the target road image, selects a plurality of unoccluded reference annotation frames from the plurality of reference annotation frames; determines the second road object set based on the candidate objects indicated by each of the plurality of unoccluded reference annotation frames; and acquires the object position relationship between the first road object set and the plurality of candidate objects in the second road object set; a first determining unit, wherein, when the object position relationship between a first candidate object in the first road object set and a second candidate object in the second road object set meets a target position condition, determining an object matching relationship between the first candidate object and the second candidate object comprises at least one of the following: determining the object matching relationship between the first candidate object and the second candidate object when an area of an overlapping area between a first annotation box corresponding to the first candidate object and a second annotation box of the second candidate object is greater than or equal to an area threshold; and determining the object matching relationship between the first candidate object and the second candidate object when an annotation box distance between the first annotation box corresponding to the first candidate object and the second annotation box of the second candidate object is less than or equal to a distance threshold; wherein the first annotation box is used to indicate a first image area of the first candidate object in a target road image, the second annotation box is used to indicate a second image area of the second candidate object in the target road image, the second image area is a projection result of point cloud information based on matching of the second candidate object on the target road image, and the first candidate object and the second candidate object having the object matching relationship are used to indicate the same target road object in the target road scene; A second determination unit is configured to determine, when the object position relationship between the first candidate object in the first road object set and any one of the candidate objects in the second road object set does not satisfy the target position condition, the second candidate object that satisfies the historical matching condition from the second road object set based on historical matching information of the first candidate object, wherein the historical matching information includes an object matching result corresponding to at least one historical road scene, and the acquisition timestamp corresponding to the historical road scene is earlier than the acquisition timestamp of the target road scene.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method described in any one of claims 1 to 8 when executed by an electronic device.
11. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 8 through the computer program.
12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Road object identification method and device, storage medium and electronic equipment
CN117576652A