Target intersection fusion method and device, electronic equipment and storage medium
By using the BEV multimodal data fusion method in the vehicle-road collaboration scenario and generating BEV multimodal data using lidar and cameras, the occlusion and multi-pole logic problems in intersection and road section fusion are solved, and the accuracy and consistency of the fusion results are improved.
Patent Information
- Application Number
- CN202510640973.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-09-05
AI Technical Summary
In the vehicle-road collaboration scenario, during the fusion process of intersections and road sections, there are problems with poor indicators caused by occlusion, inaccurate single-pole perception results, and multi-pole fusion logic, especially the inconsistency between laser perception and image perception results.
The BEV multimodal data fusion method is adopted to generate BEV multimodal data using lidar and different types of cameras. Through the multimodal data fusion of intersections and road sections, the final intersection fusion result is generated, including the target's 3D detection frame, heading angle, speed and position information, and it is converted into the world coordinate system.
It improves the accuracy of fusion results at intersections and road sections, solves occlusion and multi-pole fusion logic problems, improves the indicators of target position, speed and heading angle, and ensures the consistency of single-pole results and multi-pole results.
Smart Images

Figure CN120599554A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of vehicle-road collaboration technology, and in particular to a target intersection fusion method, device, electronic device, and storage medium. Background Art
[0002] In the vehicle-road collaboration scenario, the same target needs to be integrated and tracked at intersections and road sections.
[0003] Common fusion methods typically use poles as units, installing sensors on the poles to detect the pole's direction of approach, destination, and the area below the pole. Multiple individual poles are then post-fused across the intersection to obtain data for the entire intersection and road section. However, this presents the following challenges:
[0004] (1) Occlusions in the area facing the intersection result in low recall. For example, indicators such as position and speed are poor. Specifically for a single pole, occlusions can lead to inaccurate perception results from the single pole, which in turn affects the fusion results of multiple poles.
[0005] (2) There are post-fusion logic problems caused by the inconsistency between the laser perception results and the image perception results obtained by single-shot perception, and there are also multi-shot post-fusion logic problems caused by the inconsistency of single-shot results. Summary of the Invention
[0006] The embodiments of the present application provide a target intersection fusion method, device, electronic device, and storage medium to improve the indicators and performance of intersection and road section fusion results.
[0007] The embodiments of this application adopt the following technical solutions:
[0008] In a first aspect, an embodiment of the present application provides a target intersection fusion method, wherein the fusion method includes:
[0009] generating BEV multimodal data in response to perception results from a perception device, the perception results including a road segment result and an intersection result, the perception device including a lidar and a camera;
[0010] According to the road section results in the BEV multimodal data and the intersection results in the BEV multimodal data, the intersection results and the road section results are fused to obtain a fusion result of the entire intersection.
[0011] In some embodiments, the camera includes a first type camera and a second type camera, and generating BEV multimodal data in response to a perception result of a perception device includes:
[0012] When the perception result of the perception device includes the perception result of the first type camera, the perception result of the second type camera, and the perception result of the lidar, generating first BEV multimodal data according to the perception result of the first type camera and the perception result of the lidar;
[0013] At the same time, second BEV multimodal data is generated according to the perception results of the first type of camera, the perception results of the second type of camera, and the perception results of the lidar.
[0014] In some embodiments, the perception results include perception results of at least one intersection and perception results of multiple road segments, and generating BEV multimodal data in response to the perception results of the perception device includes:
[0015] When the perception results of the perception device include perception results of at least one intersection and perception results of multiple road sections, generating the first BEV multimodal data according to the perception results of the first type of camera and the perception results of the lidar in the perception results of the multiple road sections;
[0016] At the same time, the second BEV multimodal data is generated based on the perception results of the first type of camera, the perception results of the second type of camera, and the perception results of the lidar in the perception results of the at least one intersection.
[0017] In some embodiments, the step of fusing the intersection and segment results based on the segment results in the BEV multimodal data and the intersection results in the BEV multimodal data to obtain a fusion result of the entire intersection includes:
[0018] According to the multiple road section results in the BEV multimodal data and at least one intersection result in the BEV multimodal data, the results of an intersection and its adjacent road sections are fused to obtain a final fusion result of the entire complete intersection.
[0019] In some embodiments, the first type of camera comprises a box camera, and the second type of camera comprises a fisheye camera.
[0020] In some embodiments, before responding to the perception results of the perception device, it also includes: deploying a gun camera and a laser radar on the crossbar of the pole in directions toward the intersection and toward the road section respectively, and deploying a fisheye camera under the pole.
[0021] In some embodiments, the final fusion result of the entire intersection includes:
[0022] Calibrate each of the sensing devices into a local coordinate system;
[0023] Associating any two of the sensing devices through the local coordinate system;
[0024] According to the result of the association, the position information of the target in the local coordinate system is obtained;
[0025] Outputting the final fusion result of the entire intersection including the target's 3D detection frame, target heading angle, target speed, target category, and the target's position information in the local coordinate system; and
[0026] The position information of the target in the local coordinate system is converted to the world coordinate system to obtain the position information of the target in the world coordinate system.
[0027] In a second aspect, an embodiment of the present application further provides a target intersection fusion device, wherein the device is used to implement the method described in the first aspect.
[0028] In a third aspect, an embodiment of the present application further provides an electronic device, comprising: a processor; and a memory arranged to store computer-executable instructions, wherein the executable instructions, when executed, enable the processor to perform the above method.
[0029] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores one or more programs. When the one or more programs are executed by an electronic device including multiple application programs, the electronic device executes the above method.
[0030] At least one of the above-mentioned technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: In response to the perception results of the perception device, BEV multimodal data is generated to obtain multiple BEV multimodal data. Based on the road segment results in the BEV multimodal data and the intersection results in the BEV multimodal data, the intersection and road segment results are fused to obtain a fusion result of the entire intersection. Through the above-mentioned method, by adopting a multimodal fusion method that is split into intersections and road segments, the intersection multimodal results and the road segment multimodal results are post-fused, thereby solving the problem of single-pole occlusion toward the intersection and the indicator problem caused by multi-pole fusion logic. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0032] Figure 1 This is a schematic diagram of a scenario applicable to the target intersection fusion method in an embodiment of the present application;
[0033] Figure 2This is a flow chart of the target intersection fusion method in an embodiment of the present application;
[0034] Figure 3 This is a schematic diagram of the structure of the target intersection fusion device in an embodiment of the present application;
[0035] Figure 4 This is a structural diagram of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0036] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0037] BEV (Bird's Eye View) projects image information collected by multiple sensors into a unified 3D space. Compared to traditional camera images, BEV provides a unified space that is closer to the real physical world, providing greater convenience and possibilities for subsequent multi-sensor fusion and planning control module development.
[0038] like Figure 1 As shown, the intersection BEV multimodal data and road segment BEV multimodal data of the intersection area are included. The BEV multimodal data of one intersection and the BEV multimodal data of the adjacent road segment are combined to obtain the target intersection fusion result in the embodiment of this application. BEV multimodal data includes but is not limited to data obtained by projecting multiple perception results into the same coordinate system.
[0039] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.
[0040] The embodiment of the present application provides a target intersection fusion method, such as Figure 2 As shown, a flow chart of a target intersection fusion method according to an embodiment of the present application is provided, wherein the method comprises at least the following steps S210 to S220:
[0041] Step S210 , generating BEV multimodal data in response to perception results of a perception device, wherein the perception results include road section results and intersection results, and the perception device includes a lidar and a camera.
[0042] The target intersection fusion method in the embodiment of the present application can be provided with computing power by a roadside edge computing device, or by a cloud service device, which is not specifically limited in the embodiment of the present application.
[0043] BEV multimodal data is generated from the perception results of the perception device. Since the perception results include the perception results of intersections and road sections, the BEV multimodal data is also distinguished according to intersections and road sections.
[0044] The perception equipment is a laser radar, a gun camera and a fisheye camera deployed on a pole. Usually the fisheye camera is deployed under the pole, and the gun camera and laser radar face the direction of the incoming and outgoing traffic respectively.
[0045] It can be understood that the coverage distance of the sensing device covers multiple lanes in the horizontal direction and a certain area in the vertical direction.
[0046] Step S220 , based on the road section results in the BEV multimodal data and the intersection results in the BEV multimodal data, the intersection results and the road section results are fused to obtain a fusion result of the entire intersection.
[0047] By fusing the separate results for the road section and intersection, and then the intersection and road section results, we can obtain the final fusion result for the entire intersection. Since both the road section and intersection results are BEV multimodal data, the multimodal data processing dimensions are unified, converting multiple camera or radar data into a 3D perspective, reducing perception error. In addition, temporal information fusion is implemented in BEV multimodal data. Compared to 2D information, the 3D perspective under BEV can effectively reduce scale and occlusion issues. It can also use prior knowledge to complete occluded objects, effectively improving the accuracy of perception results.
[0048] The above method generates BEV multimodal data in response to perception results from sensing devices, including lidar and cameras, covering road sections and intersections. At intersections, the BEV fusion of multi-pole, multi-perspective laser and visual features can address issues such as single-pole occlusion at the intersection and performance issues caused by multi-pole fusion logic, improving intersection metrics such as target position, speed, heading angle, and tracking continuity.
[0049] Through the above method, based on the road section results in the BEV multimodal data and the intersection results in the BEV multimodal data, the intersection and section results are fused to obtain the final fusion result of the entire complete intersection, so that the single-pole and multi-pole laser and image perception results are consistent, while ensuring that the single-pole result does not affect the multi-pole post-fusion.
[0050] Unlike related technologies, which suffer from occlusion in areas facing intersections, resulting in low precision recall, this approach generates BEV multimodal data based on the perception results of sensing devices, improving the precision recall of multiple indicators.
[0051] In the embodiments of this application, a BEV fusion method using multi-pole, multi-view LiDAR and camera visual features at intersections can address the problem of single-pole occlusion of the intersection and the performance issues caused by multi-pole fusion logic, thereby improving multiple or single indicators such as position, speed, heading angle, and continuity at the intersection. This solves the problem of inaccurate perception results from a single pole due to occlusion.
[0052] In the embodiments of the present application, a multimodal algorithm for front fusion is used to solve the post-fusion logic problem caused by the inconsistency of single-shot and multi-shot laser and image perception results. At the same time, the post-fusion logic problem of multiple shots caused by the inconsistency of single-shot results in the front fusion process is also solved, thereby reducing the difficulty of post-fusion of single and multi-shot.
[0053] In the embodiments of the present application, the post-fusion logic problem caused by the inconsistency between the single-shot laser and image perception results, as well as the multi-shot post-fusion logic problem caused by the inconsistency of the single-shot results in the pre-fusion process, are also solved.
[0054] In one embodiment of the present application, the camera includes a first type of camera and a second type of camera, and the BEV multimodal data is generated in response to the perception results of the perception device, including: when the perception results of the perception device include the perception results of the first type of camera, the perception results of the second type of camera, and the perception results of the lidar, the first BEV multimodal data is generated according to the perception results of the first type of camera and the perception results of the lidar; at the same time, the second BEV multimodal data is generated according to the perception results of the first type of camera, the perception results of the second type of camera, and the perception results of the lidar.
[0055] Because different types of cameras are deployed on the pole, first BEV multimodal data is generated based on the perception results of the first type of camera and the lidar. This BEV multimodal data corresponds to the perception results of the road section, and the second type of camera is not required for the perception results of the road section. Second BEV multimodal data is then generated based on the perception results of the first type of camera, the second type of camera, and the lidar. This BEV multimodal data corresponds to the perception results of the intersection, and the two types of cameras are required for the perception results of the intersection.
[0056] It is understood that BEV multimodal data includes, but is not limited to, multimodal fusion algorithms for intersections and road sections. Specifically, the inputs to the multimodal fusion algorithm for intersections include multi-channel lidar perception results, multi-channel gun camera perception results, and fisheye camera perception results. After time synchronization, image and point cloud processing preprocessing is performed, followed by image feature and point cloud feature extraction to obtain point cloud image feature fusion. A multi-task head model (including 3D box, category, heading angle, etc.) is used to obtain the post-fusion result. Similarly, the inputs to the multimodal fusion algorithm for specific road sections include multi-channel lidar perception results and multi-channel gun camera perception results. The fusion process is the same as that for intersections, so it will not be repeated here.
[0057] In one embodiment of the present application, the perception results include perception results of at least one intersection and perception results of multiple road sections, and the BEV multimodal data is generated in response to the perception results of the perception device, including: when the perception results of the perception device include perception results of at least one intersection and perception results of multiple road sections, the first BEV multimodal data is generated based on the perception results of the first type of camera and the perception results of the lidar in the perception results of the multiple road sections; at the same time, the second BEV multimodal data is generated based on the perception results of the first type of camera and the perception results of the second type of camera and the perception results of the lidar in the perception results of the at least one intersection.
[0058] Since the perception results include those for intersections and road sections, in the case of perception results for at least one intersection and multiple road sections, the perception results from the first type of camera and the lidar from the perception results for the multiple road sections are used as first BEV multimodal data. Similarly, the perception results from the first type of camera, the second type of camera, and the lidar from the perception results for the at least one intersection are used as second BEV multimodal data.
[0059] It should be noted that the distinction between at least one intersection and multiple road sections is for illustrative purposes only and is not intended to limit the scope of protection of this application. It is generally believed that intersection scenarios are more complex and variable, while road section scenarios are more simple and regular. Therefore, the multimodal perception results of BEVs using different types of cameras and lidars were selected.
[0060] In one embodiment of the present application, the intersection and section results are fused based on the road section results in the BEV multimodal data and the intersection results in the BEV multimodal data to obtain the final fusion result of the entire complete intersection, including: based on multiple road section results in the BEV multimodal data and at least one intersection result in the BEV multimodal data, the results between an intersection and its adjacent road sections are fused to obtain the final fusion result of the entire complete intersection.
[0061] A scene that needs to be fused includes at least one intersection and the fusion results between its adjacent road segments, such as Figure 1 As shown in the figure, the BEV multimodal data includes the fusion of the results of an intersection and its adjacent road sections to obtain the final fusion result of the entire intersection. The specific method is post-fusion. Post-fusion refers to the fusion of target detection results processed by different sensors, rather than fusion at the data layer or feature layer.
[0062] In one embodiment of the present application, the first type of camera includes a box camera, and the second type of camera includes a fisheye camera.
[0063] Specifically, a gun camera, fisheye camera, and lidar can be deployed on the same pole. Usually, in an intersection scenario, the fusion results of all poles arranged at the intersection are included.
[0064] In one embodiment of the present application, before responding to the perception results of the perception device, it also includes: deploying a gun camera and a laser radar on the crossbar of the pole in the directions toward the intersection and toward the road section respectively, and deploying a fisheye camera under the pole.
[0065] In one embodiment of the present application, taking a crossroads as an example, please refer to Figure 1 The intersection includes north, south, west, and east poles, each equipped with a gun camera and fisheye camera. The gun camera adjusts its direction according to the intersection or road section orientation. The fisheye camera is located below the pole, and the lidar also adjusts its direction according to the intersection or road section orientation. The lidar has a longitudinal unidirectional coverage range of 20-120 meters, the fisheye camera has a longitudinal unidirectional coverage range of 0-20 meters, and the gun camera has a longitudinal unidirectional coverage range of 10-120 meters. The final fusion result is obtained by obtaining the segment BEV and the intersection BEV at the intersection or road section, and then performing post-fusion.
[0066] In one embodiment of the present application, the final fusion result of the entire complete intersection includes: calibrating each of the perception devices to the local coordinate system; associating any two of the perception devices through the local coordinate system; obtaining the position information of the target in the local coordinate system based on the association result; outputting the final fusion result of the entire complete intersection including the target's 3D detection frame, target heading angle, target speed, target category and the position information of the target in the local coordinate system; and converting the position information of the target in the local coordinate system to the world coordinate system to obtain the position information of the target in the world coordinate system.
[0067] The local coordinate system can be defined based on the actual use case. Typically, each sensing device is calibrated to the local coordinate system. Any two sensing devices are then linked using the local coordinate system. Based on the linking results, the target's position information in the local coordinate system is obtained. The target fusion results include, but are not limited to: {the target's 3D detection bounding box, target heading angle, target velocity, target category, and the target's position information in the local coordinate system}. The output target can then be tracked.
[0068] Preferably, the position information of the target in the local coordinate system is converted to the world coordinate system, and the position information of the target in the world coordinate system is obtained for visual display of the target position information in the digital twin system.
[0069] The present application also provides a target intersection fusion device 300. As shown in FIG. Target intersection fusion device, a schematic diagram of the structure of the target intersection fusion device in the present application is provided. The target intersection fusion device 300 includes at least: a response module 310 and a fusion module 320, wherein:
[0070] In one embodiment of the present application, the response module 310 is specifically used to generate BEV multimodal data in response to the perception results of the perception device, wherein the perception results include road section results and intersection results, and the perception device includes a lidar and a camera.
[0071] The target intersection fusion method in the embodiment of the present application can be provided with computing power by a roadside edge computing device, or by a cloud service device, which is not specifically limited in the embodiment of the present application.
[0072] BEV multimodal data is generated from the perception results of the perception device. Since the perception results include the perception results of intersections and road sections, the BEV multimodal data is also distinguished according to intersections and road sections.
[0073] The perception equipment is a laser radar, a gun camera and a fisheye camera deployed on a pole. Usually the fisheye camera is deployed under the pole, and the gun camera and laser radar face the direction of the incoming and outgoing traffic respectively.
[0074] In one embodiment of the present application, the fusion module 320 is specifically used to: fuse the intersection results and the road section results according to the road section results in the BEV multimodal data and the intersection results in the BEV multimodal data to obtain the fusion result of the entire intersection.
[0075] By fusing the separate results for the road section and intersection, and then the intersection and road section results, we can obtain the final fusion result for the entire intersection. Since both the road section and intersection results are BEV multimodal data, the multimodal data processing dimensions are unified, converting multiple camera or radar data into a 3D perspective, reducing perception error. In addition, temporal information fusion is implemented in BEV multimodal data. Compared to 2D information, the 3D perspective under BEV can effectively reduce scale and occlusion issues. It can also use prior knowledge to complete occluded objects, effectively improving the accuracy of perception results.
[0076] In one embodiment of the present application, the camera includes a first type camera and a second type camera, and the response module 310 is further configured to:
[0077] When the perception result of the perception device includes the perception result of the first type camera, the perception result of the second type camera, and the perception result of the lidar, generating first BEV multimodal data according to the perception result of the first type camera and the perception result of the lidar;
[0078] At the same time, second BEV multimodal data is generated according to the perception results of the first type of camera, the perception results of the second type of camera, and the perception results of the lidar.
[0079] In one embodiment of the present application, the perception result includes a perception result of at least one intersection and a perception result of multiple road sections. The response module 310 is further configured to:
[0080] When the perception results of the perception device include perception results of at least one intersection and perception results of multiple road sections, generating the first BEV multimodal data according to the perception results of the first type of camera and the perception results of the lidar in the perception results of the multiple road sections;
[0081] At the same time, the second BEV multimodal data is generated based on the perception results of the first type of camera, the perception results of the second type of camera, and the perception results of the lidar in the perception results of the at least one intersection.
[0082] In one embodiment of the present application, the fusion module 320 is further configured to:
[0083] According to the multiple road section results in the BEV multimodal data and at least one intersection result in the BEV multimodal data, the results of an intersection and its adjacent road sections are fused to obtain a final fusion result of the entire complete intersection.
[0084] In one embodiment of the present application, the first type of camera includes a box camera, and the second type of camera includes a fisheye camera.
[0085] In one embodiment of the present application, before responding to the perception result of the perception device, the method further includes:
[0086] On the crossbar of the pole, a gun camera and a laser radar are respectively deployed in the direction of the intersection and the road section, and a fisheye camera is deployed under the pole. In one embodiment of the present application, the fusion module 320 is further used to:
[0087] Calibrate each of the sensing devices into a local coordinate system;
[0088] Associating any two of the sensing devices through the local coordinate system;
[0089] According to the result of the association, the position information of the target in the local coordinate system is obtained;
[0090] Outputting the final fusion result of the entire intersection including the target's 3D detection frame, target heading angle, target speed, target category, and the target's position information in the local coordinate system; and
[0091] The position information of the target in the local coordinate system is converted to the world coordinate system to obtain the position information of the target in the world coordinate system.
[0092] It can be understood that the above-mentioned target intersection fusion device can implement each step of the target intersection fusion method provided in the above-mentioned embodiment. The relevant explanations about the target intersection fusion method are applicable to the target intersection fusion device and will not be repeated here.
[0093] Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. Figure 4 At the hardware level, the electronic device includes a processor and, optionally, an internal bus, a network interface, and memory. The memory may include internal memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for its services.
[0094] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0095] The memory is used to store programs. Specifically, the program may include program code, which includes computer operating instructions. The memory may include internal memory and non-volatile memory, and provides instructions and data to the processor.
[0096] The processor reads the corresponding computer program from the non-volatile memory into the internal memory and then runs it, forming the target intersection fusion device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:
[0097] generating BEV multimodal data in response to perception results from a perception device, the perception results including road segments and intersections, the perception device including a lidar and a camera;
[0098] According to the road section results in the BEV multimodal data and the intersection results in the BEV multimodal data, the intersection and road section results are fused to obtain a final fusion result of the entire complete intersection.
[0099] The above application Figure 1The methods performed by the target intersection fusion device disclosed in the illustrated embodiments can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor or by software instructions. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0100] The electronic device may also perform Figure 2 The target intersection fusion device executes the method, and realizes the target intersection fusion device in Figure 2 The functions of the illustrated embodiment will not be described in detail in the embodiments of the present application.
[0101] The embodiment of the present application also provides a computer-readable storage medium, which stores one or more programs, wherein the one or more programs include instructions, which, when executed by an electronic device including multiple application programs, can enable the electronic device to execute Figure 2 The method performed by the target intersection fusion device in the illustrated embodiment is specifically used to perform:
[0102] generating BEV multimodal data in response to perception results from a perception device, the perception results including road segments and intersections, the perception device including a lidar and a camera;
[0103] According to the road section results in the BEV multimodal data and the intersection results in the BEV multimodal data, the intersection and road section results are fused to obtain a final fusion result of the entire complete intersection.
[0104] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0105] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0106] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0107] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0108] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0109] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0110] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0111] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0112] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0113] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A target intersection fusion method, wherein: The fusion method comprises: generating BEV multimodal data in response to perception results from a perception device, the perception results including a road segment result and an intersection result, the perception device including a lidar and a camera; According to the road section results in the BEV multimodal data and the intersection results in the BEV multimodal data, the intersection results and the road section results are fused to obtain a fusion result of the entire intersection.
2. The method according to claim 1, wherein: The camera includes a first type camera and a second type camera, and generating BEV multimodal data in response to a perception result of a perception device includes: When the perception result of the perception device includes the perception result of the first type camera, the perception result of the second type camera, and the perception result of the lidar, generating first BEV multimodal data according to the perception result of the first type camera and the perception result of the lidar; At the same time, second BEV multimodal data is generated according to the perception results of the first type of camera, the perception results of the second type of camera, and the perception results of the lidar.
3. The method according to claim 2, wherein: The perception results include perception results of at least one intersection and perception results of multiple road sections. The generating of BEV multimodal data in response to the perception results of the perception device includes: When the perception results of the perception device include perception results of at least one intersection and perception results of multiple road sections, generating the first BEV multimodal data according to the perception results of the first type of camera and the perception results of the lidar in the perception results of the multiple road sections; At the same time, the second BEV multimodal data is generated based on the perception results of the first type of camera, the perception results of the second type of camera, and the perception results of the lidar in the perception results of the at least one intersection.
4. The method according to claim 3, wherein: The step of fusing the intersection and section results based on the section results in the BEV multimodal data and the intersection results in the BEV multimodal data to obtain a fusion result of the entire intersection includes: According to the multiple road section results in the BEV multimodal data and at least one intersection result in the BEV multimodal data, the results of an intersection and its adjacent road sections are fused to obtain a final fusion result of the entire complete intersection.
5. The method according to any one of claims 2 to 4, wherein: The first type of camera includes a box camera, and the second type of camera includes a fisheye camera.
6. The method of claim 5, wherein: Before responding to the perception result of the perception device, it also includes: On the horizontal bar of the pole, a gun camera and a lidar are deployed respectively in the direction of the intersection and the road section, and a fisheye camera is deployed under the pole.
7. The method of claim 1, wherein: The final fusion result of the entire intersection includes: Calibrate each of the sensing devices into a local coordinate system; Associating any two of the sensing devices through the local coordinate system; According to the result of the association, the position information of the target in the local coordinate system is obtained; Outputting the final fusion result of the entire intersection including the target's 3D detection frame, target heading angle, target speed, target category, and the target's position information in the local coordinate system; and The position information of the target in the local coordinate system is converted to the world coordinate system to obtain the position information of the target in the world coordinate system.
8. A target intersection fusion device, wherein: The device is used to implement the method according to any one of claims 1 to 7.
9. An electronic device comprising: processor; as well as A memory arranged to store computer executable instructions, which when executed cause the processor to perform the method of any one of claims 1 to 7.
10. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of application programs, causes the electronic device to execute the method according to any one of claims 1 to 7.