A target detection method, system, and vehicle based on multimodal pre-fusion
By synchronously acquiring and associating image data frames and laser point cloud data at the sensor end, the problem of information loss caused by sensor post-fusion is solved, the accuracy and timeliness of target detection are improved, and the safety of the autonomous driving system is ensured.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG ZEEKR INTELLIGENT TECH CO LTD
- Filing Date
- 2026-01-21
- Publication Date
- 2026-05-26
AI Technical Summary
In existing technologies, post-fusion of sensors leads to the loss of some original information, affecting the accuracy of target detection. This is especially true in complex scenarios where it can easily result in missed detections or misjudgments, posing a safety hazard.
A multimodal pre-fusion target detection method is adopted, which synchronously acquires image data frames and laser point cloud data through the eye-sensor module, performs correlation processing at the sensor end to generate correlated data frames and uncorrelated data frames, and outputs them to the domain controller for target detection.
By associating image data frames and laser point cloud data at the sensor end, link time is reduced, the integrity of sensor data is ensured, and the accuracy and timeliness of target detection are improved.
Smart Images

Figure CN122090413A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving perception technology, and more specifically, to a target detection method, system, and vehicle based on multimodal front fusion. Background Technology
[0002] In autonomous driving systems, to build accurate and reliable environmental perception capabilities, it is usually necessary to integrate information from multiple sensors such as LiDAR and cameras.
[0003] In related technologies, the fusion schemes for multiple sensors mentioned above are still mainly based on "post-fusion," which involves performing target detection on the data from each sensor separately and then associating and fusing the target detection results. This post-fusion method is widely used because it is relatively simple to implement, but it also has obvious limitations: since the association is performed at the target level, a large amount of fine-grained and complementary information in the original sensor data is lost during the fusion process, affecting the accuracy of the target detection results. As a result, the autonomous driving system cannot provide a complete and accurate perception environment in all scenarios, especially when facing some unconventional or long-tail scenarios, which are prone to missed detections or misjudgments. For example, in situations such as a truck tire falling off on a highway at night, a vehicle carrying an extra-long bamboo pole or steel pipe, or a trailer tail that is difficult to identify, the post-fusion scheme often cannot perceive and identify in a timely and accurate manner, thus creating potential safety hazards. Summary of the Invention
[0004] The problem addressed by this invention is how to solve the problem of loss of some original information due to sensor post-fusion, which in turn affects the accuracy of target detection.
[0005] To address the aforementioned issues, this invention provides a target detection method, system, and vehicle based on multimodal pre-fusion.
[0006] In a first aspect, the present invention provides a target detection system based on multimodal pre-fusion, comprising at least one eye-sensing sensor module and a domain controller; The laser sensor module is used to synchronously acquire image data frames and laser point cloud data for the same area, associate the image data frames and laser point cloud data acquired at the same acquisition timestamp to generate associated data frames, and output the unassociated image data frames and associated data frames to the domain controller. The domain controller is used to perform target detection based on the unassociated image data frames and the associated data frames.
[0007] Optionally, the vision sensor module is further configured to: At least two sub-image data frames targeting the same area are acquired simultaneously, wherein the at least two sub-image data frames originate from different viewpoints; The at least two sub-image data frames are registered and stitched together into a single image data frame with unified spatial coordinates.
[0008] Optionally, the step of associating the image data frames acquired at the same acquisition timestamp with the laser point cloud data to generate associated data frames includes: For each pixel in the image data frame, the nearest laser reflection point in the laser point cloud data is found in space, and the pixel is associated with the found laser reflection point to obtain the associated data frame.
[0009] Optionally, outputting the unassociated image data frame and the associated data frame to the domain controller includes: The laser sensor module alternately outputs the unassociated image data frames and the associated data frames to the domain controller in an n-1:1 ratio; where n is the ratio of the acquisition frequency of the image data frames to the acquisition frequency of the laser point cloud data, and n is a positive integer greater than or equal to 2.
[0010] Optionally, the domain controller includes: a multi-dimensional data fusion chip and a perception-decision chip; The target detection based on the unrelated image data frames and the related data frames includes: The multi-dimensional data fusion chip performs fusion processing on the associated pixels and laser reflection points in the associated data frame to generate a multimodal data frame, and outputs the unassociated image data frame and the multimodal data frame to the perception-decision chip; wherein, each multi-dimensional pixel in the multimodal data frame includes color channel information, distance information, velocity information, angular velocity information and acceleration information; The perception-decision chip is used to perform target detection based on the unassociated image data frames and the multimodal data frames.
[0011] Optionally, the multidimensional data fusion chip outputs the unrelated image data frames and the multimodal data frames to the perception-decision chip, including: The multidimensional data fusion chip alternately outputs the unrelated image data frames and the multimodal data frames to the perception-decision chip in an n-1:1 ratio; where n is the ratio of the acquisition frequency of the image data frames to the acquisition frequency of the laser point cloud data, and n is a positive integer greater than or equal to 2.
[0012] Optionally, the perception-decision chip is used to perform target detection based on the unassociated image data frames and the multimodal data frames, including: The unassociated image data frames and the multimodal data frames are input into the target detection model to obtain the 3D target detection result output by the target detection model.
[0013] Optionally, the eye-sensing sensor module transmits the unassociated image data frames and the associated data frames to the multidimensional data fusion chip via Ethernet.
[0014] Secondly, this invention provides a target detection method based on multimodal pre-fusion, applied to a target detection system, the target detection system including at least one gaze sensor module and a domain controller; the target detection method includes: The laser sensor module synchronously acquires image data frames and laser point cloud data for the same area, associates the image data frames and laser point cloud data acquired at the same acquisition timestamp to generate associated data frames, and outputs the unassociated image data frames and associated data frames to the domain controller. The domain controller performs target detection based on the unassociated image data frames and the associated data frames.
[0015] Thirdly, the present invention provides a vehicle comprising: a target detection system based on multimodal pre-fusion as described in any of the preceding claims.
[0016] The beneficial effects of the target detection method, system, and vehicle based on multimodal pre-fusion of the present invention are as follows: The eye-sensing sensor module synchronously acquires image data frames and laser point cloud data for the same area. Image data frames acquired at the same acquisition timestamp are correlated with laser point cloud data to generate correlated data frames. Uncorrelated image data frames and correlated data frames are output to the domain controller. Correlation of image data frames and laser point cloud data is performed at the sensor end, reducing link time and ensuring no loss of original information, thus guaranteeing the integrity of the sensing data. The domain controller performs target detection based on the uncorrelated image data frames and correlated data frames. The domain controller can perform target detection based on more accurate and timely correlated data frames, improving the accuracy of target detection. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the structure of a target detection system based on multimodal pre-fusion according to an embodiment of the present invention; Figure 2(a) is a schematic diagram of the left-side vision sensor module; Figure 2(b) is a schematic diagram of the right-side vision sensor module; Figure 2(c) is a schematic diagram of the front-side vision sensor module. Figure 3 This is a schematic diagram of the structure of a domain controller according to one embodiment; Figure 4This is a spatial schematic diagram of each pixel in an image data frame and a multimodal data frame according to one embodiment; Figure 5 This is a flowchart of a target detection method based on multimodal pre-fusion according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0018] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0019] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0020] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to"; the term "based on" means "at least partially based on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; and the term "optionally" means "optional embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first," "second," etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0021] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0022] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0023] like Figure 1As shown in the figure, an embodiment of the present invention provides a target detection system based on multimodal front fusion, comprising: at least one eye-sensing sensor module 110 and a domain controller 120.
[0024] The LiDAR sensor module 110 is used to synchronously acquire image data frames and laser point cloud data for the same area, associate the image data frames and laser point cloud data acquired at the same acquisition timestamp to generate associated data frames, and output unassociated image data frames and associated data frames to the domain controller 120.
[0025] Taking a vehicle as an example, a vehicle typically includes eight sensors: a front-view camera, a front-view LiDAR, a left front camera, a left rear camera, a left side LiDAR, a right front camera, a right rear camera, and a right side LiDAR. For example, these eight sensors can be divided into three groups according to the left, front, and right sides of the vehicle: the front-view camera and front-view LiDAR; the left front camera, left rear camera, and left side LiDAR; and the right front camera, right rear camera, and right side LiDAR.
[0026] For the three sets of sensors mentioned above, this embodiment integrates each set of sensors into a single hardware device, namely a laser-driven sensor module. This module can simultaneously acquire image frame data and laser point cloud data for the same area. For example, for the left-side area, the left-side laser-driven sensor module can acquire image frame data and laser point cloud data for the left-side area (left front and left rear). For ease of description, the three laser-driven sensor modules after the integration of the three sets of sensors are referred to as the left laser-driven sensor module, the right laser-driven sensor module, and the front laser-driven sensor module, as shown in Figure 2. Figure 2(a) shows that the left laser-driven sensor module integrates the left front camera, the left rear camera, and the left-side LiDAR. Figure 2(b) shows that the right laser-driven sensor module integrates the right front camera, the right rear camera, and the right-side LiDAR. Figure 2(c) shows that the front laser-driven sensor module integrates the front-view camera and the front-view LiDAR.
[0027] Specifically, the left-view sensor module, the right-view sensor module, and the front-view sensor module receive unified timing data from the domain controller 120 for synchronous data acquisition, so as to avoid the spatiotemporal dimensions of the final detected target object being different due to inconsistent timing and sensor acquisition frequencies.
[0028] In some embodiments, the vision sensor module 110 is further configured to: synchronously acquire at least two sub-image data frames for the same area, wherein the at least two sub-image data frames are from different viewpoints; and register and stitch the at least two sub-image data frames into a single image data frame with unified spatial coordinates.
[0029] Taking the left-side vision sensor module as an example, its synchronously acquired image data frames are the front left image data frame and the rear left image data frame of the left side of the vehicle. Therefore, the front left image data frame and the rear left image data frame need to be registered and stitched together to form a single left-side image data frame with unified spatial coordinates. This stitched left-side image data frame is then associated with the laser point cloud data of the left side of the vehicle acquired by the left-side vision sensor module at the same acquisition timestamp to generate an associated data frame. Since the acquisition frequencies of the image data frames and the laser point cloud data are different, and the acquisition frequency of the image data frames is often greater than that of the laser point cloud data (e.g., the acquisition frequency of the image data frames is 30Hz, and the acquisition frequency of the laser point cloud data is 10Hz), out of every three image data frames, one image data frame is associated with the laser point cloud data to form one associated data frame, while the other two image data frames are not associated. The vision sensor module 110 finally outputs the unassociated image data frame and the associated data frame to the domain controller 120.
[0030] In some embodiments, the image data frame is a video stream, and the associated data frame is an associated video stream.
[0031] In some embodiments, the LiDAR sensor module 110 outputs unassociated image data frames and associated data frames to the Ethernet interface of the domain controller 120 via Ethernet.
[0032] Domain controller 120 is used for target detection based on unassociated image data frames and associated data frames.
[0033] Specifically, the Domain Controller Unit (ADCU) 120 is the autonomous driving domain controller for the vehicle. It is responsible for aggregating data from various sensors and using built-in computing chips and algorithms to achieve accurate perception, fusion understanding, path planning, and driving control of the vehicle's surrounding environment.
[0034] In some embodiments, the domain controller 120 performs target detection on unassociated image data frames and associated data frames respectively. For example, the unassociated image data frames and associated data frames are input into a pre-built target detection model to obtain the target detection results output by the target detection model. It should be noted that in this embodiment, both the target detection results obtained from unassociated image data frames and the target detection results obtained from associated data frames can be input into the subsequent decision analysis of the intelligent driving system. In the decision analysis of the intelligent driving system, if only the target detection results from unassociated image frames are relied upon, the limited perceptual information and lack of depth and 3D verification may lead to biases in the judgment of obstacle position, size, or category, thus affecting the accuracy of the planning and control module's analysis. Detection results based on associated data frames, however, provide a higher confidence level of perception output due to their combination of texture and precise 3D structural information. Therefore, the system can compare and fuse these two target detection results, using the detection results from associated data frames to verify and correct errors that may arise from the detection of unassociated image frames, thereby improving the reliability of the overall environmental perception.
[0035] In this embodiment, the laser sensor module 110 simultaneously acquires image data frames and laser point cloud data for the same area. It correlates the image data frames acquired at the same acquisition timestamp with the laser point cloud data to generate correlated data frames, and outputs both uncorrelated image data frames and correlated data frames to the domain controller. This correlation between image data frames and laser point cloud data at the sensor end reduces link time and ensures no loss of original information, thus guaranteeing the integrity of the sensor data. The domain controller 120 performs target detection based on the uncorrelated image data frames and correlated data frames. The domain controller can perform target detection based on more accurate and timely correlated data frames, improving the accuracy of target detection.
[0036] Optionally, the laser sensor module will perform association processing on the image data frame and the laser point cloud data acquired at the same acquisition timestamp to generate an associated data frame, including: for each pixel in the image data frame, finding the laser reflection point closest to the pixel in the laser point cloud data in space, and associating the pixel with the found laser reflection point to obtain the associated data frame.
[0037] Specifically, for image data frames and laser point cloud data under the same acquisition timestamp, there may be some pixels that cannot be associated with laser reflection points. These unassociated pixels will be formed as unassociated pixels in the associated data frames.
[0038] In this optional embodiment, by associating pixels with laser reflection points in a spatiotemporal manner, a correlated data frame with complete information alignment is generated at the sensor end, effectively avoiding perceptual distortion caused by asynchronous and lost information during post-fusion.
[0039] Optionally, the laser sensor module 110 alternately outputs unassociated image data frames and associated data frames to the domain controller 120 in an n-1:1 ratio; where n is the ratio of the acquisition frequency of the image data frames to the acquisition frequency of the laser point cloud data, and n is a positive integer greater than or equal to 2.
[0040] For example, if the acquisition frequency of image data frames is 30Hz and the acquisition frequency of laser point cloud data is 10Hz, then n is 3. The laser sensor module 110 outputs data to the domain controller 120 via Ethernet according to 2 unassociated image data frames and 1 associated data frame.
[0041] In this optional embodiment, image data frames and associated data frames are transmitted alternately at a determined timing ratio (n-1:1). While ensuring that associated data frames are stably output to the domain controller 120, high frame rate image data frames are also taken into account. This allows the domain controller 120 to both use high-frequency image data frames for rapid response to environmental changes and obtain accurate geometric perception based on synchronously associated data frames, thereby achieving an efficient balance between system bandwidth and perception accuracy.
[0042] Optionally, such as Figure 3 As shown, the domain controller 120 includes a multi-dimensional data fusion chip 121 and a perception-decision chip 122. Furthermore, the domain controller 120 also includes an Ethernet interface 123 and a gateway 124, wherein the multi-dimensional data fusion chip 121 is disposed between the Ethernet interface 123 and the gateway 124.
[0043] The eye-sensor module 110 outputs unassociated image data frames and associated data frames alternately to the Ethernet interface 123 via Ethernet in an n-1:1 ratio. The Ethernet interface 123 then transmits the unassociated image data frames and associated data frames to the multi-dimensional data fusion chip 121.
[0044] The multi-dimensional data fusion chip 121 is used to fuse associated pixels and laser reflection points in associated data frames to generate multi-modal data frames, and output unassociated image data frames and multi-modal data frames to the perception-decision chip; wherein, each multi-dimensional pixel in the multi-modal data frame includes color channel information, distance information, velocity information, angular velocity information and acceleration information.
[0045] Specifically, such as Figure 4The above describes the spatial diagrams of each pixel in the image data frame and each pixel in the multimodal data frame. Each pixel in the image data frame includes color channel information (R / G / B), while each pixel in the multimodal data frame includes color channel information (R / G / B), distance information (L), velocity information (L), angular velocity information (H), and acceleration information (A). It can be seen that each pixel in the multimodal data frame has a 7-dimensional array of data, compared to the 3-dimensional array of data (R / G / B) in the image data frame, thus possessing more detailed information.
[0046] In some embodiments, the multidimensional data fusion chip 121 alternately outputs the unassociated image data frames and the associated data frames to the perception-decision chip 122 through the gateway 124 in an n-1:1 ratio; where n is the ratio of the acquisition frequency of the image data frames to the acquisition frequency of the laser point cloud data, and n is a positive integer greater than or equal to 2.
[0047] The perception-decision chip 122 is used for target detection based on unrelated image data frames and multimodal data frames.
[0048] Specifically, the perception-decision chip 122 is a core computing chip used in fields such as autonomous driving. It can process uncorrelated data from multiple sensors, such as cameras and radar, in real time and complete perception fusion and decision-making. The perception-decision chip 122 can be NVIDIA's Orin-X chip or Thore chip.
[0049] In some embodiments, the perception-decision chip 122 has a pre-set target detection model, which can be constructed from a Transformer model, an MLP model, or a VLA model.
[0050] Specifically, the perception-decision chip 122 inputs both unrelated image data frames and multimodal data frames into the target detection model, and the target detection model outputs three-dimensional target detection results respectively. That is, it outputs the corresponding 3D target detection results based on the unrelated image data frames and the corresponding 3D target detection results based on the multimodal data frames.
[0051] In this optional embodiment, in subsequent autonomous driving analysis, since the 3D target detection results corresponding to the multimodal data frame output have more dimensional information, the 3D target detection results corresponding to the output of unrelated image data frames can be corrected to obtain more accurate analysis results without affecting the response speed; furthermore, since both image data frames and multimodal data frames are 2D data, the target detection model converts the 2D data into 3D data, resulting in higher target detection accuracy.
[0052] like Figure 5As shown in the figure, an embodiment of the present invention provides a target detection method based on multimodal pre-fusion, which is applied to a target detection system. The target detection system includes at least one eye-sensing sensor module 110 and a domain controller 120. The target detection method includes the following steps: Step S510: The LiDAR sensor module 110 synchronously acquires image data frames and laser point cloud data for the same area, performs association processing on the image data frames and laser point cloud data acquired at the same acquisition timestamp, generates associated data frames, and outputs unassociated image data frames and associated data frames to the domain controller 120.
[0053] Step S520: Domain controller 120 performs target detection based on unassociated image data frames and associated data frames.
[0054] Optionally, the target detection method based on multimodal pre-fusion also includes: The eye-sensor module 110 synchronously acquires at least two sub-image data frames for the same area, wherein the at least two sub-image data frames come from different viewpoints; The eye-sensor module 110 registers and stitches at least two sub-image data frames into one image data frame with unified spatial coordinates.
[0055] Optionally, the laser sensor module 110 correlates image data frames acquired at the same acquisition timestamp with laser point cloud data to generate correlated data frames, including: For each pixel in the image data frame, find the nearest laser reflection point in the laser point cloud data in space, and associate the pixel with the found laser reflection point to obtain the associated data frame.
[0056] Optionally, the eye-sensor module 110 outputs unassociated image data frames and associated data frames to the domain controller 120, including: The laser sensor module 110 alternately outputs unassociated image data frames and associated data frames to the domain controller 120 in an n-1:1 ratio; where n is the ratio of the acquisition frequency of the image data frames to the acquisition frequency of the laser point cloud data, and n is a positive integer greater than or equal to 2.
[0057] Optionally, the domain controller 120 includes: a multi-dimensional data fusion chip 121 and a perception-decision chip 122; Domain controller 120 performs target detection based on unassociated image data frames and associated data frames, including: The multi-dimensional data fusion chip 121 performs fusion processing on associated pixels and laser reflection points in the associated data frame to generate a multi-modal data frame, and outputs the unassociated image data frame and the multi-modal data frame to the perception-decision chip 122; wherein, each multi-dimensional pixel in the multi-modal data frame includes color channel information, distance information, velocity information, angular velocity information and acceleration information; The perception-decision chip 122 is used for target detection based on unrelated image data frames and multimodal data frames.
[0058] Optionally, the multi-dimensional data fusion chip 121 outputs unrelated image data frames and multimodal data frames to the perception-decision chip 122, including: The multidimensional data fusion chip 121 alternately outputs unrelated image data frames and multimodal data frames to the perception-decision chip in an n-1:1 ratio; where n is the ratio of the acquisition frequency of the image data frames to the acquisition frequency of the laser point cloud data, and n is a positive integer greater than or equal to 2.
[0059] Optionally, the perception-decision chip 122 is used for target detection based on unrelated image data frames and multimodal data frames, including: Unrelated image data frames and multimodal data frames are input into the target detection model to obtain the 3D target detection results output by the target detection model.
[0060] Optionally, the eye-sensor module 110 transmits unassociated image data frames and associated data frames to the multidimensional data fusion chip 121 via Ethernet.
[0061] like Figure 6 As shown in the figure, an electronic device 600 provided in this embodiment of the invention includes a memory 610 and a processor 620; the memory 610 is used to store a computer program; the processor 620 is used to implement the target detection method based on multimodal pre-fusion as described above when the computer program is executed.
[0062] Alternatively, an electronic device 600 includes a memory 610 and a processor 620 coupled to the memory 610; the memory 610 is configured to store a computer program; and the processor 620 is configured to perform the following operations when the computer program is executed: The Lishi sensor module synchronously acquires image data frames and laser point cloud data for the same area. It performs association processing on the image data frames and laser point cloud data acquired at the same acquisition timestamp to generate associated data frames, and outputs unassociated image data frames and associated data frames to the domain controller. The domain controller performs target detection based on unassociated image data frames and associated data frames.
[0063] This invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the target detection method based on multimodal pre-fusion as described above.
[0064] Alternatively, a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the following operations: The Lishi sensor module synchronously acquires image data frames and laser point cloud data for the same area. It performs association processing on the image data frames and laser point cloud data acquired at the same acquisition timestamp to generate associated data frames, and outputs unassociated image data frames and associated data frames to the domain controller. The domain controller performs target detection based on unassociated image data frames and associated data frames.
[0065] This invention also provides a vehicle, including the multimodal front fusion-based target detection system described in the above embodiments.
[0066] Electronic device 600, which can serve as a server or client of the present invention, is described below as an example of a hardware device applicable to various aspects of the present invention. Electronic device 600 is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic device 600 can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0067] Electronic device 600 includes a computing unit that can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) or a computer program loaded from a storage unit into random access memory (RAM). The RAM may also store various programs and data required for device operation. The computing unit, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0068] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc. In this application, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of the present invention according to actual needs. Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units can be implemented in hardware or as software functional units.
[0069] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.
Claims
1. A target detection system based on multimodal pre-fusion, characterized in that, Includes at least one vision sensor module and a domain controller; The laser sensor module is used to synchronously acquire image data frames and laser point cloud data for the same area, associate the image data frames and laser point cloud data acquired at the same acquisition timestamp to generate associated data frames, and output the unassociated image data frames and associated data frames to the domain controller. The domain controller is used to perform target detection based on the unassociated image data frames and the associated data frames.
2. The target detection system based on multimodal pre-fusion according to claim 1, characterized in that, The vision sensor module is also used for: At least two sub-image data frames targeting the same area are acquired simultaneously, wherein the at least two sub-image data frames originate from different viewpoints; The at least two sub-image data frames are registered and stitched together into a single image data frame with unified spatial coordinates.
3. The target detection system based on multimodal pre-fusion according to claim 1, characterized in that, The step of associating the image data frame acquired at the same acquisition timestamp with the laser point cloud data to generate an associated data frame includes: For each pixel in the image data frame, the nearest laser reflection point in the laser point cloud data is found in space, and the pixel is associated with the found laser reflection point to obtain the associated data frame.
4. The target detection system based on multimodal pre-fusion according to claim 1, characterized in that, The output of the unassociated image data frames and the associated data frames to the domain controller includes: The laser sensor module alternately outputs the unassociated image data frames and the associated data frames to the domain controller in an n-1:1 ratio; where n is the ratio of the acquisition frequency of the image data frames to the acquisition frequency of the laser point cloud data, and n is a positive integer greater than or equal to 2.
5. The target detection system based on multimodal pre-fusion according to claim 3, characterized in that, The domain controller includes: a multi-dimensional data fusion chip and a perception-decision chip; The target detection based on the unrelated image data frames and the related data frames includes: The multi-dimensional data fusion chip performs fusion processing on the associated pixels and laser reflection points in the associated data frame to generate a multimodal data frame, and outputs the unassociated image data frame and the multimodal data frame to the perception-decision chip; wherein, each multi-dimensional pixel in the multimodal data frame includes color channel information, distance information, velocity information, angular velocity information and acceleration information; The perception-decision chip performs target detection based on the unassociated image data frames and the multimodal data frames.
6. The target detection system based on multimodal pre-fusion according to claim 5, characterized in that, The step of outputting the unassociated image data frames and the multimodal data frames to the perception-decision chip includes: The multidimensional data fusion chip alternately outputs the unrelated image data frames and the multimodal data frames to the perception-decision chip in an n-1:1 ratio; where n is the ratio of the acquisition frequency of the image data frames to the acquisition frequency of the laser point cloud data, and n is a positive integer greater than or equal to 2.
7. The target detection system based on multimodal pre-fusion according to claim 5, characterized in that, The perception-decision chip performs target detection based on the unassociated image data frames and the multimodal data frames, including: The unassociated image data frames and the multimodal data frames are input into the target detection model to obtain the 3D target detection result output by the target detection model.
8. The target detection system based on multimodal pre-fusion according to claim 5, characterized in that, The laser vision sensor module transmits the unassociated image data frames and the associated data frames to the multidimensional data fusion chip via Ethernet.
9. A target detection method based on multimodal pre-fusion, characterized in that, The target detection system is applied to a target detection system, which includes at least one visual sensor module and a domain controller. The target detection method includes: The laser sensor module synchronously acquires image data frames and laser point cloud data for the same area, associates the image data frames and laser point cloud data acquired at the same acquisition timestamp to generate associated data frames, and outputs the unassociated image data frames and associated data frames to the domain controller. The domain controller performs target detection based on the unassociated image data frames and the associated data frames.
10. A vehicle, characterized in that, include: The target detection system based on multimodal pre-fusion as described in any one of claims 1 to 9.