A multi-modal perception detection method and system thereof

CN122546202APending Publication Date: 2026-08-11SHANGHAI YUGAN MICROELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]在使用多种传感器对目标进行感知探测的过程中,会导入多种感知误差,特别是摄像头传感器获取的原始图像数据存在像差,由于原始图像数据的像差导致最终得到的多维感知信息数据存在误差,从而,导致多模态感知探测方法对目标的探测精度较低

Benefits of technology

本发明提供的一种多模态感知探测方法、系统、电子设备及计算机可读存储介质中,由于获取测量目标及其环境的多维感知信息数据,多维感知信息数据中包括若干目标感知距离以及若干多模态感知数据,每个多模态感知数据与一个目标感知距离相对应,多模态感知数据至少包括图像数据;获取若干目标感知距离对应的若干目标标定参数表;基于若干目标标定参数表,对各多模态感知数据分别进行校正,得到若干目标多模态感知数据;基于若干目标多模态感知数据,获取目标多维感知信息数据。因此,通过对若干目标感知距离对应的目标标定参数表对多模态感知数据进行校正,提高最终形成的目标多维感知信息数据的精度,从而,提高多模态感知探测方法对测量目标的探测精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122546202A_ABST
    Figure CN122546202A_ABST
Patent Text Reader

Abstract

This invention provides a multimodal sensing detection method and system. It acquires multidimensional sensing information data of the target and its environment, including several target sensing distances and several multimodal sensing data, each corresponding to a target sensing distance. The multimodal sensing data includes at least image data. The method then acquires several target calibration parameter tables corresponding to the target sensing distances. Based on these tables, each multimodal sensing data is calibrated to obtain several target multimodal sensing data. Finally, based on this multimodal sensing data, target multidimensional sensing information data is acquired. Therefore, by calibrating the multimodal sensing data using target calibration parameter tables corresponding to the target sensing distances, the accuracy of the final target multidimensional sensing information data is improved, thereby enhancing the detection accuracy of the multimodal sensing detection method for the target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sensor detection, and in particular to a multimodal sensing and detection method and system thereof. Background Technology

[0002] Multimodal perception detection methods utilize fused data from multiple sensors, such as camera sensors and radar sensors, to detect targets. These methods acquire richer environmental perception information, improving the reliability of target detection and spatial localization. Therefore, multimodal perception methods have been widely applied in fields such as autonomous driving, robotics, and industrial inspection.

[0003] In the process of using multiple sensors to perceive and detect targets, various perception errors are introduced. In particular, the raw image data acquired by the camera sensor has aberrations. Due to the aberrations of the raw image data, the final multidimensional perception information data has errors, resulting in low detection accuracy of the multimodal perception and detection method. Summary of the Invention

[0004] This invention provides a multimodal sensing detection method, system, electronic device, and computer-readable storage medium. By correcting the multimodal sensing data corresponding to the sensing distances of several targets, the accuracy of the final multidimensional sensing information data of the targets is improved, thereby improving the detection accuracy of the multimodal sensing detection method for the detected targets.

[0005] In multimodal sensing detection methods, the multimodal sensing data corresponding to the measured target and its environment includes not only image data, but also several sensing distances (object distances corresponding to optical imaging) of several image data corresponding to the measured target and its environment, and other multi-dimensional physical sensing. We use the target sensing distance of the system to calibrate and correct the aberrations of the original image data acquired by the camera module and the accuracy of the mapping between radar data and corresponding image data of radar and other subsystems for the measurement target and its environment: establish a data processing method for the sensing system corresponding to the target sensing distance through the target calibration parameter table (group) to achieve accurate detection for different target sensing distances.

[0006] According to a first aspect of the present invention, the technical solution of the present invention provides a multimodal sensing and detection method, comprising: Acquire multidimensional sensing information data of the measurement target and its environment. The multidimensional sensing information data includes several target sensing distances and several multimodal sensing data. Each multimodal sensing data corresponds to a target sensing distance. The multimodal sensing data includes at least image data. Obtain a table of target calibration parameters corresponding to the sensing distances of several targets; Based on several target calibration parameter tables, each of the multimodal sensing data is corrected to obtain several target multimodal sensing data. Based on multimodal perception data of several targets, multidimensional perception information data of the targets is obtained.

[0007] Optionally, the multimodal sensing data includes image data and radar data; Based on several target calibration parameter tables, each of the multimodal sensing data is corrected to obtain several target multimodal sensing data, including: Based on several target calibration parameter tables, the aberrations of each image data are calibrated to obtain several corrected image data, which include the original image data or the target image data.

[0008] Optionally, the multimodal sensing data includes image data and radar data; Based on several target calibration parameter tables, each of the multimodal sensing data is corrected to obtain several target multimodal sensing data, including: Based on several target calibration parameter tables, the image data and the corresponding radar data are correlated and aligned to obtain target multimodal data.

[0009] Optionally, the multimodal sensing data includes image data and radar data; Based on several target calibration parameter tables, the multimodal sensing data are corrected respectively to obtain several target multimodal sensing data, including: Based on several target calibration parameter tables, the aberrations of each image data are calibrated to obtain several corrected image data, which include the original image data or the target image data. Optionally, the multimodal sensing data includes raw image data and radar data; the pixel coordinates of each corresponding raw image data are associated and aligned with the spatial coordinates contained in the radar point cloud of each corresponding radar data, and the pixel coordinate data of the corresponding raw image data is connected based on the spatial positioning information data of the radar data.

[0010] Based on several target calibration parameter tables, the corrected image data are correlated and aligned with the corresponding radar data to obtain several target multimodal perception data.

[0011] Optionally, the aberrations of each of the image data are corrected to obtain several corrected image data, including: Using the corresponding sensing spectral frequencies, several calibration parameter tables are created based on several sensing distances, and several target calibration parameter tables corresponding to several target sensing distances at the corresponding sensing spectral frequencies are obtained. Based on several target calibration parameter tables, each image data is corrected according to its corresponding spectral frequency to obtain several corrected image data.

[0012] Optionally, the aberrations of each of the image data are corrected to obtain corrected image data, including: Based on several target calibration parameter tables, at least one aberration among distortion, field curvature, spherical aberration, chromatic aberration, and coma of each image data is corrected to obtain several target image data.

[0013] Optionally, obtain a table of target calibration parameters corresponding to several target perception distances, including: External radar equipment is used for calibration to obtain a table of target calibration parameters for the sensing range of several targets.

[0014] Optionally, based on several target calibration parameter tables, each of the multimodal sensing data is corrected to obtain several target multimodal sensing data, further comprising: Based on several target calibration parameter tables, obtain the target calibration parameter table for each target perception distance in all target perception distances, specifically including: Determine whether the current target perception distance is the target perception distance in the existing target calibration parameter table; If so, then call the target calibration parameter table corresponding to the current target perception distance; If not, then interpolate the target calibration parameter tables of the two target perception distances adjacent to the current target perception distance to obtain the target calibration parameter table corresponding to the current target perception distance.

[0015] According to a second aspect of the present invention, a multimodal sensing and detection system is provided for implementing the multimodal sensing and detection method as described above, comprising: The data acquisition module is used to acquire multidimensional sensing information data of the measurement target and its environment. The multidimensional sensing information data includes several target sensing distances and several multimodal sensing data. Each multimodal sensing data corresponds to a target sensing distance. The multimodal sensing data includes at least image data. The target calibration parameter table acquisition module is used to acquire several target calibration parameter tables corresponding to the sensing distances of several targets. The calibration module is used to calibrate each of the multimodal sensing data based on several target calibration parameter tables to obtain several target multimodal sensing data. The target data acquisition module is used to acquire multi-dimensional perception information data of targets based on multi-modal perception data of several targets.

[0016] According to a third aspect of the present invention, an electronic device is provided, including a processor and a memory, the memory being used to store code; The processor is used to execute the code in the memory to implement the multimodal sensing and detection method as described above.

[0017] According to a fourth aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the multimodal sensing and detection method as described above.

[0018] Compared with the prior art, the technical solution of the embodiments of the present invention has the following beneficial effects: This invention provides a multimodal sensing detection method, system, electronic device, and computer-readable storage medium. By acquiring multidimensional sensing information data of the measurement target and its environment, including several target sensing distances and several multimodal sensing data, each multimodal sensing data corresponding to a target sensing distance, and each multimodal sensing data including at least image data; acquiring several target calibration parameter tables corresponding to the target sensing distances; correcting each multimodal sensing data based on the target calibration parameter tables to obtain several target multimodal sensing data; and acquiring target multidimensional sensing information data based on the target multimodal sensing data. Therefore, by correcting the multimodal sensing data using target calibration parameter tables corresponding to the target sensing distances, the accuracy of the final target multidimensional sensing information data is improved, thereby enhancing the detection accuracy of the multimodal sensing detection method for the measurement target. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart illustrating a multimodal sensing and detection method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the correction principle of raw image data according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a multimodal sensing and detection system provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0021] As described in the background section, the present invention aims to solve the technical problem of improving the detection accuracy of multimodal sensing and detection methods for detection targets.

[0022] Currently, the output of mainstream sensing systems is digital data. The assembly of optical lenses with sensing sensors (such as CCD and CMOS sensing chips) introduces component processing and installation errors. For example, the optical center and optical axis of the camera sensor lens may deviate from the center and central axis of the sensing chip. Moreover, the optical systems we actually use are not pinhole imaging systems, but have a certain optical lens aperture (entrance pupil diameter). The aberrations of the original image data acquired by the corresponding camera module are also related to the lens aperture parameters and the corresponding actual imaging object distance (target sensing distance). All these factors will cause the aberrations of the image data generated and output by the system to be related to the object distance between the target and the environment sensed by the system. We also need to perform aberration correction and alignment of the corresponding output image data in the system. Furthermore, the definition of "system image distortion" in this paper differs from the definition of "distortion" in geometric optics. The definition of "system image distortion" in this paper covers the system level and includes the distortion aberration of the system output image caused by factors such as the assembly, component processing, and installation errors of optical lenses and sensing sensors (such as CCD and CMOS sensing chips). In addition, the misalignment error of the installation position of the camera and radar subsystem within the multimodal sensing system will also introduce spatial mapping and positioning deviations in the image data within the system's multimodal sensing data. We refer to all of these as "system image distortion" in this paper.

[0023] In existing technologies, aberration correction of image data obtained from image sensors mainly relies on two approaches. One approach involves physical compensation for aberrations through the optical lens design of the camera sensor. This includes techniques such as symmetrical structural designs, aspherical lenses, special dispersive materials, and floating lenses. By adjusting the shape, structure, and materials of the lens, light is directed onto the image sensor in the most ideal way possible. However, purely optical design methods result in bulky camera sensor lenses, high manufacturing difficulty, and high production costs. Even with optimal design, manufacturing errors and assembly tolerances of the various components of the camera sensor are unavoidable during production and assembly. Each camera sensor has physical tolerances, making complete consistency impossible. The other approach involves calibration and correction of image distortion through subsequent algorithms. Typical methods include introducing fixed calibration objects into the image data and fitting the distortion process based on a polynomial distortion model (such as radial distortion coefficients) before inversely solving the distortion. Alternatively, in the absence of calibration objects, distortion parameters can be calculated by identifying geometric features such as lines and circles in the image data. However, the distortion parameters derived by these methods have limited accuracy and are prone to misalignment in complex scenes, making it difficult for the system's reliability and accuracy to meet practical application requirements. More importantly, optical lens aberrations (including distortion, spherical aberration, chromatic aberration, field curvature, coma, etc.) arise because the camera sensor's lens has different refractive indices and magnifications for light with different field of view angles or different spectral frequencies, causing distortion in object imaging. This aberration exists before the light enters the sensor and is converted into an electrical signal, representing a physical error at the level of the camera sensor's raw image data. Existing calibration methods mainly process the two-dimensional image plane, failing to fully consider the strong correlation between aberrations and target perception distance, and also failing to perform high-precision spatial alignment of three-dimensional physical space position perception information (such as radar data) with the corresponding image data pixel information. Therefore, for applications such as autonomous driving, robotics, and industrial inspection that require the simultaneous perception of multiple measurement targets in three-dimensional space and the maintenance of high-precision alignment at different perception distances, existing solutions cannot meet the system's requirement to accurately connect the image data of the measurement targets and the environment with their physical three-dimensional positions. They also lack a means of unifying and adaptively correcting three-dimensional spatial perception information with two-dimensional image information based on a multimodal perception system.

[0024] In view of this, the present invention proposes a multimodal sensing detection method. This method acquires multidimensional sensing information data of the target and its environment, including several target sensing distances and several multimodal sensing data, each corresponding to a target sensing distance. The multimodal sensing data includes at least image data. The method then acquires several target calibration parameter tables corresponding to the target sensing distances. Based on these calibration parameter tables, each multimodal sensing data is corrected to obtain several target multimodal sensing data. Finally, based on these target multimodal sensing data, target multidimensional sensing information data is acquired. Therefore, by correcting the multimodal sensing data using the target calibration parameter tables corresponding to the target sensing distances, the accuracy of the final target multidimensional sensing information data is improved, thereby enhancing the detection accuracy of the multimodal sensing detection method for the target.

[0025] In this invention, the first multimodal sensing detection method can correct the image data based on the target sensing distance in the multidimensional sensing information data acquired by the sensing system, and then map the radar point cloud data information of the radar data to the pixels of the corresponding corrected image data to complete the multimodal sensing based on the image data. The second multimodal sensing detection method can use a target calibration parameter table based on the target sensing distance (i.e., target object distance) to perform extrinsic parameter calibration. This involves using a target calibration parameter table based on several associated object distances. Each radar data point cloud is mapped to the corresponding image data by calling the corresponding target sensing distance's extrinsic parameter table based on its target sensing distance. The point cloud data of each radar data point is then mapped onto the original image data by performing translation, rotation, and scaling on its spatial positioning data based on its target sensing distance (corresponding to the camera's object distance) using the corresponding target calibration parameter table. This completes the docking of multimodal fusion sensing radar data and image data. These two methods can also be combined; however, this increases the system's computational load but results in more accurate output sensing data.

[0026] For the first multimodal sensing detection method described above, the multimodal sensing detection method may include: First, based on the target calibration parameter tables, calibrating the original image data in the multimodal sensing data to obtain target image data, correcting each pixel of the original image data to its corresponding spatial position to obtain target image data; then, using the target image data with these pixels having spatial positioning data to fit a high-order polynomial model to correct the deformation of the original image; or using interpolation or other methods to generate the corresponding geometric transformation of the entire image, and then aligning the calibrated target image data with the point cloud data of the corresponding original radar data to obtain several multimodal sensing data; finally, obtaining target multidimensional sensing information data based on several multimodal sensing data.

[0027] In the above embodiments, after the mapping and docking of the corresponding target perception distance and the original image data are completed in the current frame, the entire image is projected and transformed to eliminate the aberrations of the target image. In particular, for the "system image distortion" mentioned in this article, the elimination of the "system image distortion" aberration can achieve a very high accuracy.

[0028] For the second multimodal sensing detection method mentioned above, we also have another option: we can retain the three-dimensional spatial position data of the corresponding pixels in the original image data in the output data. The system can then use the corresponding image correction algorithm for further processing as needed. Since we have already completed the accurate spatial position sensing and localization of the corresponding pixels in the image (these corresponding image pixels have the accurate spatial position data information of the radar-sensed point cloud), even if we do not eliminate aberrations and output updated target image data in real time based on image deformation algorithms (the system retains the original image data in the output multimodal sensing data), the task of correcting the multimodal sensing data can actually be considered as completed because we have already completed the accurate spatial position sensing of the corresponding pixels in the image and retained the synchronous output of the corresponding spatial position data (the corresponding relevant correction data have been generated and output).

[0029] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0030] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0031] Please refer to Figure 1 The embodiments of the present invention provide a multimodal sensing and detection method, which may include: S100: Acquire multidimensional sensing information data of the measurement target and its environment. The multidimensional sensing information data includes several target sensing distances and several multimodal sensing data. Each multimodal sensing data corresponds to a target sensing distance. The multimodal sensing data includes at least image data.

[0032] In one embodiment, the multimodal sensing data may further include radar data.

[0033] In this embodiment, the radar data may include a combination of power measurement information data corresponding to the four dimensions of the radar 4D tensor (distance, Doppler, horizontal azimuth, and elevation angle) sensed by the radar sensor at each sampling point in the three-dimensional space where the radar sensor's detection domain is located. The power measurement information data includes: data on signal reflection energy, or data on radar cross section or signal-to-noise ratio information corresponding to the signal reflection energy.

[0034] The multimodal perception system contains spatial location point cloud information output by radar sensors for target and environment perception and detection (the role of radar sensors in the multimodal perception system is to output relatively accurate coordinate data of the target and its environment in physical space). The point cloud information output by radar sensors for target and environment perception and detection includes the physical perception of the spatial location dimension of the target perception distance (corresponding to the object distance in optical imaging). We use the system's target perception distance to correct the aberration errors of the image data: a parameter table for extrinsic parameter calibration is made based on the target perception distance (the system performs one-to-one translation, rotation and scaling of image pixels and radar-sensed point cloud data), and the corresponding pixels of the multimodal fused perception image data are precisely aligned with the point cloud of the radar point cloud data based on the corresponding target perception distance (object distance).

[0035] In one example, radar sensors may include lidar, millimeter-wave radar, or terahertz radar.

[0036] In this embodiment, image data represents the data acquired by the image sensor.

[0037] Based on the point cloud information of radar sensors that detect and perceive the target and its environment in a multimodal perception system, which includes the spatial position dimension of the target perception distance (corresponding to the object distance in optical imaging), we use the target perception distance of the system to complete the precise image deformation correction of the target digital image imaging based on the object distance. Then, we project the radar point cloud information onto the corresponding output image pixels to complete the multimodal precise perception based on the image pixels. Aberrations in optical lenses (taking distortion as an example; similar aberrations exist for other aberrations such as spherical aberration and field curvature) arise because lenses have different magnifications for light at different field angles, leading to distortion in the image of the object. The assembly of the optical lens with the sensing sensor (such as a CCD or CMOS sensor chip) also introduces errors. For example, the optical center and optical axis of the lens may deviate from the center and central axis of the sensing chip. Furthermore, the optical systems we actually use are not pinhole imaging systems; they must have a certain optical lens aperture (entrance pupil diameter). The corresponding image projection aberrations are simultaneously correlated with the lens aperture parameters and the corresponding actual imaging object distance. All these factors contribute to the aberrations in the digital image generated and output by the system, which are also related to the object distance of the measured target and its environment. Based on this implementation method, we correct these aberrations in the image imaging subsystem.

[0038] In this embodiment, the multimodal perception system includes a radar sensor and a camera sensor. The radar sensor outputs point cloud information showing the spatial location of the target and its environment in physical space. The point cloud information output by the radar sensor includes physical perception data corresponding to the target perception distance (i.e., the object distance in optical imaging) in the spatial location dimension. The camera sensor is used to acquire image data. This embodiment utilizes the target perception distance acquired by the system to correct aberrations in the image data: a target calibration parameter table is constructed based on the target perception distance, and the pixels of the image data are mapped to the radar point cloud of the radar data through translation, rotation, and scaling to achieve accurate alignment between image data and radar data in multimodal fusion perception.

[0039] S200: Obtain a table of target calibration parameters corresponding to the sensing distances of several targets.

[0040] The system collects measurement data from radar and camera sensors for the target and its environment. It then analyzes the measurement values ​​from the radar and camera sensors for the calibration tool (the measurement values ​​can be derived based on multiple calibration points (calibration groups). The tolerances are calculated based on these calibration points (calibration groups), and the calibration data with the smallest tolerance is used as the calibration measurement value): namely, the point cloud 3D coordinate values ​​and pixel planar coordinate values. The system calculates the corresponding scaling, rotation, and translation matrix parameters (usually including a 3×3 rotation matrix and a 3×1 translation matrix). Since the aberration of the original image data acquired by the camera module is related to the target perception distance (object distance), the aberration of the image data of the measured target and its environment at different target perception distances (projected onto the same imaging plane, i.e., the pixel plane of the same imaging sensing chip (such as CCD or CMOS chip)) is different. We import multiple sets (N sets) of calibration objects with different object distances Li (i=1,2,3...N, N>1, and N is an integer) for extrinsic parameter calibration, and generate corresponding calibration parameter combination tables for different object distances Li (i.e., scaling, rotation, and translation matrix parameter tables distributed within the image frame for different object distances) for extrinsic parameter calibration. Then, in the use of the multimodal sensing system, based on the object distance Li of the target and environment perception, the calibration parameter combination table of object distance Li (i.e., scaling, rotation, and translation matrix parameter tables for the same object distance) is called to map and connect the corresponding radar data with the image data of the same target, so as to complete the precise alignment of the point cloud data of the multimodal fusion sensing image data and the radar data.

[0041] In this embodiment, the target calibration parameter table is a dataset containing the target perception distance and the corresponding correction mapping relationship.

[0042] In an optional implementation, S200, obtaining a table of target calibration parameters corresponding to several target perception distances may include: External radar equipment is used for calibration to obtain a table of target calibration parameters for the sensing range of several targets.

[0043] External radar equipment is used for calibration to obtain several target calibration parameter tables for the sensing distance of several targets. In this embodiment, firstly, external radar equipment typically has higher measurement accuracy and more stable signal output capabilities, and can acquire more reliable raw data than in-house developed vehicle or airborne sensors, thus significantly improving the accuracy and reliability of the calibration parameter tables; secondly, using external radar equipment for calibration decouples the calibration process from the operation of the main system, meaning the main system does not need to bear the additional burden of data acquisition and processing during the calibration phase, thereby shortening the calibration time and reducing the debugging complexity of the main system; thirdly, as an independent external reference source, the data and parameter table format of the external radar equipment can be cross-validated with the calibration parameter tables used internally by the main system, facilitating the discovery and correction of potential system deviations in the main system's own sensors; fourthly, when the main system's operating environment changes (such as temperature, humidity, installation location offset, etc.) and recalibration is required, the external radar equipment can be quickly deployed. Furthermore, it can be reused without requiring hardware modifications to the main system, improving the flexibility and maintainability of the calibration process. Fifth, the external dedicated calibration radar can also be used to calibrate the radar subsystems within the multimodal perception system. For example, the internal radar subsystem also uses optical lenses—a popular structure in current lidar—which can introduce optical aberrations that affect the radar's detection of the corresponding target. Importing an external calibration radar can optimize this problem by calibrating the internal radar subsystem, making the radar data within the system more accurate. The system also imports multiple sets of calibration parameter tables based on the target distance (object distance) to achieve more accurate perception. Finally, the calibration parameter tables obtained through external radar equipment can be stored and traced as a baseline version, facilitating comparative analysis of calibration results at different times and under different environments, providing reliable data support for subsequent algorithm optimization and troubleshooting.

[0044] In this embodiment, radar data and image data are collected to measure the target and its environment, and the measurement values ​​of the radar data and image data are analyzed respectively. The measurement values ​​can be derived based on a calibration group consisting of multiple calibration points. Specifically, the coordinates and pixel coordinates of the radar point cloud corresponding to each calibration point are calculated, scaling, rotation, and translation matrix parameters are calculated, tolerances are calculated based on these calibration points, and the calibration data with the smallest tolerance is derived as the target calibration parameter table used by the calibration group.

[0045] In this embodiment, it is difficult to obtain a target calibration parameter table corresponding to each target sensing distance among all target sensing distances. However, in order to correct each pixel in the image data, it is necessary to obtain a target calibration parameter table corresponding to each target sensing distance. This results in higher accuracy of the corrected image data.

[0046] In this embodiment, for any target sensing distance, multiple sets of candidate parameters are calculated using multiple calibrations, and the set with the smallest tolerance (most accurate alignment) is selected as the target calibration parameter table for that target sensing distance.

[0047] In this embodiment, based on several target calibration parameter tables, each of the multimodal sensing data is corrected to obtain several target multimodal sensing data, which may further include: Based on several target calibration parameter tables, obtain the target calibration parameter table for each target perception distance in all target perception distances, specifically including: Determine whether the current target perception distance is the target perception distance in the existing target calibration parameter table; If so, then call the target calibration parameter table corresponding to the current target perception distance; If not, then interpolate the target calibration parameter tables of the two target perception distances adjacent to the current target perception distance to obtain the target calibration parameter table corresponding to the current target perception distance.

[0048] Interpolation algorithms include linear interpolation and polynomial interpolation.

[0049] In this embodiment, by means of the above method, while ensuring calibration accuracy, the number of target calibration parameter tables that need to be originally acquired is reduced, and continuous coverage of the sensing distance of any target is achieved, thereby improving the accuracy of the final sensing data.

[0050] S300: Based on several target calibration parameter tables, each of the multimodal sensing data is corrected to obtain several target multimodal sensing data.

[0051] In one specific embodiment, S300, based on several target calibration parameter tables, each of the multimodal sensing data is corrected to obtain several target multimodal sensing data, which may include: Based on several target calibration parameter tables, the phase difference of each image data is calibrated to obtain several corrected image data, which include the original image data or the target image data.

[0052] In this embodiment, the aberrations of each image data are corrected separately. The correction methods for the corrected image data can be divided into two types. The first correction method is to correct each pixel of the image data to the correct spatial position to obtain the target image data (e.g., ...). Figure 2(As shown). The second correction method is to align the image data with the corresponding radar data based on several target calibration parameter tables to obtain the original correct spatial position of each pixel as the output of the corrected image data, while retaining the original image data.

[0053] In this embodiment, since the image aberrations of the corresponding image data originally acquired by the optical lens and the digital imaging system are related to the corresponding imaging object distance, different target perception distances will cause different aberrations in the image data generated and output by the system. In this embodiment, the above-mentioned method based on this embodiment can correct these aberrations in the image imaging subsystem.

[0054] In another specific embodiment, S300, based on several target calibration parameter tables, each of the multimodal sensing data is corrected to obtain several target multimodal sensing data, which may further include: Based on several target calibration parameter tables, each image data and its corresponding radar data are correlated and aligned to obtain target multimodal data.

[0055] In this embodiment, associated alignment includes translation, rotation, and scaling.

[0056] In this embodiment, the corresponding target distance extrinsic parameter table is called to map and interface the corresponding radar point cloud data with the image data of the same target. The radar-sensed point cloud data is translated, rotated, and scaled one-to-one to align with the image data, thus completing the interface between the multimodal fusion sensing radar point cloud data and the image pixel data. The pixel coordinates of the image data have been corrected to the correct spatial position, obtaining the target multidimensional sensing information data. The system-aligned image data and radar point cloud data are output as sensing data. At the same time, the spatial position data information of all radar point clouds (including the spatial position coordinate sensing data of the corresponding radar point cloud based on the sensing target) is retained in the output aligned target image pixel data and radar point cloud sensing data. Under this processing method, the system has actually completed the accurate alignment of the sensing image pixel data and radar point cloud data. Subsequently, the system can adjust accordingly based on whether the image needs to be deformed and corrected before outputting (aligning the target image pixel data and the corresponding radar point cloud, while retaining and outputting the spatial position data information of the corresponding radar point cloud, means that the system has completed the spatial position correction of the image).

[0057] In this embodiment, on the one hand, the image data is aligned by target-by-target sensing distance correction of the target alignment parameters, eliminating the nonlinear variation of alignment error at different target distances caused by factors such as differences in sensor installation position, lens distortion, and changes in radar beam angle. Compared with the traditional alignment method using a single fixed transformation matrix, this method has higher accuracy and range adaptability. On the other hand, the target multimodal sensing data obtained after precise alignment retains the high-resolution texture features of the image data and the depth, Doppler, and angle information of the radar. This provides a data foundation with good spatial consistency and strong information complementarity for subsequent target detection, tracking, recognition, and other fusion sensing tasks, thereby effectively improving the target positioning accuracy and reliability of the entire sensing system in complex environments.

[0058] In another specific embodiment, S300, based on several target calibration parameter tables, the multimodal sensing data are corrected respectively to obtain several target multimodal sensing data, which may further include: Based on several target calibration parameter tables, each of the multimodal sensing data is corrected to obtain several target multimodal sensing data, including: Based on several target calibration parameter tables, the aberrations of each image data are calibrated to obtain several corrected image data, which include the original image data or the target image data. Based on several target calibration parameter tables, the corrected image data are correlated and aligned with the corresponding radar data to obtain several target multimodal perception data.

[0059] In this embodiment, the two methods described above are used in combination. Of course, this will increase the amount of data processing in the system, but the output multimodal sensing data of the target can be more accurate.

[0060] In this embodiment, due to the manufacturing and assembly tolerances of the optical lens and digital imaging system, i.e., the image sensor, the original image data contains certain errors in aberrations such as distortion, field curvature, spherical aberration, chromatic aberration, and coma. Therefore, in this embodiment, the aberrations of each of the original image data are corrected to obtain corrected image data, which may include: Based on several target calibration parameter tables, at least one aberration among distortion, field curvature, spherical aberration, chromatic aberration, and coma of each of the original image data is corrected to obtain several corrected image data.

[0061] In this embodiment, it applies to multimodal sensing systems (including multimodal sensing modules and multimodal sensing chips such as RGB-IR / RCB-IR / RGB-D). Multimodal sensing systems can be divided into: 1) Non-coaxial multimodal sensing systems, where the main axes of the target sensing areas of each sensing module (camera, radar, infrared sensing module, etc.) do not coincide—that is, they are non-coaxial—but they correspond to a common sensing area, and multimodal sensing is established within their corresponding common sensing area; 2) Coaxial (each sensing module is based on the same system main axis) multimodal sensing systems. Let's illustrate our corresponding innovative applications with examples: 1) Multimodal sensing systems that are not coaxial but share a common sensing region (field of view): Non-coaxial multimodal sensing systems require spatial alignment of the various sensing components (such as visible light camera sensors and corresponding radar sensor systems) with respect to the target and its environment. However, image aberrations (including system image distortion) cause distortion in the imaging of objects at different spatial locations, and the degree of distortion varies. Furthermore, the positional tolerances of each system installation can lead to deviations in the spatial positioning data for the corresponding target, resulting in inaccurate perception.

[0062] Non-coaxial multimodal sensing systems require establishing a unified detection domain for the system and transforming the detection space mapping relationship of multiple sensors to the sensor perception space with the most information content (or the most important one) determined during system design (which we call the reference sensing unit for multimodal sensing fusion in this paper). (For example, the 3D detection perception space corresponding to a visible light camera sensor is used as the reference (i.e., the object space corresponding to the camera). In some applications, the 3D detection space of a radar sensor (including lidar, millimeter-wave, or terahertz radar) is also used as the reference. The specific operation is as follows: the detection domains of each sensor are associated in a standard 3D Euclidean solid geometric space using geometric space transformation, by transforming their respective coordinates.) The system coordinate axes are translated, rotated, and scaled to unify them into the three-dimensional detection space coordinates of the camera sensor, aligning them with the detection domain center axis (e.g., the optical axis of the camera) of the reference sensing unit for multimodal perception fusion, thus establishing a unified detection space and a common detection perspective for the system. Then, based on the established mapping relationship, the system determines the sensory data docking (data fusion) of the target object corresponding to each detection data of the target object in the detection domain of the reference sensing unit for multimodal perception fusion. The mapping relationship is used to characterize the mapping relationship between the detection data of different positions in different detection dimensions of the other sensors and the basic sensing unit (e.g., image pixel) of the reference sensing unit for multimodal perception fusion.

[0063] If optical cameras are used in the sensors of a multimodal fusion system, the aberrations (including system image distortion) of the perceived image data obtained by the camera sensors need to be calculated based on the distance (target perception distance) of the corresponding measurement target (there may be multiple measurement targets) in physical space to obtain the aberrations (including system image distortion) of the original image corresponding to the image data and then eliminate the error. Otherwise, the corresponding target perception data will differ from their positioning in physical space. This difference will lead to problems such as misjudgment of the target by the system, which needs to be corrected.

[0064] In this embodiment, a calibration parameter method for multimodal perception spatial alignment (including image system image distortion calibration) is introduced based on the target perception distance of the measured target and its environment. According to the present invention, the method of the above embodiment can effectively solve the problem and achieve "accurate detection of perception spatial distance and angle for targets at different distances".

[0065] 2) Coaxial (each sensing module is based on the same system main axis) multimodal sensing system: The 3D spatial coordinate information (XYZ spatial coordinates, or the corresponding target perception distance, horizontal axis angle, and vertical axis angle) is fused with the image pixel data (RGB or YUV data group) corresponding to the perceived image data. Each point output carries information such as visible light brightness, color, and spatial position, achieving pixel-level spatial position alignment. Coaxial systems establish a unified detection domain and map the detection space of multiple sensors onto the same sensing axis, eliminating the need for spatial coordinate transformations. However, if the system uses a single optical camera, image distortion arises from the different refractive indices of the lens for different spectral frequencies (e.g., visible light, far-infrared, ultraviolet light), resulting in varying magnifications. This causes optical distortion between the visible light sensing unit and the target distance sensing unit (currently, the industry typically uses 905nm or 1550nm infrared bands). Even with chip-level fusion, at high resolutions (where each photosensitive unit has a small area), the lens optical distortion due to different refractive indices for different spectral frequencies exceeds the distance between the visible light sensing unit and the target distance sensing unit. This results in a misalignment between the target sensing data in the visible light RGB spectrum and the target distance sensing unit, leading to spatial positioning data deviation. This deviation is also related to the target's spatial distance (target sensing distance). In this embodiment, the method described above is used to correct this deviation.

[0066] Regardless of whether it is coaxial or not, in this embodiment, regarding the elimination of lens aberrations introduced by the optical sensing system, for most applications (where the accuracy of physical space perception does not necessarily require pixel-level precision across different spectra within the corresponding visible light spectrum), visible light aberrations can be calibrated using white light. However, for applications requiring pixel-level precision across the entire field of view, we need to calibrate and align the corresponding RGB light aberrations for different target perception distances (the measurement target and its surrounding environment are related) for different corresponding RGB spectra. This is necessary to achieve accurate calibration across the entire spectrum. Of course, for systems containing sensing modules beyond the visible light spectrum (e.g., infrared sensing units), we need to perform image aberration calibration and alignment for targets at different distances based on the spectral frequencies sensed by the corresponding sensing modules beyond the visible light spectrum, and then fuse these calibrations into the corresponding calibration unit (the calibration unit can be based on a hardware accelerator or implemented by CPU, GPU, or other computing units calling corresponding software).

[0067] In this embodiment, the geometric distortion and edge blurring of the image data are reduced by the above-described method, thereby improving the overall clarity and uniformity of the image data. The resulting image data shows significant improvement in geometric accuracy and image quality, providing accurate and reliable image data for subsequent spatial alignment and multimodal fusion with the original radar data.

[0068] In this embodiment, because the optical lens has different refractive indices for light of different spectral frequencies (such as visible light, far-infrared light, or ultraviolet light), it results in different magnifications of the detected target. Therefore, the image distortion of different objects varies depending on the spectral frequency of the light. Thus, it is necessary to correct for different spectral frequencies (such as visible light, far-infrared light, or ultraviolet light) corresponding to the sensing spectrum used by the sensing module at the target sensing distance. Therefore, in another optional embodiment, correcting the aberrations of each of the original image data to obtain target image data may include: Based on several target calibration parameter tables, several calibration parameter tables based on several sensing distances are made according to the sensing spectral frequencies used by the corresponding sensing modules, and several target calibration parameter tables corresponding to several target sensing distances at the corresponding sensing spectral frequencies are obtained. Based on several target calibration parameter tables, each of the original image data is corrected according to its corresponding spectral frequency.

[0069] In this embodiment, the above steps utilize pre-stored correction parameters in the target calibration parameter table, which are associated with the target sensing distance and various spectral frequencies, to perform spectral correction on the original image data, eliminating image scale inconsistencies and geometric distortions caused by different spectral characteristics. This corresponds to image correction for monochromatic light (infrared imaging, ultraviolet imaging) imaging systems, as well as image correction (also monochromatic images) for lidar based on optical lens imaging. As long as the system has the sensing of optical lens + target distance information, this invention can perform aberration correction on the original image data corresponding to different target sensing distances based on the corresponding spectral frequencies to improve sensing accuracy. The target image data obtained thereby can improve accuracy at different spectral frequencies, providing accurate and consistent image input for subsequent spatial alignment and multimodal fusion with the original radar data, effectively improving the target matching accuracy and overall reliability of the sensing system in multispectral scenarios.

[0070] Because focusing adjustments of optical lenses introduce changes in aberrations into the original image data (including distortion, field curvature, spherical aberration, chromatic aberration, and coma—at least one aberration), if the lens is further modified into an optical zoom lens, the adjustments will result in even greater aberrations. The aberrations in the original image data are related to the focal length and aperture (entrance pupil diameter) of the optical lens. We utilize multimodal sensing data (containing combined spatial location sensing information of original image data and original radar data) to correct the distortions of each original image data frame during zooming for different combinations of optical focal lengths and apertures. This includes simultaneously correcting for the aberrations in the original image data caused by component and system projection docking errors in the optical lens and sensing sensors (such as CCD and CMOS sensing chips) during focusing and zooming.

[0071] S400: Based on multimodal sensing data of several targets, acquire multidimensional sensing information data of the targets.

[0072] In this embodiment, the multimodal sensing and detection method is applicable to application scenarios where the main axes of the target sensing areas of image sensors, radar sensors, and other sensors do not coincide, i.e., they are not coaxial, and it is also applicable to application scenarios where the main axes of the target sensing areas of various sensors coincide.

[0073] In this embodiment, after the system completes the mapping and docking of the target and the original radar data acquired by environmental perception with the original image data according to the corresponding object distance in the current frame, the system generates a pixel combination of the original image data with accurate three-dimensional spatial perception data of radar data. We then perform projection transformation on the entire image generated from the original image data (e.g., Figure 2As shown, this process eliminates aberrations in image data (especially aberrations caused by image system distortion). Because the pixels in these image data possess precise three-dimensional spatial location information, aberration elimination can be highly accurate. During real-time system perception, we can perform image transformation processing based on the pixel coordinates of the aligned image data output by the system in real time and the perceived target area of ​​the spatial point cloud of the radar data. This can be done using methods such as polynomial fitting, where pixels with spatial positioning data are used to fit a high-order polynomial model (e.g., a quadratic binomial), correcting the original image distortion into a regular plane; or using interpolation to generate the corresponding geometric transformation for the entire image, and then outputting the aligned target multimodal perception image. This processing makes the sensing perspective of the output image more consistent with the target position correspondence in object space, which is equivalent to system image distortion calibration. Since the system completes the mapping and docking of the acquired radar data with the corresponding object distance and image data, the system generates several original image data pixels with precise three-dimensional spatial perception data from the radar data. The system can support projection algorithms based on stereo space to complete the corresponding image distortion calibration processing.

[0074] In this embodiment, the image system's image distortion calibration method can also be used to process the target calibration parameters, i.e., calibration parameters, for a given object distance. The key point is that our method imports these parameters (parameter table) and associates them with the corresponding target sensing distances perceived by the system (i.e., multiple sets of parameters correspond to image aberration correction tables for different sensing distances). Then, using the system's target sensing distance, based on the target calibration parameter tables for different target sensing distances, we can accurately correct the image deformation corresponding to the image data at the target sensing distance. Specific implementation: A fixed calibration object is added to the imaging image, and then the image is captured. Then, the corresponding calibration object image is found in the two-dimensional plane image. Based on the geometric position information of the calibration object, a mathematical model is constructed to fit the output image deformation (mainly system image distortion) process. Then, the image deformation parameters are derived by reverse solving. For example, a polynomial distortion model: a polynomial function (usually using radial distortion coefficients k1, k2, k3, etc.) is used to model and measure the "pincushion" or "barrel" distortion on the two-dimensional plane of the image. We perform image pixel traversal calculations and docking on the full-frame coordinate system of the image. This process is equivalent to establishing a correspondence between the aberration-calibrated image pixel coordinates and the pixel coordinates of the original input distorted image. The aberration-calibrated image pixel coordinates of the image data find a mapping in the pixel image of the original input distorted image. We call this mapping table the image distortion calibration parameter table. In practical use, for each frame of the input original image, the multimodal perception system uses the corresponding mapping table to traverse the input distorted image for each output image pixel according to the corresponding projection relationship in the mapping table, assigns values, and then outputs the corrected original image data.

[0075] In this embodiment, multiple sets (N sets) of calibration objects with different target perception distances Li (i=1,2,3...N) are imported for calibration. The more sets of N, the better, so that the actual target and environment perception of the system will be more accurate. For the target and its environment with different target perception distances, the target calibration parameter table with the image deformation that is closest to its target perception value can be found. The calibration accuracy will be higher in this way.

[0076] As can be seen, in this embodiment, multi-dimensional sensing information data of the target and its environment is acquired. This multi-dimensional sensing information data includes several target sensing distances and several multi-modal sensing data, each multi-modal sensing data corresponding to a target sensing distance. The multi-modal sensing data includes at least the original image data. Several target calibration parameter tables corresponding to the target sensing distances are acquired. Based on these target calibration parameter tables, each multi-modal sensing data is corrected to obtain several target multi-modal sensing data. Based on these target multi-modal sensing data, target multi-dimensional sensing information data is acquired. Therefore, by correcting the multi-modal sensing data using target calibration parameter tables corresponding to the target sensing distances, the accuracy of the final target multi-dimensional sensing information data is improved, thereby enhancing the detection accuracy of the multi-modal sensing detection method for the target.

[0077] Accordingly, please refer to Figure 3 The embodiments of the present invention also provide a multimodal sensing and detection system for implementing the multimodal sensing and detection method as described above, which may include: a data acquisition module 100, a target calibration parameter table acquisition module 200, a correction module 300, and a target data acquisition module 400.

[0078] The data acquisition module 100 is used to acquire multidimensional sensing information data of the measurement target and its environment. The multidimensional sensing information data includes several target sensing distances and several multimodal sensing data. Each multimodal sensing data corresponds to a target sensing distance. The multimodal sensing data includes at least image data.

[0079] The target calibration parameter table acquisition module 200 is used to acquire several target calibration parameter tables corresponding to the perception distances of several targets.

[0080] The calibration module 300 is used to calibrate each of the multimodal sensing data based on several target calibration parameter tables to obtain several target multimodal sensing data.

[0081] The target data acquisition module 400 is used to acquire multi-dimensional perception information data of targets based on multi-modal perception data of several targets.

[0082] In addition, please refer to Figure 4 An embodiment of the present invention also provides an electronic device 500, including a processor 510 and a memory 520, wherein the memory 520 is used to store code; The processor 510 is used to execute the code in the memory 520 to implement the multimodal sensing and detection method as described above.

[0083] In this embodiment, the processor 510 can communicate with the memory 520 via the bus 530.

[0084] Embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multimodal sensing and detection method as described above.

[0085] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned computer-readable storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0086] In summary, in the above embodiments, by acquiring multi-dimensional sensing information data of the target and its environment, including several target sensing distances and several multi-modal sensing data, each multi-modal sensing data corresponding to a target sensing distance, and the multi-modal sensing data including at least the original image data; acquiring several target calibration parameter tables corresponding to several target sensing distances; correcting each multi-modal sensing data based on the several target calibration parameter tables to obtain several target multi-modal sensing data; and acquiring target multi-dimensional sensing information data based on the several target multi-modal sensing data. Therefore, by correcting the multi-modal sensing data using target calibration parameter tables corresponding to several target sensing distances, the accuracy of the final target multi-dimensional sensing information data is improved, thereby improving the detection accuracy of the multi-modal sensing detection method for the target.

[0087] While the present invention has been disclosed above, it is not limited thereto. Any person skilled in the art can make various modifications and alterations without departing from the spirit and scope of the invention; therefore, the scope of protection of the present invention should be determined by the scope defined in the claims.

Claims

1. A multimodal sensing and detection method, characterized in that, include: Acquire multidimensional sensing information data of the measurement target and its environment. The multidimensional sensing information data includes several target sensing distances and several multimodal sensing data. Each multimodal sensing data corresponds to a target sensing distance. The multimodal sensing data includes at least image data. Obtain a table of target calibration parameters corresponding to the sensing distances of several targets; Based on several target calibration parameter tables, each of the multimodal sensing data is corrected to obtain several target multimodal sensing data. Based on multimodal perception data of several targets, multidimensional perception information data of the targets is obtained.

2. The multimodal sensing and detection method according to claim 1, characterized in that, The multimodal sensing data includes image data and radar data; Based on several target calibration parameter tables, each of the multimodal sensing data is corrected to obtain several target multimodal sensing data, including: Based on several target calibration parameter tables, the aberrations of each image data are calibrated to obtain several corrected image data, which include the original image data or the target image data.

3. The multimodal sensing and detection method according to claim 1, characterized in that, The multimodal sensing data includes image data and radar data; Based on several target calibration parameter tables, each of the multimodal sensing data is corrected to obtain several target multimodal sensing data, including: Based on several target calibration parameter tables, the image data and the corresponding radar data are correlated and aligned to obtain target multimodal data.

4. The multimodal sensing and detection method according to claim 1, characterized in that, The multimodal sensing data includes image data and radar data; Based on several target calibration parameter tables, the multimodal sensing data are corrected respectively to obtain several target multimodal sensing data, including: Based on several target calibration parameter tables, the aberrations of each image data are calibrated to obtain several corrected image data, which include the original image data or the target image data. Based on several target calibration parameter tables, the corrected image data are correlated and aligned with the corresponding radar data to obtain several target multimodal perception data.

5. The multimodal sensing and detection method according to any one of claims 2 to 4, characterized in that, The aberrations of each of the image data are corrected to obtain several corrected image data, including: Using the corresponding sensing spectral frequencies, several calibration parameter tables are created based on several sensing distances, and several target calibration parameter tables corresponding to several target sensing distances at the corresponding sensing spectral frequencies are obtained. Based on several target calibration parameter tables, each image data is corrected according to its corresponding spectral frequency to obtain several corrected image data.

6. The multimodal sensing and detection method according to any one of claims 2 to 4, characterized in that, The aberrations of each of the image data are corrected to obtain corrected image data, including: Based on several target calibration parameter tables, at least one aberration among distortion, field curvature, spherical aberration, chromatic aberration, and coma of each image data is corrected to obtain several target image data.

7. The multimodal sensing and detection method according to claim 1, characterized in that, Obtain a table of target calibration parameters corresponding to the perception distances of several targets, including: External radar equipment is used for calibration to obtain a table of target calibration parameters for the sensing range of several targets.

8. The multimodal sensing and detection method according to claim 1, characterized in that, Based on several target calibration parameter tables, each of the multimodal sensing data is corrected to obtain several target multimodal sensing data, which also includes: Based on several target calibration parameter tables, obtain the target calibration parameter table for each target perception distance in all target perception distances, specifically including: Determine whether the current target perception distance is the target perception distance in the existing target calibration parameter table; If so, then call the target calibration parameter table corresponding to the current target perception distance; If not, then interpolate the target calibration parameter tables of the two target perception distances adjacent to the current target perception distance to obtain the target calibration parameter table corresponding to the current target perception distance.

9. A multimodal sensing and detection system, characterized in that, The multimodal sensing and detection method applicable to any one of claims 1 to 8 includes: The data acquisition module is used to acquire multidimensional sensing information data of the measurement target and its environment. The multidimensional sensing information data includes several target sensing distances and several multimodal sensing data. Each multimodal sensing data corresponds to a target sensing distance. The multimodal sensing data includes at least image data. The target calibration parameter table acquisition module is used to acquire several target calibration parameter tables corresponding to the sensing distances of several targets. The calibration module is used to calibrate each of the multimodal sensing data based on several target calibration parameter tables to obtain several target multimodal sensing data. The target data acquisition module is used to acquire multi-dimensional perception information data of targets based on multi-modal perception data of several targets.

10. An electronic device, characterized in that, Includes a processor and a memory, wherein the memory is used to store code; The processor is configured to execute code in the memory to implement the multimodal sensing and detection method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the multimodal sensing and detection method according to any one of claims 1 to 8.