Apparatus and method for targetless lidar-camera calibration

WO2026197507A1PCT designated stage Publication Date: 2026-09-24SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/015694
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-05-23
Filing Date
2025-10-01
Publication Date
2026-09-24

Smart Images

  • Figure KR2025015694_24092026_PF_FP_ABST
    Figure KR2025015694_24092026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments disclosed in the present specification relate to an apparatus and a method for targetless LiDAR-camera calibration. According to one embodiment, the method for targetless LiDAR-camera calibration is disclosed, the method comprising the steps of: generating, on the basis of data obtained by a LiDAR sensor, an anchor Gaussian having a fixed center position, and a plurality of auxiliary Gaussians adjacent to the anchor Gaussian; rendering a 2D image of respective cameras on the basis of a 3D Gaussian including the anchor Gaussian and the plurality of auxiliary Gaussians and camera poses obtained on the basis of LiDAR-camera extrinsic parameters; and calculating loss on the basis of the rendered 2D image of the respective cameras and actual images obtained by the cameras, and jointly optimizing the LiDAR-camera extrinsic parameters and the 3D Gaussian on the basis of the calculated loss.
Need to check novelty before this filing date? Find Prior Art

Description

Non-target LiDAR-Camera Matching Device and Method

[0001] The embodiments disclosed in this specification relate to a lidar-camera matching device and method, and more specifically, to a non-target lidar-camera matching device and method for matching coordinates between a lidar and a camera without a dedicated object for matching.

[0002] This study was conducted as a result of the research on the "Artificial Intelligence Graduate School Support (Seoul National University) project of the Information and Communications Broadcasting Innovation Talent Development Project (IITP-II211343)" and the "Cm-level Error Space Restoration Technology Using Unstructured Images project of the Immersive Content Core Technology Development Project (IITP-00227993)" of the Ministry of Science and ICT and the Institute of Information and Communications Planning and Evaluation (IITP).

[0003] This study was conducted as a result of the "Realistic Rendering for Movable Omnidirectional Dynamic Scenes" project (NRF-00358701) of the Science and Engineering Research Infrastructure Construction Project of the Ministry of Education and the National Research Foundation of Korea (NRF).

[0004] With the recent advancement of Novel View Synthesis (NVS), various methods for reconstructing 3D scenes from 2D images are being researched. Specifically, the introduction of Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) techniques has enabled more precise scene rendering than before, and 3DGS, in particular, provides faster rendering performance compared to existing methods.

[0005] Despite these innovations, multi-sensor fusion of LiDAR and cameras, rather than utilizing only single-camera data, is essential for more precise image rendering and accurate acquisition of geometry information. However, to effectively apply neural rendering technology in a multi-sensor environment, the precise mounting position and orientation of each sensor—that is, sensor attitude information—is essential. In the case of autonomous vehicles or mobile robots, however, the initially corrected sensor attitude information may become slightly deformed over time due to vibration, shock, or temperature changes. This can degrade overall sensor alignment and adversely affect the perception or spatial reconstruction performance of autonomous driving.

[0006] To solve these problems, precise sensor alignment between the LiDAR and the camera is required, but conventional technology had the disadvantage of requiring a large-scale calibration infrastructure or a large space. A related prior art is Korean Patent Registration No. 10-2309608 (Title of Invention: Method for Coordinate System Alignment between LiDAR and Stereo-Camera).

[0007] Meanwhile, the aforementioned background technology is technical information that the inventor possessed for the derivation of the present invention or acquired during the process of deriving the present invention, and it cannot be considered as prior art disclosed to the general public prior to the filing of the present invention.

[0008] The embodiments disclosed in this specification aim to provide a non-target LiDAR-camera matching apparatus and method that perform scene representation optimization and LiDAR-camera matching together without a separate target.

[0009] The embodiments disclosed in this specification aim to provide a non-target LiDAR-camera matching device and method capable of stabilizing the global scale and translation amount, and resolving scene saturation that may occur during the process of co-optimizing the scene and camera attitude.

[0010] The embodiments disclosed in this specification are intended to provide a non-target LiDAR-camera matching device and method that can maintain consistency of the 3D Gaussian scale and stably provide high performance in indoor, outdoor, and various lighting environments.

[0011] According to one embodiment, as a technical means for achieving the technical problem described above, a non-target lidar-camera matching method is disclosed, comprising the steps of: generating an anchor Gaussian having a fixed center position and a plurality of auxiliary Gaussians around said anchor Gaussian based on data acquired by a lidar sensor based on data acquired by a lidar sensor; rendering a 2D image for each of said cameras based on a 3D Gaussian including said anchor Gaussian and said plurality of auxiliary Gaussians and the camera pose of said cameras acquired based on lidar-camera extrinsic parameters; calculating a loss based on the rendered 2D image for each of said cameras and the actual image acquired from said cameras, and jointly optimizing said lidar-camera extrinsic parameters and said 3D Gaussian based on the calculated loss.

[0012] According to another embodiment, a non-target LiDAR-camera matching device is disclosed, comprising a neural network model for generating a Gaussian, a memory for storing a program and data for LiDAR-camera matching, and at least one processor, and operates by executing a program stored in the memory, and generates an anchor Gaussian having a fixed center position and a plurality of auxiliary Gaussians around the anchor Gaussian based on data acquired by a LiDAR sensor, and renders a 2D image for each of the cameras based on a 3D Gaussian including the anchor Gaussian and the plurality of auxiliary Gaussians and the camera pose of the cameras acquired based on LiDAR-camera extrinsic parameters, and calculates a loss based on the rendered 2D image for each of the cameras and the actual image acquired from the cameras, and jointly optimizes the LiDAR-camera extrinsic parameters and the 3D Gaussian based on the calculated loss.

[0013] According to another embodiment, a computer-readable recording medium is disclosed that stores a program for performing a non-target lidar-camera matching method, which is performed by a non-target lidar-camera matching device, wherein the non-target lidar-camera matching method comprises the steps of: generating an anchor Gaussian having a fixed center position and a plurality of auxiliary Gaussians around the anchor Gaussian based on data acquired by a lidar sensor; rendering a 2D image for each of the cameras based on a 3D Gaussian including the anchor Gaussian and the plurality of auxiliary Gaussians and the camera pose of the cameras acquired based on lidar-camera extrinsic parameters; and calculating a loss based on the rendered 2D image for each of the cameras and the actual image acquired from the cameras, and jointly optimizing the lidar-camera extrinsic parameters and the 3D Gaussian based on the calculated loss.

[0014] According to another embodiment, a computer program stored on a computer-readable recording medium for performing a non-target lidar-camera matching method is disclosed, which is performed by a non-target lidar-camera matching device, and the non-target lidar-camera matching method comprises the steps of: generating an anchor Gaussian having a fixed center position and a plurality of auxiliary Gaussians around the anchor Gaussian based on data acquired by a lidar sensor; rendering a 2D image for each of the cameras based on a 3D Gaussian including the anchor Gaussian and the plurality of auxiliary Gaussians and the camera pose of the cameras acquired based on lidar-camera extrinsic parameters; and calculating a loss based on the rendered 2D image for each of the cameras and the actual image acquired from the cameras, and jointly optimizing the lidar-camera extrinsic parameters and the 3D Gaussian based on the calculated loss.

[0015] According to any one of the aforementioned means for solving the problem, by performing scene representation optimization and LiDAR-camera alignment together without a separate target, hardware installation costs are reduced and automated alignment is possible.

[0016] In addition, according to any one of the aforementioned means for solving the problem, the global scale and translation amount can be stabilized by fixing the voxel center obtained based on the point cloud based on the LiDAR data as the anchor point, which is the center point of the anchor Gaussian, and by introducing an auxiliary Gaussian, scene saturation phenomena that may occur locally during the process of jointly optimizing the scene and camera attitude can be prevented.

[0017] In addition, according to any one of the aforementioned problem-solving means, consistency of the 3D Gaussian scale can be maintained and no separate target is required, so high performance can be stably provided in indoor, outdoor, and various lighting environments.

[0018] The effects obtainable from the disclosed embodiments are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the disclosed embodiments belong from the description below.

[0019] Figure 1 is a diagram illustrating an example of a LiDAR-camera system.

[0020] FIG. 2 is a block diagram illustrating a non-target lidar-camera matching device according to one embodiment.

[0021] FIG. 3 is a schematic diagram of the lidar-camera matching process according to one embodiment.

[0022] FIG. 4 is a flowchart illustrating a non-target lidar-camera matching method according to one embodiment.

[0023] Various embodiments are described in detail below with reference to the attached drawings. The embodiments described below may be implemented in various different forms. In order to explain the features of the embodiments more clearly, detailed descriptions of matters widely known to those skilled in the art to which the following embodiments belong have been omitted. Additionally, parts of the drawings unrelated to the description of the embodiments have been omitted, and similar parts throughout the specification have been given similar reference numerals.

[0024] Throughout the specification, when a configuration is described as being "connected" to another configuration, this includes not only cases where they are "directly connected," but also cases where they are "connected with another configuration in between." Furthermore, when a configuration is described as "including" another configuration, this means that, unless specifically stated otherwise, it does not exclude other configurations but may include additional configurations.

[0025] The embodiments will be described in detail below with reference to the attached drawings.

[0026] Before that, the terms used in this specification will be explained.

[0027] Lidar-camera calibration refers to aligning multiple different coordinate systems. In other words, it refers to the process of aligning extrinsic parameters between a camera and a lidar (hereinafter referred to as “Lidar-camera extrinsic parameters”). Lidar-camera extrinsic parameters are parameters that describe the transformation relationship between the camera coordinate system and the lidar coordinate system, and may consist of a rotation matrix and a translation vector between the two coordinate systems. Lidar-camera calibration in this specification refers to non-target Lidar-camera calibration that does not use a dedicated target. Unless otherwise specified below, the Lidar-camera calibration device in this specification is a non-target Lidar-camera calibration device, and the Lidar-camera calibration method in this specification is a non-target Lidar-camera calibration method.

[0028] 'Camera pose' is information regarding the position and orientation of the camera in 3D space, and can be defined by translation in 3D space and rotation about the coordinate axes (i.e., camera pose is defined as pose(rotation, translation). 'Translation' is a vector representing the translation between the coordinates before and after the change when the camera's position in 3D space is changed, and 'rotation' can mean a rotation value about the xyz axes.

[0029] From now on, unless otherwise noted, the term "pose" shall refer to the camera's pose.

[0030] A 'world coordinate system' refers to a coordinate system in real space used as a reference when expressing the position of an object, meaning a coordinate system with an arbitrary point as the origin, while a 'camera coordinate system' refers to a coordinate system based on the camera. For example, in a camera coordinate system, the camera's focus can be the origin, the front optical axis can be the Z-axis, the direction to the right of the camera the X-axis, and the direction downward the Y-axis.

[0031] LiDAR (Light Detection and Ranging) is a distance measuring device that measures the distance to a target. Distance can be measured by sending out short laser pulses and recording the time elapsed between the outgoing light pulse and the reflected (backscattered) light pulse. Since LiDAR senses objects using light reflected from a laser, the laser points reflecting the target object are arranged like points, which is called a point cloud. In this specification, LiDAR and LiDAR sensor may be used interchangeably.

[0032] A LiDAR-camera matching device is a device that performs matching between multiple cameras and LiDAR sensors positioned at different locations and facing different directions. In other words, it can mean matching the camera's attitude based on the LiDAR sensor.

[0033] Figure 1 is a diagram illustrating an example of a LiDAR-camera system.

[0034] Referring to FIG. 1, the lidar-camera system (10) may include a plurality of camera sensors (30A, 30B, 30C, hereinafter referred to as 'cameras') installed at different locations, a lidar sensor (20), and a lidar-camera matching device (100).

[0035] The lidar-camera system (10) may be a robot (10) that recognizes objects, moves, and performs tasks based on data obtained from a plurality of cameras (30A, 30B, 30C) and a lidar sensor (20) as shown in FIG. 1, but is not limited thereto and may be a vehicle capable of autonomous driving.

[0036] The LiDAR sensor (20) is a distance sensor that measures the position coordinates of an object by measuring the time it takes for light to be emitted and reflected, and can provide 3D spatial information. The LiDAR sensor (20) can provide measurement data along with a timestamp indicating the measurement time.

[0037] Multiple cameras (30A, 30B, 30C) are installed in the front, rear, side, etc. of the robot and can capture objects or landscapes located in the direction of the cameras to generate video or images. The video from the multiple cameras (30A, 30B, 30C) may include multiple frame images, and at least one of the frame images included in the video or the captured image may be used to calculate loss with a rendered 2D image.

[0038] The LiDAR-camera matching device (100) can create a 3D scene (hereinafter referred to as "scene") based on data obtained from multiple cameras (30A, 30B, 30C) and a LiDAR sensor (20). There is an advantage that a more precise scene can be created when using a multi-sensor fusion technique that uses multiple cameras and LiDAR sensors together, compared to using a single camera or LiDAR sensor.

[0039] However, for this purpose, accurate orientation of the lidar sensor and camera is required. However, in the case of the movable robot or autonomous vehicle of Fig. 1, the orientation information of the lidar sensor and camera initially set may be slightly deformed over time due to vibration, shock, temperature changes, etc. Therefore, it is necessary to align the orientation information of the lidar sensor and the orientation information of the camera in real time.

[0040] A lidar-camera matching device (100) according to one embodiment can perform matching between a lidar sensor and a plurality of camera sensors in a targetless manner without using a dedicated target. Specifically, the lidar-camera matching device can calibrate the camera attitude relative to the lidar coordinate system by using the lidar sensor as a reference sensor and the lidar coordinate system as a global reference.

[0041] In particular, a LiDAR-camera matching device (100) according to one embodiment can improve both rendering fidelity and alignment accuracy of the pose by using a rendering-based approach that utilizes a rendering pipeline for automatic correction between multiple cameras and LiDAR sensors to refine the scene representation and the sensor pose in common.

[0042] That is, the lidar-camera matching device (100) can perform scene implementation and update the camera attitude for each of the plurality of cameras simultaneously based on sensor data obtained from the lidar sensor and images obtained through the camera.

[0043] As described above, the LiDAR-camera matching device (100) can be implemented as an electronic terminal or a server-client system.

[0044] At this time, the electronic terminal can be implemented as a laptop, portable terminal, wearable device, etc., which can connect to a remote server (20) via a network or connect to other electronic terminals and servers. Here, the laptop includes, for example, a laptop equipped with a web browser, etc., and the portable terminal can include, for example, a wireless communication device that guarantees portability and mobility, such as a PCS (Personal Communication System), PDC (Personal Digital Cellular), PHS (Personal Handyphone System), PDA (Personal Digital Assistant), GSM (Global System for Mobile communications), IMT (International Mobile Telecommunication)-2000, CDMA (Code Division Multiple Access)-2000, W-CDMA (W-Code Division Multiple Access), Wibro (Wireless Broadband Internet), a smartphone, a mobile WiMAX (Mobile Worldwide Interoperability for Microwave Access), etc., all kinds of handheld-based wireless communication devices. In addition, wearable devices are types of information processing devices that can be worn directly on the human body, such as watches, glasses, accessories, clothing, and shoes, and can connect to a server at a remote location or to another terminal via a network, either directly or through another information processing device.

[0045] And the server can be implemented as a computing device capable of communicating with electronic terminals and networks, or as a cloud computing server.

[0046] FIG. 2 is a block diagram illustrating the configuration of a non-target lidar-camera matching device according to one embodiment.

[0047] Referring to FIG. 2, a lidar-camera matching device (100) according to one embodiment may include a memory (110), a control unit (120), a communication unit (130), and an input / output unit (140).

[0048] A program for LiDAR-camera alignment can be installed and stored in the memory (110). Specifically, a program for attitude alignment or alignment between the LiDAR and the camera can be installed and stored in the memory (110), and a neural network model for generating a Gaussian, a small multilayer perceptron implemented in two layers to generate an auxiliary Gaussian, a program for 2D image rendering, a program for calculating optical loss, scale normalization loss and total loss, etc., and a program for optimizing scene representation and camera attitude based thereon can be installed and stored in the memory (110).

[0049] Additionally, data for LiDAR-camera matching may be stored in the memory (110). For example, LiDAR data obtained from a LiDAR sensor, loss function and learning rate obtained from multiple cameras for LiDAR-camera matching, number of auxiliary Gaussians, threshold of scale normalization loss, etc. may be stored.

[0050] The control unit (120) includes at least one processor such as a CPU, GPU, etc., and can perform the LiDAR-camera matching method presented below by executing a program stored in memory (110). For example, the control unit (120) can perform the method by executing a program stored in memory (110) by a processor.

[0051] Additionally, the control unit (120) can control other components included in the LiDAR-camera matching device (100). For example, the control unit (120) can provide the matching result to the user through the input / output unit (140).

[0052] The communication unit (130) can receive training data required for lidar-camera matching or receive a program for lidar-camera matching by performing wired or wireless communication with another device or network.

[0053] To this end, the communication unit (130) may include a communication module that supports at least one of various wired and wireless communication methods, and the communication module may be implemented in the form of a chipset. The wireless communication supported by the communication unit (130) may be, for example, WiFi (Wireless Fidelity), Wi-Fi Direct, Bluetooth, UWB (Ultra Wide Band), or NFC (Near Field Communication).

[0054] The input / output unit (140) may include an output device such as a display panel or wearable display device and a speaker for displaying an interface, and may include various types of input devices (e.g., keyboard, touchscreen, camera, etc.) for receiving input from a user.

[0055] Hereinafter, a LiDAR-camera alignment method according to an embodiment in which the control unit (120) operates by executing a program is described in detail. Unless otherwise specifically stated, the processes described below are performed by the control unit (120) executing a program stored in the memory (110).

[0056] FIG. 3 is a schematic diagram of the lidar-camera matching process according to one embodiment.

[0057] The control unit (120) uses a global reference for the lidar coordinate system and can relatively correct the coordinate systems of all cameras based on the lidar coordinate system.

[0058] This is because LiDAR sensors generally provide a 360-degree field of view (FoV), enabling rich overlapping coverage between multiple cameras, and because LiDAR sensors provide highly accurate 3D geometric information, they have higher reliability than cameras in estimating consistent motion between frames.

[0059] Referring to FIG. 3, in the initialization step, when a lidar sensor and a plurality of cameras are arranged to have a specific view, the control unit (120) can obtain the attitude of the arranged lidar sensor and the initial camera attitudes of each of the plurality of cameras.

[0060] At this time, the control unit (120) can acquire an initial camera attitude based on the position on the CAD or acquire an initial camera attitude based on the position (attitude) of the lidar sensor. The control unit (120) can store lidar-camera external parameters by performing rough calibration between the lidar and the camera based on the acquired initial camera attitude and the attitude of the lidar sensor.

[0061] Next, the control unit (120) has a series of timestamps ( A LiDAR scan can be acquired in ), and the acquired scans can be aggregated to acquire point clouds. For example, the control unit (120) uses a LiDAR odometry to each timestamp A LiDAR scan for can be obtained.

[0062] This can be expressed as a formula as shown in Mathematical Formula 1 below.

[0063] [Mathematical Formula 1]

[0064]

[0065] In mathematical formula 1 ( ) is a point cloud, and can mean the length of the time being integrated, and is a timestamp It could be a LiDAR scan.

[0066] For example, the control unit (120) can use a LiDAR SLAM (Chen et al, “ig-lio: An incremental gicp-based tightly-coupled lidar-inertial odometry.”, IEEE Robotics and Automation Letters, 9(2): 1883-1890, 2024) technique or an ICP (Ignacio Vizzo et al, Kiss-icp: In defense of point-to-point icp-simple, accurate, and robust registration if done the right way IEEE Robotics and Automation Letters, 8(2):1029-1036, 2023) technique to align consecutive lidar scans and then integrate the lidar scans to generate a point cloud. Accordingly, global alignment and geometric consistency can be ensured prior to registration.

[0067] Additionally, the control unit (120) can estimate the LiDAR SLAM technique or the pose of the LiDAR sensor.

[0068] In one embodiment, since a LiDAR sensor is set as a reference sensor, the integrated point cloud can be considered as prior information of the global shape. Accordingly, the control unit (120) can generate 3D Neural Gaussians based on the integrated point cloud. The 3D Neural Gaussians may include frozen anchor Gaussians and a plurality of auxiliary Gaussians.

[0069] 3D Gaussian It can be expressed as shown in mathematical formula 2 below.

[0070] [Mathematical Formula 2]

[0071]

[0072] In mathematical formula 2 ( ) is the center point of the 3D Gaussian, ( ) is the anisotropic covariance matrix, ( ) is transparency (opacity), represents the spherical harmonic coefficients that model the view-dependent appearance, i.e., color. Covariance matrix It can be expressed as in mathematical formula 3.

[0073] [Mathematical Formula 3]

[0074]

[0075] In mathematical formula 3 ( ) is a rotation matrix, and ( ) is a scale matrix, and is a diagonal matrix (diag; ) and the scale factor for each axis ( Includes ).

[0076] The control unit (120) can set an anchor point, which is the center point of the anchor Gaussian, based on the integrated point cloud, and generate an anchor Gaussian at the set anchor point using a learned neural network model.

[0077] The neural network model is other learnable parameters of the anchor Gaussian excluding the center point. , , , By predicting, an anchor Gaussian according to mathematical formula 2 can be generated.

[0078] For example, the control unit (120) is Instant-NGP ( The model proposed by et al., “Instant neural graphics primitives with a multiresolution hash encoding”, ACM TOG, vol. 41, no. 4, pp. 1-15, 2022.) is used as a neural network model, and this model adds four heads to the output of a multilayer perceptron (MLP). , , , It is possible to predict and generate an anchor Gaussian based on the input anchor points.

[0079] To determine the anchor point, the control unit (120) uses an integrated point cloud voxelized into a set of voxel centers It is possible to obtain the voxel centers and determine the obtained centers as anchor points, which are the centers (center points) of the anchor Gaussian.

[0080] Specifically, the control unit (120) calculates the scale of the overall 3D scene based on bounding boxes extracted from LiDAR center points, and the control unit (120) calculates the size of the voxels based on the scale of the overall scene. It can produce.

[0081] [Mathematical Formula 4]

[0082]

[0083] In mathematical formula 4 The scale of the entire scene, is a fixed constant.

[0084] In addition, the control unit (120) has a voxel size point cloud integrated into Partitioning and the set of voxel centers according to mathematical formula 5 You can obtain.

[0085] [Mathematical Formula 5]

[0086]

[0087] In mathematical formula 2 means element-wise floor operation, and Is i-th Voxel Center represents a point belonging to a point cloud. The control unit (120) is each voxel center By setting it as the center point of the anchor Gaussian, a globally consistent reference can be provided in the real-world coordinate system.

[0088] As described above, unlike conventional techniques that dynamically adjust or expand the position of an anchor point, in one embodiment, by fixing the anchor point, which is the center point of the anchor Gaussian, at the center of the voxel, the scene scale can be maintained and significant translational drift can be prevented during the alignment process.

[0089] Meanwhile, the control unit (120) can maintain global alignment while minimizing artifacts by removing anchors caused by LiDAR noise when anchor Gaussians with consistently low transparency occur during training, by classifying them with a plotter and removing them from the set.

[0090] The control unit (120) is the anchor point, which is the center point of the anchor Gaussian. Auxiliary Gaussians can be generated for each. This is intended to refine the local geometry and mitigate the risk of convergence to a non-optimal solution while the anchor Gaussian remains static.

[0091] Auxiliary Gaussians, like anchor Gaussians, can be placed around the anchor Gaussians as 3D Gaussians according to Equation 2, and for example, K auxiliary Gaussians can be placed around each anchor Gaussian. For example, K can be 8 or 10.

[0092] The control unit (120) can predict the parameters of the auxiliary Gaussian for each anchor Gaussian using a small multilayer perceptron (MLP). In this case, a separate multilayer perceptron may be provided for each parameter of the auxiliary Gaussian.

[0093] For example, the control unit (120) is a small multilayer perceptron Using [it] according to mathematical formula 6 A set of offsets defined as It can predict.

[0094] [Mathematical Formula 6]

[0095]

[0096] In mathematical formula 6 is the anchor Gaussian It is an anchor-associated feature vector, and means encoding factors dependent on the camera view (e.g., relative distance or field of view, camera direction, etc.), and means a trainable scale.

[0097] For example, the control unit (120) It can be obtained by sampling points around the anchor point within the point cloud immediately after selecting the anchor point, and encoding the extracted points. In addition, as an example, can mean the relative distance between the camera and the anchor point and the direction of the camera, and the control unit (120) can be directly encoded based on the position of the camera and the position of the anchor point.

[0098] Similar to the offset set, the other auxiliary Gaussian parameters, such as color, transparency, and rotation, also It can be predicted by a multilayer perceptron under the condition, where the covariance matrix, which is the parameter of another auxiliary Gaussian, , color , transparency Each can be predicted in an individual multilayer perceptron.

[0099] The offset is used to determine the position of auxiliary Gaussians to be placed around the anchor Gaussian, for example, the center of each auxiliary Gaussian Is It could be.

[0100] These auxiliary Gaussians act as adjustable buffers around the anchor Gaussians to adapt local geometry and appearance when the camera pose deviates from the initial prediction. Even if the system starts with an inaccurate external configuration, the control unit (120) can move or rescale the auxiliary Gaussians to adjust for discrepancy between the rendered 2D image and the observed image, thereby inducing the system to move away from an inappropriate local minimum.

[0101] Meanwhile, the control unit (120) receives camera poses obtained by existing LiDAR-camera external parameters and generates a rendered 2D image by splatting a 3D Gaussian, that is, an anchor Gaussian and an auxiliary Gaussian. It can calculate a loss based on the generated image and a measured image obtained from an actual camera, and then jointly optimize the LiDAR-camera external parameters and the 3D Gaussian based on the calculated loss. At this time, the control unit (120) can update the camera pose based on the generated image and the measured image obtained from an actual camera, and update the LiDAR-camera external parameters based on the updated camera pose.

[0102] The control unit (120) renders a 2D image with the camera pose of each of the plurality of cameras, and can update the attributes of the anchor Gaussian and auxiliary Gaussian and the LiDAR-camera extrinsic parameters based on the total loss between the image of the camera actually captured and the rendered 2D image. As described above, the control unit (120) can obtain the camera pose for each of the plurality of cameras by the LiDAR-camera extrinsic parameters set (initialized) based on the CAD or LiDAR pose, and render a 2D image based on the obtained camera pose.

[0103] General camera pose This refers to the position of the camera in the world coordinate system, It can also be expressed as (i.e., ).Camera posture Is It can be defined as, ( ) is a rotation matrix, ( ) can mean a translated vector.

[0104] Each of the multiple cameras in FIG. 1 may have a camera attitude. In this case, the camera attitude for each of the multiple cameras refers to a relative position in a LiDAR coordinate system based on the attitude of the LiDAR sensor. That is, while a typical camera attitude is the position of the camera in a world coordinate system, the camera attitude in one embodiment may be the relative position of the camera in a LiDAR coordinate system where the position of the LiDAR sensor (i.e., the LiDAR attitude) is the origin (i.e., a world coordinate system where the LiDAR attitude is the origin), and the camera attitudes for each of the multiple cameras may be different from each other.

[0105] Accordingly, the control unit (120) controls the camera positions of N multiple cameras. It can be obtained from LiDAR-to-Camera extrinsic parameters having N elements, and the LiDAR-to-Camera extrinsic parameters of each camera Is It can be defined as follows. Here, is a LiDAR pose, which can be obtained from SLAM, and is defined as the camera pose in the world coordinate system. Consequently, represents a matrix that converts the LiDAR coordinate system to the camera coordinate system.

[0106] Meanwhile, the control unit (120) is a camera position Based on, 3D Gaussians such as anchor Gaussians and auxiliary Gaussians 2D Gaussian on the image plane It can be projected as.

[0107] Below, for convenience, one camera A process for generating (rendering) one 2D image for a plurality of cameras will be described. A 2D image for a plurality of cameras can be generated by the control unit (120) carrying out the process described below simultaneously or sequentially.

[0108] The control unit (120) is the camera position Based on this, the camera center can be transformed into the world coordinate system (i.e., the LiDAR coordinate system), and the center of 3D Gaussians, such as the anchor Gaussian and auxiliary Gaussian, can be transformed into the 2D camera coordinate system.

[0109] For example, camera 3D Gaussian centers such as anchor Gaussians and auxiliary Gaussians in a 2D image plane for silver It can be converted to the camera coordinate system, and the camera center in the world coordinate system. silver It can be given as follows.

[0110] Additionally, the control unit (120) may perform frustum filtering (or culling) before projecting the 3D Gaussian onto the 2D Gaussian. Frustum filtering is performed on the camera Only 3D Gaussians within the frustum, which is the area that can actually be seen, are retained, and 3D Gaussians outside the frustum, namely anchor Gaussians and auxiliary Gaussians outside the frustum, can be excluded from the projection target.

[0111] Next, the control unit (120) is the Jacobian of the Local Projective Transformation Using, A 2D covariance matrix defined as Calculate the 3D Gaussian that became the projection target 2D Gaussian on a 2D image plane It can be projected as.

[0112] Finally, the control unit (120) determines the pixel position through sequential alpha blending illustrated in Equation 7. The final rendered color You can generate a rendered 2D image by calculating it.

[0113] [Mathematical Formula 7]

[0114]

[0115] Here is the transparency of the Gaussian, is a 3D Gaussian color, is a projected 2D Gaussian pixel The value at, The previous Gaussians are pixels It means the probability of passing through.

[0116] The control unit (120) is a camera The 2D image rendered for (Rendered) and the actual camera Based on the image captured by, i.e., the actual measurement image (Reference), a 3D Gaussian as shown in Equation 8 class LiDAR-camera extrinsic parameters of dogs cameras It can be jointly optimized.

[0117] [Mathematical Formula 8]

[0118]

[0119] In mathematical formula 8 is a rendered 2D image, is total loss, each camera Obtained by With dog images, It can be expressed as.

[0120] To this end, the control unit (120) has the parameters of the anchor Gaussian and the parameters of the auxiliary Gaussian and the LiDAR-camera external parameters in a direction in which the total loss is minimized. You can update it.

[0121] Meanwhile, the LiDAR-camera extrinsic parameters in Equation 8 It can be updated as in mathematical formula 9.

[0122] [Mathematical Formula 9]

[0123]

[0124] In mathematical formula 9 can mean the step size, that is, the learning rate, and This is photometric loss, which can refer to the difference between a rendered 2D image and an image captured by a real camera.

[0125] That is, the control unit (120) controls the LiDAR-camera external parameters in a direction that minimizes photometric loss between the rendered 2D image and the actual image according to mathematical formula 9. Optimize, but the step size (learning rate) You can update as much as possible.

[0126] one side, is camera posture ( It is a differentiable function with respect to ). As previously mentioned, LiDAR-camera extrinsic parameter Is Therefore, the control unit (120) determines the camera position of each of the plurality of cameras based on the process described below. Update and camera position LiDAR-camera extrinsic parameters based on The camera pose can be updated in the direction of the gradient through an analytical Jacobian derivation method, as described in Optical Loss Gaussian Splatting SLAM (Matsuki et al, “Gaussian splatting slam”. In IEEE Conf. Comput. Vis. Pattern Recog. (CVPR), pages 18039-18048, 2024). Specifically, the control unit (120) updates the camera pose in the direction of the gradient through an analytical Jacobian derivation method. You can update it, and then the updated camera attitude LiDAR-camera extrinsic parameters based on You can update it.

[0127] Using the chain rule, camera position Optical loss for The gradient of can be expressed as in Equation 10. Unless otherwise noted, silver It is considered to be the same as.

[0128] [Mathematical Formula 10]

[0129]

[0130] In mathematical formula 10 can be the center of a 2D Gaussian in the 2D image domain, that is, in a 2D plane. class It is a Jacobian that can follow the Standard Rigid Body Transformation as shown in Equation 11 below.

[0131] [Mathematical Formula 11]

[0132]

[0133] Also, the 2D covariance matrix of Equation 10 is the projection Jacobian and camera rotation Since it depends on, as in mathematical equation 12 camera stance The gradient for is given by Equation 12.

[0134] [Mathematical Formula 12]

[0135]

[0136] In mathematical formula 12, Is Standard basis vector in the group ( It can be derived as in mathematical formula 13 using ).

[0137] [Mathematical Formula 13]

[0138]

[0139] one side, Since it depends on the viewing direction, in Equation 10 camera stance Gradient for It is equal to mathematical formula 14.

[0140] [Mathematical Formula 14]

[0141]

[0142] In summary, the control unit (120) has a 6D tangent vector as described in Equation 15 ( Using ) Direct camera position in space It can be updated. And the control unit (120) updates the camera position. and LiDAR attitude (i.e., attitude of the LiDAR sensor) LiDAR-camera extrinsic parameters based on You can update it.

[0143] [Mathematical Formula 15]

[0144]

[0145] In mathematical formula 15 is the learning rate, which can be a hyperparameter, and the step size in Equation 8. It can be the same as, and the learning rate can be dynamically adjusted to achieve optimal alignment between the lidar point cloud and the camera image.

[0146] Meanwhile, the camera of When an update is triggered by the nth image, the control unit (120) [controls] the corresponding camera The camera pose can be updated in a rig-based manner, where the updated camera pose is consistently applied to all rendered 2D images. This ensures that the camera pose maintains internal consistency within the set of images of that camera, thereby maintaining the geometric alignment of the rig.

[0147] Meanwhile, the total loss function of Equation 8 is the luminance loss that measures image reconstruction quality, as in Equation 16. Scale normalization loss, a normalization term that prevents over-degenerate Gaussian shapes It may include two terms.

[0148] [Mathematical Formula 16]

[0149]

[0150] Scale normalization loss can be introduced to limit the anisotropy of each 3D Gaussian, and scale normalization loss can prevent the 3D Gaussian from collapsing into an excessively thin or sharp shape during the training process.

[0151] Scale normalization loss can be defined by imposing a restriction on the ratio between the maximum scale component and the minimum scale component of a 3D Gaussian, as shown in Equation 17.

[0152] [Mathematical Formula 17]

[0153]

[0154] In mathematical formula 17 ( ) is a Gaussian scale vector, and each element can produce a spatial scale (magnitude, size) for each corresponding axis. represents the set of Gaussians remaining after frustum filtering, and is a predefined scale normalization threshold in the hyperparameter settings, used when there are no remaining 3D Gaussians after frustum filtering (i.e., =0), this term can be evaluated to 0.

[0155] Scale normalization loss is applied only to 3D Gaussians that survive the frustum filtering (or culling) step performed before the projection of 2D Gaussians, which ensures that it affects only active contributing Gaussians.

[0156] As described above, a LiDAR-camera matching device according to one embodiment can prevent degenerate Gaussians that become extremely thin along one or multiple axes by applying smooth constraints to the spectral ratios of each Gaussian through scale normalization. In addition, by maintaining balance in scale, numerical stability is maintained during the optimization process, while allowing the Gaussian to flexibly adapt to the local geometry of the scene.

[0157] As described above, a LiDAR-camera matching device according to one embodiment can stabilize the global scale and translation amount by acquiring a point cloud based on data acquired by a LiDAR sensor, voxelizing the point cloud, and fixing the acquired voxel center as the anchor point, which is the center point of an anchor Gaussian. In addition, by introducing an auxiliary Gaussian, scene saturation phenomena that may occur locally during the process of jointly optimizing the scene and LiDAR-camera external parameters can be prevented.

[0158] In addition, by introducing scale normalization loss in addition to optical loss, which is the difference between the actual image and the rendered 2D image, consistency of the 3D Gaussian scale can be enforced, and since no separate target is required, high performance can be reliably provided in indoor, outdoor, and various lighting environments.

[0159] In the embodiments above, the term 'part' refers to a software or hardware component such as an FPGA (field programmable gate array) or an ASIC, and the 'part' performs certain roles. However, the meaning of 'part' is not limited to software or hardware. The 'part' may be configured to reside in an addressable storage medium or may be configured to run one or more processors. Accordingly, as an example, the 'part' includes components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables.

[0160] The functions provided within the components and 'parts' can be combined into fewer components and 'parts' or separated from additional components and 'parts'.

[0161] In addition, the components and '~parts' may be implemented to play one or more CPUs within the device or secure multimedia card.

[0162] Meanwhile, FIG. 4 is a flowchart illustrating a non-target lidar-camera alignment method according to one embodiment.

[0163] The non-target lidar-camera matching method illustrated in FIG. 4 includes steps processed sequentially in the non-target lidar-camera matching device (100) illustrated in FIG. 1 to 3. Therefore, even if details are omitted below, the description above regarding the non-target lidar-camera matching device (100) illustrated in FIG. 1 to 3 can also be used in the non-target lidar-camera matching method according to one embodiment illustrated in FIG. 4.

[0164] Referring to FIG. 4, the lidar-camera matching device (100) can generate an anchor Gaussian having a fixed center position and a plurality of auxiliary Gaussians around the anchor Gaussian based on data acquired by the lidar sensor (S410).

[0165] Specifically, the lidar-camera matching device (100) receives lidar scan values ​​measured by the lidar sensor at timestamp t, and can align and integrate lidar scan values ​​over a predetermined period of time to generate a point cloud.

[0166] The LiDAR-camera matching device (100) can voxelize the generated point cloud according to Equation 5 to obtain a set of voxel centers, and determine the obtained voxel centers as anchor points, which are the center positions of the anchor Gaussian.

[0167] The LiDAR-camera matching device (100) can generate an anchor Gaussian at an anchor point according to Equation 2 using a neural network model, and generate K multiple auxiliary Gaussians around the anchor Gaussian. The auxiliary Gaussians are also 3D Gaussians according to Equation 2, and the LiDAR-camera matching device (100) can obtain the properties of the auxiliary Gaussians, i.e., the parameters of the auxiliary Gaussians, using individual small multilayer perceptrons (MLPs). The parameters of the auxiliary Gaussians are as described in Equation 6 Under the conditions of each, it can be predicted by a separate multilayer perceptron (MLP).

[0168] Next, the lidar-camera matching device (100) can render a 2D image for each of the cameras based on the camera pose of the cameras obtained by a 3D Gaussian including an anchor Gaussian and a plurality of auxiliary Gaussians and lidar-camera external parameters (S420).

[0169] Specifically, the lidar-camera matching device (100) can project a 3D Gaussian, i.e., an anchor Gaussian and a plurality of auxiliary Gaussians, onto a 2D image plane corresponding to each of the plurality of cameras based on the camera attitude for each of the plurality of cameras obtained by the existing lidar-camera external parameters, and render a 2D image for each of the plurality of cameras by calculating a pixel value for a pixel position according to Equation 7.

[0170] Next, the lidar-camera matching device (100) can jointly optimize lidar-camera extrinsic parameters and 3D Gaussian based on the loss obtained from the rendered 2D image and the actual image obtained from the cameras (S430).

[0171] Specifically, the LiDAR-camera matching device (100) can calculate a scale normalization loss defined by a maximum scale component and a minimum scale component according to Equation 17, and calculate an optical loss calculated based on a rendered 2D image and a measured image to calculate a total loss according to Equation 16.

[0172] And the LiDAR-camera matching device (100) can jointly optimize the 3D Gaussian, including the anchor Gaussian and auxiliary Gaussian, and the LiDAR-camera extrinsic parameters in a direction in which the total loss is minimized according to Equation 8. At this time, the LiDAR-camera matching device (100) can update the camera attitudes of the cameras in a direction in which optical loss is minimized, and update the LiDAR-camera extrinsic parameters based on the updated camera attitudes. In addition, the LiDAR-camera matching device (100) can jointly optimize the LiDAR-camera extrinsic parameters and the 3D Gaussian in a direction in which the total loss is minimized, and can optimize the 3D Gaussian by updating the parameters of the 3D Gaussian, that is, the parameters of the anchor Gaussian and the parameters of a plurality of auxiliary Gaussians.

[0173] At this time, the LiDAR-camera matching device (100) can be updated in a rig-based manner in which the camera pose updated for a specific camera is consistently applied to the entire rendered 2D image. This ensures that the camera pose maintains internal consistency within the set of images of the camera, thereby maintaining the geometric alignment of the rig.

[0174] As described above, a LiDAR-camera matching method according to one embodiment can stabilize the global scale and translation amount by acquiring a point cloud based on data acquired by a LiDAR sensor, voxelizing the point cloud, and fixing the acquired voxel center as the anchor point, which is the center point of an anchor Gaussian. In addition, by introducing an auxiliary Gaussian, scene saturation phenomena that may occur locally during the process of jointly optimizing the scene and camera attitude can be prevented.

[0175] In addition, by introducing scale normalization loss in addition to optical loss, which is the difference between the actual image and the rendered 2D image, consistency of the 3D Gaussian scale can be enforced, and since no separate target is required, high performance can be reliably provided in indoor, outdoor, and various lighting environments.

[0176] A LiDAR-camera matching method according to one embodiment described through FIG. 4 may also be implemented in the form of a computer-readable medium that stores instructions and data executable by a computer. In this case, the instructions and data may be stored in the form of program code, and when executed by a processor, may generate a specific program module to perform a specific operation. Furthermore, the computer-readable medium may be any available medium accessible by a computer and includes both volatile and non-volatile media, as well as removable and non-removable media. Additionally, the computer-readable medium may be a computer recording medium, which may include both volatile and non-volatile, removable and non-removable media implemented by any method or technique for storing information such as computer-readable instructions, data structures, program modules, or other data. For example, the computer recording medium may be a magnetic storage medium such as an HDD and an SSD, an optical recording medium such as a CD, DVD, and Blu-ray disc, or a memory included in a server accessible via a network.

[0177] In addition, the LiDAR-camera matching method according to one embodiment described through FIG. 4 may be implemented as a computer program (or computer program product) comprising instructions executable by a computer. The computer program includes programmable machine instructions processed by a processor and may be implemented in a high-level programming language, an object-oriented programming language, assembly language, or machine language, etc. Additionally, the computer program may be recorded on a tangible computer-readable recording medium (e.g., memory, hard disk, magnetic / optical medium, or SSD (Solid-State Drive), etc.).

[0178] Accordingly, the LiDAR-camera matching method according to one embodiment described through FIG. 4 can be implemented by executing a computer program as described above by a computing device. The computing device may include at least some of a processor, memory, a storage device, a high-speed interface connected to the memory and a high-speed expansion port, and a low-speed interface connected to a low-speed bus and a storage device. Each of these components is connected to one another using various buses and may be mounted on a common motherboard or mounted in other suitable ways.

[0179] Here, the processor can process instructions within the computing device, such as instructions stored in memory or storage devices to display graph information for providing a Graphic User Interface (GUI) on external input and output devices, such as a display connected to a high-speed interface. In another embodiment, a plurality of processors and / or a plurality of buses may be used together with a plurality of memories and memory types as appropriate. Additionally, the processor may be implemented as a chipset comprising a plurality of independent analog and / or digital processors.

[0180] In addition, memory stores information within a computing device. For example, memory may consist of volatile memory units or a set thereof. As another example, memory may consist of non-volatile memory units or a set thereof. Furthermore, memory may be other forms of computer-readable media, such as magnetic or optical discs.

[0181] And memory can provide a large amount of storage space to a computing device. Memory may be a computer-readable medium or a configuration containing such a medium, and may include, for example, devices or other configurations within a Storage Area Network (SAN), and may be a floppy disk device, a hard disk device, an optical disk device, or a tape device, flash memory, or other similar semiconductor memory device or device array.

[0182] The embodiments described above are for illustrative purposes only, and those skilled in the art will understand that the embodiments described above can be easily modified into other specific forms without altering the technical concept or essential features of the embodiments described above. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, each component described as a single unit may be implemented in a distributed manner, and components described as distributed may likewise be implemented in a combined form.

[0183] The scope of protection sought through this specification is defined by the claims set forth below rather than by the detailed description above, and should be interpreted to include all modifications or variations derived from the meaning and scope of the claims and the concept of equivalents.

Claims

In a non-target lidar-camera matching method performed by a non-target lidar-camera matching device, A step of generating an anchor Gaussian having a fixed center position and a plurality of auxiliary Gaussians around the anchor Gaussian based on data acquired by a LiDAR sensor; A step of rendering a 2D image for each of the cameras based on the camera poses of the cameras obtained based on the 3D Gaussian including the anchor Gaussian and the plurality of auxiliary Gaussians and the LiDAR-camera extrinsic parameters; and A LiDAR-camera matching method comprising the step of calculating a loss based on a 2D image for each of the rendered cameras and a measured image obtained from the cameras, and jointly optimizing the LiDAR-camera extrinsic parameters and the 3D Gaussian based on the calculated loss. In paragraph 1, The step of generating the above is, A step of integrating the point cloud acquired by the above-mentioned LiDAR sensor; A step of voxelizing an integrated point cloud to obtain voxel centers, setting the voxel centers as anchor points which are the center positions of the anchor Gaussian, and generating the anchor Gaussian using a neural network model based on the anchor points; and A non-target LiDAR-camera matching method comprising the step of obtaining parameters of the plurality of auxiliary Gaussians using a multilayer perceptron, and generating the plurality of auxiliary Gaussians around the anchor Gaussian based on the obtained parameters of the plurality of auxiliary Gaussians and the anchor point. In paragraph 1, The above-mentioned optimization step is, A step of calculating the loss based on optical loss calculated based on the rendered 2D image and actual images acquired from the cameras, and scale normalization loss calculated based on the ratio between the maximum scale component and the minimum scale component of the 3D Gaussian; and A non-target LiDAR-camera matching method comprising the step of updating the LiDAR-camera extrinsic parameters and the parameters of the 3D Gaussians to minimize the loss. In paragraph 1, The above-mentioned optimization step is, A non-target LiDAR-camera matching method comprising the step of updating the camera attitudes of the cameras in a direction that minimizes optical loss calculated based on the rendered 2D image and the actual image obtained from the cameras. In a non-target LiDAR-camera matching device, Memory for storing a neural network model for Gaussian generation and programs and data for LiDAR-camera matching; and A non-target LiDAR-camera matching device comprising at least one processor and a control unit that operates by executing a program stored in the memory, generates an anchor Gaussian having a fixed center position and a plurality of auxiliary Gaussians around the anchor Gaussian based on data acquired by a LiDAR sensor, renders a 2D image for each of the cameras based on a 3D Gaussian including the anchor Gaussian and the plurality of auxiliary Gaussians and the camera pose of the cameras acquired based on LiDAR-camera extrinsic parameters, calculates a loss based on the rendered 2D image for each of the cameras and the actual image acquired from the cameras, and jointly optimizes the LiDAR-camera extrinsic parameters and the 3D Gaussian based on the calculated loss. In paragraph 5, The above control unit is, The point cloud acquired by the above LiDAR sensor is integrated, the voxel center obtained by voxelizing the integrated point cloud is set as the anchor point, which is the center position of the above anchor Gaussian, and the above anchor Gaussian is generated using a neural network model based on the above anchor point. A non-target LiDAR-camera matching device that obtains parameters of the plurality of auxiliary Gaussians using a multilayer perceptron, and generates the plurality of auxiliary Gaussians around the anchor Gaussian based on the obtained parameters of the plurality of auxiliary Gaussians and the anchor point. In paragraph 5, The above control unit is, The loss is calculated based on optical loss calculated based on the rendered 2D image and actual images acquired from the cameras, and scale normalization loss calculated based on the ratio between the maximum scale component and the minimum scale component of the 3D Gaussian. A non-target LiDAR-camera matching device that updates the parameters of the LiDAR-camera extrinsic parameters and the 3D Gaussians to minimize the above loss. In paragraph 5, The above control unit is, A non-target LiDAR-camera matching device that updates the camera attitudes of the cameras in a direction that minimizes optical loss calculated based on the rendered 2D image and the actual image obtained from the cameras. A computer-readable recording medium having a computer program that performs the method described in paragraph 1.