RGBD thundersight all-in-one machine based on separation optical filter

By separating the filter and directly superimposing the RGBD radar imager into a single model, the calibration process is simplified, the fusion accuracy and efficiency of the RGBD radar imager are improved, and the problems of complex process and poor fusion effect in traditional solutions are solved.

CN121999059APending Publication Date: 2026-05-08SUZHOU SAITONG INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU SAITONG INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2026-01-29
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

The existing RGBD laser radar-camera integrated machine has a complex calibration process for fusion between LiDAR and camera, high resolution dependence, and difficulty in recognizing the boundary between black and white grids, resulting in poor fusion accuracy and failing to meet the needs of practical applications.

Method used

An RGBD radar-visual integrated machine based on a separate filter is adopted. The infrared filter is installed separately from the lens. During the calibration and fusion stage, the filter is removed, and the RGB camera captures infrared light spots to obtain pixel information. During normal use, the infrared light is filtered out. The direct overlap model is used to realize the association between point cloud and RGB information, simplifying the calibration process and reducing resolution dependence.

Benefits of technology

The calibration process has been simplified, the fusion accuracy has been improved, the usage threshold has been lowered, and it has been ensured that even low-resolution LiDAR can be matched efficiently, with the fusion effect achieving an overlap of over 90% and a color accuracy of within 5.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121999059A_ABST
    Figure CN121999059A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of RGBD cameras, and particularly relates to an RGBD thundersight all-in-one machine based on a separation optical filter. The system comprises a flash laser radar, an RGB camera, a lens matched with the RGB camera and a detachable infrared optical filter, the detachable infrared optical filter and the lens are installed separately, the detachable infrared optical filter and the lens are detached in a calibration fusion stage, and the detachable infrared optical filter and the lens are installed in a normal use stage. The flash laser radar emits infrared light spots and outputs ranging data, the RGB camera shoots the infrared light spots and RGB color images in stages, fusion is completed through position correspondence by adopting a direct coincidence model, and coordinate matrix transformation is not needed. The all-in-one machine simplifies the calibration process, reduces the resolution dependence, improves the fusion accuracy, and solves the problems that a traditional scheme is complex and poor in effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of RGBD cameras, specifically relating to an RGBD radar-visual integrated camera based on a separate filter. Background Technology

[0002] In the field of RGBD laser-sensor integrated systems, the fusion calibration of LiDAR and camera is a core technical challenge. Existing fusion solutions use a black-and-white grid for pose calibration, requiring a complex process to find the transformation matrix for fusing the two. This approach has significant drawbacks: the data acquisition, calculation, and output processes are cumbersome, and calibration is difficult; for low-resolution LiDAR, identifying the boundary areas between the black and white grids is extremely difficult, resulting in poor fusion accuracy and failing to meet practical application requirements. Based on these issues, there is an urgent need for a technical solution that simplifies the calibration process, reduces resolution dependence, and improves fusion accuracy, addressing the core problems of the complexity, inefficiency, and poor fusion results of traditional solutions. Summary of the Invention

[0003] The purpose of this invention is to overcome the shortcomings of existing technologies and propose an RGBD radar-visual integrated machine based on a separate filter. This machine includes a flash laser radar, an RGB camera, a lens that works with the RGB camera, and a detachable infrared filter. The detachable infrared filter is installed separately from the lens. It is removed during the calibration and fusion phase and installed during normal use. The flash laser radar emits infrared light spots and outputs ranging data. The RGB camera captures infrared light spots to obtain pixel information during the calibration and fusion phase and captures RGB color images during normal use. The RGBD radar-visual integrated machine uses a direct overlap model, which associates point cloud and RGB information by matching the position of pixel information with the ranging data. The direct overlap model does not require coordinate matrix transformation and achieves fusion only through image recognition and data matching.

[0004] Preferably, the flash lidar has a point-to-point architecture, with each infrared spot corresponding to a detection channel. The position of the infrared spot illumination corresponds one-to-one with the position of the point cloud ranging data output by the detection channel. During the calibration and fusion stage, the RGB camera captures the infrared spot to obtain the pixel information corresponding to each detection channel on the RGB camera. During normal use, the detachable infrared filter filters out the infrared light emitted by the flash lidar, the RGB camera acquires the RGB color image, and the RGBD lidar integrated machine assigns the RGB information of the corresponding pixels in the RGB color image to the point cloud corresponding to the detection channel.

[0005] In a further preferred embodiment, the flash lidar features a uniform light architecture, where a single emitted infrared spot, after uniform light diffusion, corresponds to multiple ranging channels. During the calibration and fusion phase, an RGB camera captures an image of the infrared spot, performs binarization on the image, and calculates the spot centroid. The flash lidar is set to a grayscale output mode, which outputs black and white image data. The black and white image data is then binarized, and the centroid is calculated. The two types of centroids are matched one-to-one to determine the center ranging channel corresponding to each spot centroid. During normal use, after installing a detachable infrared filter, the RGB camera acquires RGB color images. The RGBD lidar integrated machine associates the RGB information with the ranging channels based on the centroid correspondence, and then performs uniform light processing.

[0006] In a further preferred embodiment, the filtering band of the detachable infrared filter is completely consistent with the infrared light band emitted by the flash lidar. The lens and the detachable infrared filter adopt a plug-in disassembly structure. The detachable infrared filter is installed at the front or rear end of the RGB camera lens. After installation, the optical parameters of the lens are not changed. The disassembly process does not damage the lens and the photosensitive element of the RGB camera. The detachable infrared filter uses optical glass as a substrate, and an infrared cut-off film layer is coated on the surface of the substrate.

[0007] Further preferably, the environment in which the RGB camera captures the infrared light spot is a dark environment with an ambient light intensity of no more than 10 lux, retaining only the infrared light emitted by the flash laser radar as the light source; the shooting operation steps are as follows: remove the detachable infrared filter, adjust the distance between the RGBD laser radar integrated machine and the target to 0.5 meters to 5 meters, start the flash laser radar to emit the infrared light spot, adjust the exposure parameters of the RGB camera so that the standard deviation of the gray value of the infrared light spot in the image is greater than 30, the exposure parameters include an exposure time of 1ms to 10ms and a sensitivity of ISO 100 to ISO 400, take 10 images and select the image with the largest standard deviation of gray value in the light spot area as the calibration image, and save the image data after shooting for subsequent pixel extraction or centroid calculation.

[0008] Further preferably, the threshold of 128 for binarization processing is the boundary value between black and white pixels. The threshold is determined based on the difference in gray values ​​between the infrared spot and the background in the image captured by the RGB camera. When the difference is greater than 50 gray levels, the threshold of 128 is used. When the difference is less than or equal to 50 gray levels, the threshold range is adjusted to 100 to 150. The adjustment method is to statistically analyze the average gray value of the spot area and the average gray value of the background area in the image, and calculate the median value of the two as the adjusted threshold. The adjusted threshold must meet the requirement that the proportion of white pixels in the spot area is 80% to 95%, and the proportion of white pixels in the background area does not exceed 5%.

[0009] A further preferred approach is to map the centroid of the light spot to the ranging channel as follows: Calculate the centroid coordinates of the RGB camera image and the centroid coordinates of the flash LiDAR grayscale image, and perform a one-to-one matching using the principle of minimum Euclidean distance. For each centroid coordinate in the RGB camera image, find the centroid coordinate with the smallest Euclidean distance in the set of centroid coordinates of the LiDAR grayscale image as the corresponding coordinate. If the minimum Euclidean distance is less than 5 pixels, the matching is considered successful. After successful matching, the ranging channel corresponding to the centroid coordinates of the LiDAR grayscale image is the ranging channel corresponding to the centroid coordinates of the RGB camera image. Establish a mapping table between the centroid coordinates and the ranging channel number for subsequent association process calls.

[0010] Further preferably, the resolution of the flash LiDAR grayscale image output mode is consistent with the resolution of the image captured by the RGB camera, with selectable resolutions of 320×240 pixels, 640×480 pixels, or 1280×720 pixels; the data format is an 8-bit grayscale image format, with the grayscale value of each pixel ranging from 0 to 255; the data storage format is BMP or PNG; the grayscale value of each pixel in the grayscale image is positively correlated with the laser reflection intensity at the corresponding location, with higher reflection intensity resulting in a larger grayscale value; and the frame rate of the grayscale image is consistent with the shooting frame rate of the RGB camera, ranging from 15fps to 30fps.

[0011] Further optimized, the quantitative evaluation indicators of the fusion effect include overlap and color accuracy; overlap is the proportion of spatial overlap between the point cloud and the RGB image, calculated as the percentage of pixels in the overlapping area to the total corresponding pixels in the point cloud, with an overlap value of not less than 90%; color accuracy uses the color difference evaluation of the CIE1976Lab color space, with a color difference value not exceeding 5; the test data includes overlap and color accuracy values ​​under different distances and lighting conditions, with 10 tests under each condition, and the average value is taken as the final test result. The test distances cover 0.5 meters, 1 meter, 2 meters, 3 meters, and 5 meters, and the test lighting conditions cover dark environments, weak ambient light, and indoor normal light.

[0012] A further preferred method for calibrating and compensating camera parameters after the detachable infrared filter is installed is as follows: when the detachable infrared filter is installed at the front of the lens, the installation and removal process does not change the camera's intrinsic parameters, and no additional calibration is required; when the detachable infrared filter is installed at the rear of the lens, color calibration is performed by shooting a standard color chart image after each installation and removal. The calibration steps are as follows: shoot a standard 24-color chart image, calculate the average RGB value of each color block on the color chart, compare it with the standard RGB value of the standard color chart to obtain a color deviation matrix, and use a color correction algorithm to correct the deviation of the subsequently shot RGB images. The color deviation of the camera before and after installation and removal does not exceed 3 RGB channel values.

[0013] Technical Effects: The core inventive technology of this invention is the use of a separate infrared filter and a directly coincident model, eliminating the need for coordinate matrix transformation and black-and-white grid calibration. The separate filter allows for flexible switching between calibration and usage stages, while the directly coincident model achieves fusion through image recognition and data matching. This solves the core problems of traditional solutions, such as complex processes, high resolution dependence, and poor fusion effects, simplifying the calibration process, lowering the usage threshold, and improving fusion accuracy. Attached Figure Description

[0014] Figure 1 This is a block diagram showing the core modules and working stages of the RGBD all-in-one camera in this application; Figure 2 This is a block diagram showing the connection between the processing module, calibration compensation, and fusion evaluation in this application. Figure 3 This is the output diagram of the lidar in this application. Detailed Implementation

[0015] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0016] Traditional technical solutions rely on black and white grid boards for fusion calibration, which is a complex process and makes it difficult for low-resolution LiDAR to identify the boundaries between black and white grids, resulting in poor fusion performance.

[0017] Based on this, please refer to Figures 1-3 This embodiment provides an RGBD laser-guided all-in-one camera based on a separate filter, comprising a flash laser radar, an RGB camera, a lens that works with the RGB camera, and a detachable infrared filter. The detachable infrared filter is separately mounted from the lens, and the two are stably assembled and disassembled through a compatible mechanical structure. During the calibration and fusion phase, the detachable infrared filter needs to be removed from the lens. At this time, the photosensitive element of the RGB camera can directly receive the infrared light spot emitted by the flash laser radar. After entering the normal use phase, the detachable infrared filter is installed back into the corresponding mounting position on the lens. Its filtering band perfectly matches the infrared light band emitted by the flash laser radar, which can block infrared light from entering the photosensitive element of the RGB camera and avoid interference with RGB color image acquisition.

[0018] After the flash lidar is activated, it continuously emits infrared light spots and outputs ranging data corresponding to the position of the light spot. This ranging data directly reflects the distance information of the light spot's illumination point. The core task of the RGB camera in the calibration and fusion stage is to capture the infrared light spots emitted by the flash lidar. The image sensor records the position information of the light spots on the camera's imaging plane, and then extracts the pixel information corresponding to each light spot. This pixel information will serve as the basis for subsequent data matching. During normal use, since the detachable infrared filter is already installed, the RGB camera only receives visible light from the environment, thus acquiring clear RGB color images.

[0019] This all-in-one machine employs a direct overlap model for fusion. Its core logic involves matching the pixel information acquired during the calibration fusion phase with the ranging data output by the flash LiDAR. Since the position of the infrared spot emitted by the flash LiDAR and the spatial position corresponding to the ranging data are fixedly correlated, and the pixel information of the spot captured by the RGB camera directly corresponds to the spatial coordinates of the ranging data, there is no need for complex coordinate matrix transformations to convert data between different coordinate systems. Only simple image recognition technology is needed to confirm the pixel position corresponding to the spot, and then data matching technology is used to associate the color information in the RGB color image corresponding to that pixel position with the ranging data to complete the fusion of point cloud and RGB information. The core of this design lies in fully utilizing the spot emission characteristics of the flash LiDAR and the imaging characteristics of the RGB camera to establish a direct correspondence between pixel information and ranging data. This eliminates the cumbersome coordinate transformation process of traditional solutions, simplifying the calibration operation from the root, and reducing the dependence on LiDAR resolution. Even low-resolution LiDARs can complete the matching by recognizing their own emitted spot, ensuring the accuracy of the fusion.

[0020] In traditional technical solutions, when point-to-point LiDAR and camera are fused, there is a lack of simple and efficient methods for matching pixels with detection channels.

[0021] Based on this, the flash lidar adopts a point-to-point architecture. The core design of this architecture is that each emitted infrared spot uniquely corresponds to a detection channel. The specific position of the infrared spot on the target object is completely consistent with the spatial position pointed to by the point cloud ranging data output by the detection channel. This one-to-one correspondence provides the basis for subsequent matching.

[0022] During the calibration and fusion phase, the detachable infrared filter is first removed to ensure that the RGB camera can accurately receive the infrared light spots emitted by the flash LiDAR. Then, the RGB camera is activated to capture images containing these infrared light spots. Through image acquisition and preliminary processing, the pixel information corresponding to each infrared light spot on the RGB camera's imaging plane is extracted. This pixel information forms a one-to-one mapping relationship with the flash LiDAR's detection channels, meaning that each detection channel can have its specific pixel position on the RGB camera determined through calibration.

[0023] Once in normal use, the detachable infrared filter is reinstalled. Its filtering characteristics precisely filter out the infrared light emitted by the flash lidar, preventing infrared light from interfering with the RGB camera's acquisition of ambient visible light. At this point, the RGB camera can capture clear RGB color images. The RGBD lidar integrated machine utilizes the mapping relationship between the detection channel and pixel position established during the calibration and fusion phase, accurately assigning the RGB information of the corresponding pixel position in the RGB color image to the point cloud data output by the corresponding detection channel.

[0024] The image recognition process employs binarization, specifically implemented as follows: First, the image containing the infrared spot, captured during the calibration and fusion stage, is read. Since this image may be color, it must be converted to grayscale. In the grayscale image, each pixel's grayscale value ranges from 0 to 255, where 0 represents pure black and 255 represents pure white. A fixed threshold of 128 is set as the boundary between black and white pixels. When the grayscale value of a pixel in the grayscale image is greater than 128, that pixel is determined to be white, and the area formed by these white pixels represents the location of the infrared spot in the image. Pixels with grayscale values ​​less than or equal to 128 are determined to be black pixels, corresponding to the image background area. This binarization method is simple and efficient, quickly and accurately separating the spot area from the image, providing clear recognition results for matching the detection channel and pixel information. It eliminates the need for complex image segmentation algorithms, significantly simplifying the calculation logic of the matching process.

[0025] In traditional technical solutions, the threshold for binarization is fixed, which cannot be adapted to different cameras and scenes, resulting in low accuracy of spot recognition.

[0026] Based on this, the threshold of 128 for binarization processing is the dividing line between black and white pixels. This threshold is set according to the grayscale difference between the infrared spot and the background in typical scenes, effectively distinguishing the spot from the background in most cases. However, in practical applications, differences in the photosensitivity and image quality of different cameras, as well as variations in the degree of ambient light interference in the shooting scene, can cause changes in the grayscale difference between the infrared spot and the background. Therefore, the threshold needs to be dynamically adjusted to ensure the accuracy of spot recognition.

[0027] When the difference in grayscale values ​​between the infrared spot and the background in an image captured by an RGB camera is greater than 50 grayscale levels, it indicates that the spot and the background have high distinguishability, and an accurate identification can be achieved using a threshold of 128. When the difference in grayscale values ​​is less than or equal to 50 grayscale levels, the distinguishability between the spot and the background decreases, and a fixed threshold of 128 is prone to incomplete identification of the spot area or misidentification of background pixels as the spot. In this case, the threshold adjustment range is set to 100 to 150. This range is determined based on the grayscale value distribution characteristics and can cover the optimal threshold range for most low-discrimination scenarios.

[0028] The threshold adjustment method involves statistically analyzing the average gray values ​​of the light spot region and the background region in the image, and calculating the median value as the adjusted threshold. The core formula is that the adjusted threshold equals half the sum of the average gray values ​​of the light spot and the background. This formula is logically derived based on the fundamental principle of grayscale image binarization segmentation. The core objective of binarization is to find an optimal threshold that maximizes the inter-class difference and minimizes the intra-class difference between the segmented foreground and background regions. Theoretically, the average gray value of the light spot region represents the gray level of the foreground, and the average gray value of the background region represents the gray level of the background. The median value lies precisely at the boundary between the foreground and background grayscale distributions. Using this as the threshold maximizes the distinction between foreground and background, conforming to the optimal criterion for binarization segmentation.

[0029] In practical implementation, preliminary image analysis is first used to determine the approximate ranges of the spot and background regions. Then, the grayscale values ​​of all pixels within these two regions are statistically analyzed, and their respective arithmetic mean values ​​are calculated: the average grayscale value of the spot and the average grayscale value of the background. These two average values ​​are then added together and divided by 2 to obtain the adjusted threshold. The innovation of this adjustment method lies in abandoning the limitations of a fixed threshold. By dynamically statistically analyzing the grayscale characteristics of the spot and background in the image, the optimal threshold is determined. This allows for adaptation to different cameras and scenes without manual intervention, ensuring that the proportion of white pixels in the spot region remains between 80% and 95%, and the proportion of white pixels in the background region does not exceed 5%. This guarantees the accuracy and stability of spot recognition and provides a reliable foundation for subsequent data matching.

[0030] In traditional technical solutions, the mapping logic between the centroid of the light spot and the ranging channel in the uniform light architecture is unclear, resulting in low matching accuracy.

[0031] Based on this, the mapping logic between the spot centroid and the ranging channel is to calculate the centroid coordinates of the RGB camera image and the grayscale image of the flash LiDAR, and perform one-to-one matching using the principle of minimum Euclidean distance. When the minimum Euclidean distance is less than 5 pixels, the matching is considered successful. A mapping table between the centroid coordinates and the ranging channel number is established for subsequent association process calls.

[0032] The formula for calculating Euclidean distance is: The logical derivation of this formula is based on the geometric definition of the distance between points in two-dimensional space. In RGB camera images and flash LiDAR grayscale images, the centroid exists in the form of two-dimensional coordinates. The straight-line distance between two points in two-dimensional space is the most direct and accurate indicator of the correlation between their positions. Euclidean distance is the classic calculation method for describing the straight-line distance between two points in two-dimensional space, and its theoretical basis comes from the Pythagorean theorem in plane geometry. Applying the Pythagorean theorem to the two-dimensional coordinate system, for any two points, the square root of the sum of the squares of the differences in their x-coordinates and y-coordinates is the straight-line distance between the two points. This calculation method can accurately quantify the positional deviation between the coordinates of two centroids, providing an objective numerical basis for matching judgment.

[0033] In the formula, and Represents the coordinates of the centroid of the light spot in an image captured by an RGB camera. and This represents the coordinates of the centroid of the light spot in the grayscale image of the flash LiDAR. All coordinates are in pixels, and pixels, as the basic unit of an image, accurately reflect the position of the centroid on the imaging plane. During the calculation, the centroid coordinates of all light spots in both types of images are first obtained, forming two coordinate sets. Then, for each centroid coordinate in the RGB camera image coordinate set, it is successively substituted with each centroid coordinate in the LiDAR grayscale image coordinate set into the formula to calculate the Euclidean distance. Each calculation result is recorded, and finally, the smallest Euclidean distance value is selected.

[0034] The formula is implemented by extracting the centroid coordinates through image processing algorithms and then using programming logic to perform distance calculation and minimum value selection. Its core innovation lies in combining the classic method for calculating the distance between two points in two-dimensional space with the fusion requirements of a uniform-light architecture LiDAR, solving the technical challenge of accurately corresponding the light spot with multiple ranging channels after uniform light diffusion. Traditional solutions cannot clearly define the correspondence between the light spot centroid and the ranging channel, while this formula, by quantifying the positional deviation between centroid coordinates, can objectively determine the correspondence between the two types of centroids. This ensures that the light spot centroid in each RGB camera image can find a unique corresponding LiDAR grayscale image centroid, thereby determining the corresponding center ranging channel. This provides a precise and reliable matching basis for the subsequent association of RGB information and ranging channels, ensuring the accuracy of fusion under a uniform-light architecture.

[0035] Traditional technical solutions lack clear quantitative evaluation standards for fusion effects, making it impossible to objectively judge the fusion quality.

[0036] Based on this, the quantitative evaluation indicators of the fusion effect include overlap and color accuracy. The overlap is the percentage of the number of pixels in the overlapping area to the total number of corresponding pixels in the point cloud, and the value is not less than 90%. The color accuracy adopts the color difference evaluation of the CIE1976Lab color space, and the value does not exceed 5. The test data covers different distances and lighting conditions, and each condition is tested 10 times and the average value is taken.

[0037] The formula for calculating color difference for color accuracy is: ; The logical derivation of this formula is based on the characteristics of the CIE 1976Lab color space. The CIE 1976Lab color space is a uniform color space designed to objectively reflect the human eye's perception of color differences by specifying the distance between two points. This characteristic makes it an ideal choice for evaluating color accuracy. Compared to other color spaces, the color differences corresponding to the same distance in the CIE 1976Lab color space are consistent in human visual perception, avoiding the mismatch between distance and visual perception found in non-uniform color spaces. This provides a scientific theoretical basis for quantifying color differences.

[0038] In the formula, , , The Lab value represents the color of the merged point cloud. , , The Lab value represents the actual target color. The L channel represents the brightness of the color, ranging from 0 to 100, where 0 corresponds to pure black and 100 to pure white; a higher value indicates a brighter color. The a channel represents the red-green hue of the color; positive values ​​lean towards red, and negative values ​​lean towards green; a higher absolute value indicates a more pronounced red-green hue. The b channel represents the blue-yellow hue of the color; positive values ​​lean towards yellow, and negative values ​​lean towards blue; a higher absolute value indicates a more pronounced blue-yellow hue.

[0039] The formula calculates the difference between the blended color and the actual target color in the L, a, and b channels. Then, it sums the squares of each channel's difference, and finally takes the square root of the sum. The result is the color difference value. The magnitude of the color difference value directly reflects the degree of deviation between the blended color and the actual target color. The smaller the value, the smaller the color difference and the higher the color accuracy; conversely, the larger the value, the greater the color difference and the lower the color accuracy.

[0040] The formula is implemented by acquiring the Lab values ​​of the fused point cloud color and the actual target color using professional color acquisition equipment, and then substituting these values ​​into the formula for calculation. Its core innovation lies in abandoning the traditional method of evaluation through visual observation, and adopting a quantitative calculation method based on a scientific color space to achieve an objective and precise evaluation of color accuracy. In traditional solutions, the evaluation of the fusion effect relies on manual visual observation, which is highly subjective and makes it difficult to establish a unified evaluation standard. This formula, by quantifying color differences, provides a verifiable and comparable evaluation indicator for the fusion effect, ensuring the objectivity and consistency of the fusion quality judgment, and also providing a clear direction for optimizing the technical solution.

[0041] In traditional technical solutions, the removal and installation of filters may cause changes in camera parameters, affecting the fusion accuracy.

[0042] Based on this, when the detachable infrared filter is installed at the front of the lens, the disassembly and assembly process will not change the camera's intrinsic parameters, and no additional calibration is required; when it is installed at the rear of the lens, a standard 24-color chart image needs to be taken after each disassembly and assembly, the RGB average value of each color block on the color chart is calculated, and compared with the standard RGB value of the standard color chart to obtain a color deviation matrix. The deviation is then corrected for the RGB images taken subsequently by the color correction algorithm. The color deviation of the camera before and after disassembly and assembly does not exceed 3 RGB channel values.

[0043] The core formula for color correction is that the corrected RGB value equals the original RGB value multiplied by the inverse of the deviation matrix, i.e. ,in Represents the color deviation matrix. This represents the inverse matrix of the deviation matrix. The logical derivation of this formula is based on the theory of linear color transformation. When a detachable infrared filter is installed at the rear of the lens, the installation and removal process may cause a slight shift in the relative position between the filter and the camera's image sensor, or slightly affect the optical properties of the filter itself, thus causing a linear deviation in the color data acquired by the RGB camera. This linear deviation can be corrected through matrix transformation.

[0044] From a theoretical perspective, the color deviation matrix... It is a 3×3 matrix, where each element represents the deviation coefficient between different RGB channels. The element values ​​in the matrix are calculated by comparing the actual collected values ​​of the standard color chart with the standard values. Assume the RGB standard value of a certain color patch in the standard color chart is... The camera captures the original RGB values ​​of this color patch. Then a linear relationship exists. This formula indicates that the original acquired values ​​are the result of standard values ​​transformed by a deviation matrix. To recover accurate color data from the original acquired values, an inverse operation needs to be performed on this linear relationship, i.e. This is the core derivation logic of the color correction formula.

[0045] In the formula, This represents the uncorrected RGB color data obtained from the camera, with each channel having a value ranging from 0 to 255. It is the inverse matrix of the deviation matrix obtained through matrix inversion. This represents the accurate RGB color data obtained after correction. The formula is implemented by first obtaining the deviation matrix through calibration using a standard color chart, then calculating the inverse matrix using the matrix inversion method in linear algebra, and finally performing matrix multiplication between the original RGB data obtained from each shot and the inverse matrix to obtain the corrected RGB data.

[0046] The core innovation of this formula lies in its use of matrix transformation to achieve precise correction of color deviation. It establishes a scientific correction model to address the issue of misalignment during filter installation at the rear of the lens. Traditional solutions cannot effectively solve color deviation caused by installation and removal. This formula, however, quantifies the deviation into a matrix form through linear transformation theory, and then achieves reverse compensation of the deviation through inverse matrix operations. This ensures the consistency of color data acquired by the camera before and after installation and removal, guaranteeing the stability of fusion accuracy. Furthermore, the matrix transformation method is computationally efficient and provides accurate correction, making it suitable for real-time processing requirements in practical applications.

[0047] In traditional technical solutions, the light spot of a uniform-beam lidar is difficult to accurately match with multiple ranging channels after it is diffused.

[0048] Based on this, the flash lidar adopts a uniform light architecture. After the single emitted infrared light spot is processed by uniform light diffusion, it will cover multiple ranging channels. This architecture design can expand the detection range of lidar, but it also brings the challenge of accurately matching the light spot with the ranging channel.

[0049] During the calibration and fusion phase, the detachable infrared filter is first removed to ensure that the RGB camera can clearly receive the infrared light spots emitted by the flash LiDAR after uniform diffusion. Then, the RGB camera is activated to capture images containing these light spots. Since the light spots have a large area after uniform diffusion and correspond to multiple ranging channels, they cannot be directly matched with the ranging channels by pixel position. Therefore, the captured images need to be binarized. By setting an appropriate threshold, the light spot area and background area are separated. Then, the centroid of each light spot area is calculated to obtain the geometric center coordinates of each diffused light spot. These coordinates represent the core position of the light spot, providing a precise positional reference for subsequent matching.

[0050] Simultaneously, the flash LiDAR is set to grayscale output mode. In this mode, the LiDAR outputs black and white image data synchronized with the images captured by the RGB camera. This black and white image data reflects the distribution of infrared spots on the LiDAR detection plane. The same binarization and centroid calculation steps as those for the RGB camera images are performed on this black and white image data to obtain the centroid coordinates of the spots at the LiDAR end. Since the two types of images are captured synchronously and both reflect the same infrared spot distribution, there is a one-to-one correspondence between the two types of centroid coordinates. This correspondence allows us to determine the center ranging channel at the LiDAR end corresponding to the centroid of the spot in each RGB camera image. That is, among the multiple ranging channels covered by each diffused spot, the ranging channel corresponding to the centroid at the LiDAR end is used as the core channel, establishing a mapping relationship between the centroid of the RGB camera image and the core ranging channel.

[0051] Upon entering normal operation, the detachable infrared filter is reinstalled in the lens position to filter out the infrared light emitted by the flash lidar. The RGB camera then begins acquiring RGB color images of the environment. The RGBD lidar integrated machine invokes the mapping relationship between the centroid and the core ranging channel established during the calibration and fusion phase. It associates the RGB information at the corresponding centroid position in the RGB color image with the point cloud data corresponding to that core ranging channel and the multiple ranging channels it covers. After association, uniform light processing is performed on the associated point cloud data to ensure that the distribution of the point cloud data conforms to the detection characteristics of the uniform light architecture, ultimately achieving accurate fusion of point cloud and RGB information under the uniform light architecture.

[0052] The core of this scheme lies in establishing the correspondence between the two types of images through centroid calculation. By utilizing the position reference provided by the grayscale output mode of the LiDAR, the matching problem between the light spot and multiple ranging channels after uniform light diffusion is solved. At the same time, through a phased calibration and fusion process, the accuracy of matching and the stability of fusion effect are ensured.

[0053] In traditional technical solutions, filters and lenses are integrated and installed, which cannot meet the different optical requirements of calibration and use stages, and disassembly and assembly may affect equipment performance.

[0054] Based on this, the detachable infrared filter has a precisely designed filtering band that is completely consistent with the infrared light band emitted by the flash lidar, which can achieve efficient filtering of infrared light and prevent infrared light from interfering with the color acquisition of the RGB camera during normal use.

[0055] The lens and detachable infrared filter adopt a plug-in disassembly structure. This structure design is based on the principles of convenience and stability of mechanical assembly. The assembly parts of the filter and lens are equipped with matching slots and positioning structures to ensure that the filter can be stably fixed after insertion and will not be displaced due to factors such as equipment vibration. At the same time, the plug-in design makes the filter disassembly and assembly operations simple and quick without the need for complicated tools, thus improving the efficiency of switching between calibration and use stages.

[0056] The detachable infrared filter can be installed at either the front or rear of the lens. When installed at the front of the lens, the filter is separated from the camera's image sensor by the lens's optical structure. During installation and removal, the filter will not directly contact the image sensor, nor will it change the lens's optical parameters. Therefore, no additional calibration steps are required to ensure the camera's image quality. When installed at the rear of the lens, the filter is close to the camera's image sensor. Although the installation and removal process will not damage the equipment, subsequent color calibration steps are needed to compensate for any minor deviations that may occur in order to ensure the accuracy of color acquisition.

[0057] The filter uses optical glass as its substrate. Optical glass possesses excellent light transmittance and optical stability, ensuring the smooth passage of visible light during normal use while providing a stable support for the filter layer. An infrared cutoff film is deposited on the substrate surface. This film is prepared using optical coating technology. By controlling the material, thickness, and number of layers, it achieves reflection and absorption of specific infrared wavelengths, thereby filtering out infrared light. The film's design is based on the principle of optical interference, causing destructive interference of the infrared light emitted by the lidar at the film surface. Most of the infrared light is reflected or absorbed, while visible light can pass through the film smoothly into the camera's image sensor, ensuring the quality of RGB color image acquisition.

[0058] The core of this design lies in clearly defining the filter parameters, disassembly and assembly structure, and installation position of the filter. This balances the infrared light reception requirements during the calibration phase with the infrared light filtering requirements during normal use, while ensuring the stability and lifespan of the equipment. It solves the problem that traditional integrated filters cannot flexibly adapt to the optical needs of different stages.

[0059] In traditional technical solutions, the camera is greatly affected by environmental interference when capturing light spots, and the lack of clear shooting parameters and operating procedures leads to poor light spot acquisition results.

[0060] Therefore, the environment for RGB cameras to capture infrared light spots is strictly limited to a dark environment with an ambient light intensity not exceeding 10 lux. This light intensity standard is determined based on the characteristics of the infrared light spot and the camera's photosensitivity, minimizing interference from ambient visible light and ensuring that the RGB camera can accurately capture the outline and position information of the infrared light spot. Only the infrared light emitted by the flash lidar is retained as the sole light source in the shooting environment, further avoiding interference from other light sources.

[0061] The shooting operation process is clearly designed according to a logical sequence. First, the detachable infrared filter is removed to ensure that infrared light can smoothly enter the photosensitive element of the RGB camera. Then, the distance between the RGBD laser radar and the target is adjusted to 0.5 meters to 5 meters. This distance range is determined by comprehensively considering the emission power of the flash laser radar, the diffusion range of the infrared spot, and the imaging resolution of the RGB camera. Within this distance range, the infrared spot can form a clear image on the target surface, and the RGB camera can capture a sufficiently large spot image to facilitate subsequent pixel extraction and centroid calculation. After the flash LiDAR stably emits an infrared light spot, the exposure parameters of the RGB camera are adjusted. The exposure time is set to 1ms to 10ms, and the sensitivity is set to ISO 100 to ISO 400. This parameter range is optimized based on the light intensity of the infrared light spot and the light sensitivity characteristics of the camera. It can ensure that the light spot image will not lose details due to underexposure, nor will it be overexposed and blended due to overexposure. This ensures that the infrared light spot meets the requirement that the standard deviation of the gray value in the image is greater than 30. This standard deviation can reflect the distinction between the light spot and the background. The larger the value, the clearer the light spot.

[0062] Ten images are captured consecutively during the shooting process. The image with the largest standard deviation of gray value in the spot area is selected as the calibration image through image quality analysis. This design can reduce the random errors that may occur in a single shooting, ensure that the selected calibration image has the highest clarity and discrimination, provide high-quality data support for subsequent pixel extraction, centroid calculation and data matching, and ensure the accuracy of calibration fusion.

[0063] In traditional technical solutions, the parameters of the grayscale image output mode of LiDAR are not clear, resulting in a mismatch between the timing and resolution of the image and the camera image.

[0064] Based on this, the resolution of the grayscale image output mode of the flash LiDAR is consistent with the resolution of the image captured by the RGB camera, providing three selectable resolutions: 320×240 pixels, 640×480 pixels, or 1280×720 pixels. These resolutions are commonly used by mainstream RGB cameras and can meet the application scenarios with different precision requirements, ensuring that the LiDAR grayscale image and the RGB camera image have good compatibility in the pixel dimension, which is convenient for subsequent centroid calculation and coordinate matching.

[0065] The LiDAR grayscale image data format adopts an 8-bit grayscale image format, with each pixel's grayscale value ranging from 0 to 255, where 0 corresponds to pure black and 255 corresponds to pure white. This data format conforms to mainstream image processing standards, facilitating subsequent binarization processing and centroid calculation. The grayscale image is stored in either BMP or PNG format, both of which are lossless image storage formats that can completely preserve the pixel information of the grayscale image, avoiding data loss or distortion during storage and ensuring image quality.

[0066] The grayscale value of each pixel in the grayscale image is positively correlated with the laser reflection intensity at the corresponding location. The higher the reflection intensity, the larger the grayscale value. This design is based on the detection principle of lidar. After the infrared light emitted by lidar illuminates the target surface, the intensity of the reflected light is related to factors such as the material, roughness, and distance of the target surface. By converting the reflection intensity into grayscale values, the reflection characteristics of the target surface can be intuitively reflected in the grayscale image, providing a more accurate position reference for centroid calculation.

[0067] The frame rate of the LiDAR grayscale image is set to 15fps to 30fps, consistent with the frame rate of the RGB camera. The frame rate unit is frames per second. This frame rate range ensures temporal synchronization between the LiDAR grayscale image and the RGB camera image, guaranteeing that both images capture the infrared spot distribution at the same moment and avoiding deviations in centroid coordinate matching due to temporal differences. Temporal synchronization is a crucial prerequisite for accurate fusion. This design ensures consistency between the two types of images in the temporal dimension through a unified frame rate, providing a reliable temporal foundation for subsequent centroid matching and data association.

[0068] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. An RGBD laser-guided all-in-one machine based on a separate filter, comprising a flash laser radar, an RGB camera, and a lens that works with the RGB camera, characterized in that, It also includes a detachable infrared filter, which is installed separately from the lens. The detachable infrared filter is removed during the calibration and fusion stage and installed during normal use. The flash lidar is used to emit infrared light spots and output ranging data. The RGB camera captures infrared light spots to obtain pixel information during the calibration and fusion phase, and captures RGB color images during normal use. The RGBD radar-visual integrated machine adopts a direct overlap model, which completes the association between point cloud and RGB information by matching the position of pixel information with ranging data. The direct overlap model does not require coordinate matrix transformation, and achieves fusion only through image recognition and data matching.

2. The RGBD radar-visual integrated machine based on a separate filter according to claim 1, characterized in that, The flash lidar has a point-to-point architecture, with each infrared spot corresponding to a detection channel. The position of the infrared spot illumination corresponds one-to-one with the position of the point cloud ranging data output by the detection channel. In the calibration and fusion stage, the RGB camera captures the infrared spot to obtain the pixel information corresponding to each detection channel on the RGB camera. During normal use, the detachable infrared filter filters out the infrared light emitted by the flash lidar, the RGB camera acquires RGB color images, and the RGBD lidar integrated machine assigns the RGB information of the corresponding pixels in the RGB color images to the point cloud corresponding to the detection channel.

3. The RGBD radar-visual integrated machine based on a separate filter according to claim 1, characterized in that, The flash lidar has a uniform light architecture. After a single emitted infrared spot is diffused by uniform light, it corresponds to multiple ranging channels. In the calibration and fusion stage, the RGB camera captures the infrared spot to obtain an image. After the image is binarized, the centroid of the spot is calculated. The flash lidar is set to grayscale output mode. In grayscale output mode, black and white image data is output. The black and white image data is binarized and the centroid is calculated. The two types of centroids correspond one-to-one to determine the center ranging channel corresponding to the centroid of each spot. During normal use, after installing the detachable infrared filter, the RGB camera acquires RGB color images. The RGBD radar-vision integrated machine associates the RGB information with the ranging channel based on the centroid correspondence, and then performs uniform light processing.

4. The RGBD radar-visual integrated machine based on a separated filter according to claim 1, characterized in that, The detachable infrared filter has the same filtering band as the infrared light emitted by the flash lidar. The lens and the detachable infrared filter adopt a plug-in disassembly structure. The detachable infrared filter can be installed at the front or rear of the RGB camera lens. After installation, it does not change the optical parameters of the lens. The disassembly process does not damage the lens and the photosensitive element of the RGB camera. The detachable infrared filter uses optical glass as a substrate, and the surface of the substrate is coated with an infrared cut-off film layer.

5. The RGBD radar-visual integrated machine based on a separate filter according to claim 1, characterized in that, The RGB camera captures infrared light spots in a dark environment with an ambient light intensity not exceeding 10 lux, retaining only the infrared light emitted by the flash laser radar as the light source. The shooting operation steps are as follows: remove the detachable infrared filter, adjust the distance between the RGBD laser radar and the target to 0.5 meters to 5 meters, activate the flash laser radar to emit infrared light spots, adjust the exposure parameters of the RGB camera so that the standard deviation of the grayscale value of the infrared light spot in the image is greater than 30, the exposure parameters include an exposure time of 1ms to 10ms and an ISO sensitivity of 100 to 400, take 10 images and select the image with the largest standard deviation of grayscale value in the light spot area as the calibration image, and save the image data after shooting for subsequent pixel extraction or centroid calculation.

6. The RGBD radar-visual integrated machine based on a separated filter according to claim 2 or 3, characterized in that, The threshold of 128 for binarization is the dividing value between black and white pixels. The threshold is determined based on the difference in gray values ​​between the infrared spot and the background in the image captured by the RGB camera. When the difference is greater than 50 gray levels, the threshold of 128 is used. When the difference is less than or equal to 50 gray levels, the threshold range is adjusted to 100 to 150. The adjustment method is to statistically analyze the average gray value of the spot area and the average gray value of the background area in the image, and calculate the median value of the two as the adjusted threshold. The adjusted threshold must meet the requirement that the proportion of white pixels in the spot area is 80% to 95%, and the proportion of white pixels in the background area does not exceed 5%.

7. The RGBD radar-visual integrated machine based on a separate filter according to claim 3, characterized in that, The mapping logic between the centroid of the light spot and the ranging channel is as follows: Calculate the centroid coordinates of the RGB camera image and the centroid coordinates of the flash LiDAR grayscale image, and perform a one-to-one matching using the principle of minimum Euclidean distance. For each centroid coordinate in the RGB camera image, find the centroid coordinate with the smallest Euclidean distance in the set of centroid coordinates of the LiDAR grayscale image as the corresponding coordinate. When the minimum Euclidean distance is less than 5 pixels, the matching is considered successful. After successful matching, the ranging channel corresponding to the centroid coordinates of the LiDAR grayscale image is the ranging channel corresponding to the centroid coordinates of the RGB camera image. Establish a mapping table between the centroid coordinates and the ranging channel number for subsequent association process calls.

8. The RGBD radar-visual integrated machine based on a separated filter according to claim 3, characterized in that, The resolution of the grayscale image output mode of the flash LiDAR is consistent with the resolution of the image captured by the RGB camera. The resolution can be selected as 320×240 pixels, 640×480 pixels, or 1280×720 pixels. The data format is 8-bit grayscale image format, with the grayscale value of each pixel ranging from 0 to 255. The data storage format is BMP or PNG format. The grayscale value of each pixel in the grayscale image is positively correlated with the laser reflection intensity at the corresponding position. The higher the reflection intensity, the larger the grayscale value. The frame rate of the grayscale image is consistent with the shooting frame rate of the RGB camera, which is 15fps to 30fps.

9. The RGBD radar-visual integrated machine based on a separate filter according to claim 1, characterized in that, The quantitative evaluation indicators of the fusion effect include overlap and color accuracy. Overlap is the proportion of spatial overlap between the point cloud and the RGB image. It is calculated as the percentage of pixels in the overlapping area to the total number of corresponding pixels in the point cloud. The overlap value should not be less than 90%. Color accuracy is evaluated using the color difference of the CIE1976Lab color space. The color difference value should not exceed 5. The test data includes overlap and color accuracy values ​​under different distances and lighting conditions. Each condition is tested 10 times, and the average value is taken as the final test result. The test distances cover 0.5 meters, 1 meter, 2 meters, 3 meters, and 5 meters. The test lighting conditions cover dark environment, weak ambient light, and indoor normal light.

10. The RGBD radar-visual integrated machine based on a separate filter according to claim 1, characterized in that, The calibration and compensation measures for camera parameters after the detachable infrared filter is installed are as follows: When the detachable infrared filter is installed at the front of the lens, the installation and removal process does not change the camera's intrinsic parameters and no additional calibration is required; when the detachable infrared filter is installed at the rear of the lens, color calibration is performed by shooting a standard color chart image after each installation and removal. The calibration steps are as follows: shoot a standard 24-color chart image, calculate the average RGB value of each color block on the color chart, compare it with the standard RGB value of the standard color chart to obtain a color deviation matrix, and use a color correction algorithm to correct the deviation of the subsequently shot RGB images. The color deviation of the camera before and after installation and removal does not exceed 3 RGB channel values.