Color image and depth image fusion processing method and computer equipment

By using a fusion processing method of color and depth images and registering grayscale images from image sensors and LiDAR sensors, accurate color depth images are generated, solving the problems of inaccurate image registration and color misalignment in existing technologies and improving the accuracy of environmental perception.

CN121746445APending Publication Date: 2026-03-27CHENGDU FUSHI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing color and depth image fusion methods suffer from inaccurate image registration and color misalignment in complex environments, especially when there are significant differences between the data from image sensors and LiDAR sensors.

Method used

By constructing a fusion processing method for color and depth images, including image registration and color assignment steps, grayscale images acquired by image sensors and LiDAR sensors are registered to generate a color depth image, and the transformed color image information is used to assign color to the depth image.

Benefits of technology

It improves the accuracy of image registration, reduces color misalignment in color depth images, and enhances the precision of environmental perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746445A_ABST
    Figure CN121746445A_ABST
Patent Text Reader

Abstract

The invention discloses a color image and depth image fusion processing method and computer equipment, and the method comprises the steps: obtaining a first two-dimensional color image and a first two-dimensional grayscale image of a current field of view from an image sensor, and obtaining a second two-dimensional grayscale image of the current field of view according to the output of a laser radar sensor; acquiring a depth image and a second two-dimensional grayscale image of the current view field; carrying out registration on the first two-dimensional gray scale image and the second two-dimensional gray scale image so as to obtain registration information; according to the registration information, converting the first two-dimensional color image to obtain a second two-dimensional color image; and according to the second two-dimensional color image, endowing color information to corresponding pixels on the depth image so as to generate a color depth image. By implementing the technical scheme of the invention, more accurate registration information can be obtained, and color dislocation of the generated color depth image is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of software processing, and in particular to a color image and depth image fusion processing method and computer equipment. BACKGROUND

[0002] In the field of automatic driving, robots, etc., due to the need to navigate or operate in a complex environment, high-precision environmental perception is relied on. Currently, the depth data of a laser radar and the color data sensed by a camera need to be fused to improve the acquired environmental information.

[0003] The existing fusion method relies on one-time projection registration or single two-dimensional geometric transformation, for example, using a calibration external parameter or a single two-dimensional transformation to map a color image to a depth coordinate system, or projecting a depth to a color image coordinate system, to complete the coloring of a depth image. This kind of method can only achieve certain effect under simple conditions such as static, synchronization, approximate rigidity, etc., but due to the differences in illumination, reflectivity, texture distribution between the color image from the image sensor and the depth image from the laser radar sensor, i.e., the image property difference between the two is large, the feature visibility is inconsistent, causing sparse features or unstable matching, which easily leads to inaccurate image registration and color misplacement. SUMMARY

[0004] The technical problem to be solved by the present application is to provide a color image and depth image fusion processing method and computer equipment to solve the technical problems of inaccurate image registration and color misplacement in the prior art.

[0005] The technical solution adopted by the present application to solve the technical problem is: a color image and depth image fusion processing method is constructed, comprising: Step S1, acquiring a first two-dimensional color image and a first two-dimensional gray image of a current field of view from an image sensor, and acquiring a depth image and a second two-dimensional gray image of the current field of view according to the output of a laser radar sensor; Step S2, registering the first two-dimensional gray image and the second two-dimensional gray image to obtain registration information; Step S3, transforming the first two-dimensional color image according to the registration information to obtain a second two-dimensional color image; Step S4, assigning color information to corresponding pixels on the depth image according to the second two-dimensional color image to generate a color depth image.

[0006] Optionally, the laser radar sensor comprises a transmitting module for transmitting a sensing light signal to the current field of view, and a receiving module for receiving a return echo signal from the current field of view; In step S1, acquiring the second two-dimensional grayscale image of the current field of view includes: The echo signal is detected and a corresponding light-sensing signal is output. The second two-dimensional grayscale image of the current field of view is obtained based on the number of light-sensing signals counted within a preset statistical time period.

[0007] Optionally, between step S1 and step S2, the method further includes: The first two-dimensional color image, the depth image, the first two-dimensional grayscale image, and the second two-dimensional grayscale image are preprocessed, wherein the preprocessing includes at least one of the following: resolution adjustment processing, field of view adjustment processing, distortion correction processing, affine transformation processing, and timestamp synchronization processing.

[0008] Optionally, the registration information includes: a global homography transformation matrix and a local deformation field; step S2 includes: Step S21: Perform image matching on the first two-dimensional grayscale image and the second two-dimensional grayscale image to obtain the global homography transformation matrix and the set of interior point pairs; Step S22: Filter the set of interior point pairs to obtain a set of soft anchor points, and determine the local deformation field based on the set of soft anchor points.

[0009] Optionally, step S21 includes: Step S211: Extract feature points from the second two-dimensional grayscale image and the first two-dimensional grayscale image, and calculate the descriptor of the feature points; Step S212: By matching the descriptors of the feature points, determine the corresponding point set of the second two-dimensional grayscale image and the first two-dimensional grayscale image; Step S213: Estimate the global homography transformation matrix based on the corresponding point set, and generate an inner point pair set based on the point pairs in the corresponding point set that conform to the global homography transformation matrix.

[0010] Optionally, after step S213, the method further includes: Step S215: Calculate the interior point rate and / or residual distribution based on the set of interior point pairs and the set of corresponding points; Step S216: Evaluate the global homography transformation matrix based on the inlier rate and / or the residual distribution; In step S217, if the evaluation result does not meet the preset conditions, step S211 is re-executed by adjusting the feature density and / or matching threshold.

[0011] Optionally, in step S22, the step of filtering the set of interior point pairs to obtain a set of soft anchor points includes: The set of interior point pairs is used to remove interior point pairs whose reprojection error is greater than a set value in order to generate a set of soft anchor points.

[0012] Optionally, in step S22, determining the local deformation field based on the set of soft anchor points includes: Step S221: Construct a parameterized local deformation field; Step S222: Determine the data constraint information of the local deformation field based on the set of soft anchor points; Step S223: Determine the regularization constraint information of the local deformation field by regularizing the local deformation field; Step S224: By analyzing the structure of the first two-dimensional grayscale image and the second two-dimensional grayscale image, the structural constraint information of the local deformation field is determined; Step S225: Optimize the parameters of the local deformation field based on the data constraint information, the regularization constraint information, and the structural constraint information to determine the optimal parameters of the local deformation field.

[0013] Optionally, step S222 includes: Step S2221: Calculate the quality confidence of each soft anchor point in the set of soft anchor points; Step S2222: Determine the soft-matching term of the local deformation field based on the quality confidence level; Step S2223: Determine the soft field term of the local deformation field based on the mass confidence level; Step S2224: Determine the data constraint information of the local deformation field based on the soft matching term and the soft field term.

[0014] Optionally, step S223 includes: Step S2231: Determine the smoothing term of the local deformation field by determining the degree of curvature of the local deformation field; Step S2232: Determine the stiffness term of the local deformation field by determining the degree to which the local deformation deviates from the rigid body transformation; Step S2233: Determine the regular constraint information of the local deformation field based on the smoothing term and the rigidity term.

[0015] Optionally, step S223 further includes: Step S2234: Determine the topological barrier term of the local deformation field by determining the rate of change of the local area before and after deformation; Step S2233 includes: The regular constraint information of the local deformation field is determined based on the smoothing term, the rigidity term, and the topological barrier term.

[0016] Optionally, step S224 includes: Step S2241: Extract the edge set, line segment set, and corner point set from the first two-dimensional grayscale image and the second two-dimensional grayscale image, respectively; Step S2242: Calculate the edge field, line segment field, and corner field based on the extracted edge set, line segment set, and corner point set; Step S2243: Calculate the edge soft penalty term, the line segment soft penalty term, and the corner soft penalty term based on the edge field, the line segment field, and the corner field, respectively. Step S2244: Determine the main direction consistency term based on the main directions of the local regions in the first two-dimensional grayscale image and the second two-dimensional grayscale image; Step S2245: Determine the structural constraint information of the local deformation field based on the edge soft penalty term, the line segment soft penalty term, the corner soft penalty term, and the main direction consistency term.

[0017] Optionally, step S225 includes: Step S2251: Construct the overall objective function of the local deformation field based on the data constraint information, the regularization constraint information, and the structural constraint information; Step S2252: A multi-scale optimization strategy is adopted to optimize the parameters of the local deformation field in order to determine the optimal parameters of the local deformation field.

[0018] Optionally, step S2252 includes: Based on the soft field term and the smoothing term, the parameters of the local deformation field are coarsely optimized; Based on the edge soft penalty term, the line segment soft penalty term, the corner soft penalty term, and the rigidity term, the parameters of the local deformation field are optimized at a mid-level. Based on the soft matching term, the parameters of the local deformation field are optimized in a fine-grained manner.

[0019] Optionally, in step S2252, the parameters of the local deformation field are optimized according to the following manner: The parameters of the local deformation field are optimized using the Gauss-Newton or L-BFGS algorithm.

[0020] Optionally, step S3 includes: Step S31: Perform a global transformation on the first two-dimensional color image according to the global homography transformation matrix to obtain a third two-dimensional color image; Step S32: Based on the local deformation field, perform local deformation on the third two-dimensional color image to obtain a second two-dimensional color image.

[0021] Optionally, it also includes: Step S5: Determine the registration confidence map based on the set of interior point pairs; Step S5 includes: Step S51: Calculate the matching density, local residual, and texture energy of each local region based on the set of interior point pairs. Step S52: Calculate the confidence score for each local region based on the matching density and / or local residual and / or texture energy and / or soft field term consistency of each local region. Step S53: Determine the registration confidence map based on the confidence score of each local region.

[0022] Optionally, step S32 includes: Step S321: Based on the local deformation field, perform local deformation on each local region of the third two-dimensional color image; Step S322: Based on the registration confidence map, local areas with confidence scores less than a set threshold are identified as backtracking regions. For backtracking regions, the results of global homography are used to replace the results of local deformation to obtain a second two-dimensional color image.

[0023] Optionally, step S4 includes: Step S41: Detect the geometric boundary region of the scene on the depth image; Step S42: Determine the color assignment weight map based on the geometric boundary region, the registration confidence map, and the echo quality of the lidar sensor; Step S43: For each pixel to be colored in the depth image, perform the following steps to generate a color depth image: If its color assignment weight is greater than 0, then the color value of the corresponding pixel in the second two-dimensional color image is used as its initial color value. If its color assignment weight is 0, then its color value is recorded as empty.

[0024] Optionally, step S43 includes: For pixels whose color assignment weight is greater than or equal to the preset color assignment weight threshold, their initial color value is directly used as the final color value. For pixels whose color assignment weight is less than the preset color assignment weight threshold, the initial color value of the pixel to be assigned is fused with the initial color values ​​of other pixels in the neighboring area based on their respective color assignment weights, and the fused color value is used as the final color value. For pixels with missing color values, the initial color value of one other pixel in the neighboring region with a color assignment weight greater than the preset color assignment weight threshold is used, or the initial color values ​​of at least two other pixels in the neighboring region are fused based on their respective color assignment weights, and the fused color value is used as the final color value.

[0025] Optionally, step S42 includes: Step S421: Based on the geometric boundary region, the registration confidence map, and the echo quality of the lidar sensor, determine the registration confidence weight, edge suppression weight, and echo quality weight for each pixel, respectively. Step S422: Normalize the registration confidence weight, the edge suppression weight, and the echo quality weight respectively, and multiply the normalized registration confidence weight, edge suppression weight, and echo quality weight to obtain the color assignment weight for each pixel.

[0026] Optionally, step S4 further includes: Step S44: Filter the color depth image.

[0027] Optionally, after step S4, the method further includes: Each pixel of the color depth image is back-projected into a three-dimensional space, and its three-dimensional coordinates are determined to generate a 3D image; The color value of each pixel in the color depth image is assigned to the corresponding pixel in the 3D image to generate a color point cloud map.

[0028] Optionally, it also includes: The color point cloud image is subjected to additional processing, wherein the additional processing includes at least one of the following: outlier noise removal processing; color consistency optimization processing.

[0029] The present invention also constructs a computer device, including an image sensor, a lidar sensor, a processor, and a memory storing a computer program, wherein the processor implements the steps of the fusion processing method described above when executing the computer program.

[0030] Through the technical solution of this invention, since the first two-dimensional grayscale image from the image sensor and the second two-dimensional grayscale image from the lidar sensor have the same image properties and strong feature consistency, registering the first two-dimensional grayscale image with the second two-dimensional grayscale image can obtain more accurate registration information. Then, based on this registration information, the first two-dimensional color image from the image sensor is transformed to obtain a second two-dimensional color image that is more closely matched to the depth image originating from the second two-dimensional grayscale image. Using the color information of the transformed second two-dimensional color image to assign color to the corresponding pixels in the depth image can greatly reduce color misalignment in the generated color depth image. Attached Figure Description

[0031] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 This is a flowchart of a color image and depth image fusion processing method according to an embodiment of the present invention; Figure 2 This is a flowchart of a method for generating and fusing color images and depth images according to an embodiment of the present invention; Figure 3 This is a logical structure diagram of a computer device according to an embodiment of the present invention. Detailed Implementation

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] Figure 1 This is a flowchart of a color image and depth image fusion processing method according to an embodiment of the present invention. This fusion processing method is used to fuse a color image and a depth image, that is, to map the color information of the color image to the coordinate system of the depth image to obtain a color depth image. In this embodiment, the color image and depth image fusion processing method includes: Step S1: Obtain a first two-dimensional color image and a first two-dimensional grayscale image of the current field of view from the image sensor, and obtain a depth image and a second two-dimensional grayscale image of the current field of view based on the output of the lidar sensor. In this step, the image sensor can output corresponding color and grayscale images within the current field of view (FOV), namely, a first two-dimensional color image and a first two-dimensional grayscale image. The lidar sensor includes a transmitting module (TX) and a receiving module (RX), wherein the transmitting module is used to transmit sensing light signals into the current field of view for sensing. In some embodiments, the lidar performs ranging sensing using the Time Correlated Single Photon Counting (TCSPC) principle, where the sensing light signal consists of multiple rounds of light pulses emitted for each sensing operation. The receiving module is used to receive echo signals returned from the current field of view. Therefore, the depth image and laser intensity image of the current field of view can be obtained by analyzing the echo signals. Since the laser intensity of the echo signal is proportional to the grayscale of the field of view image, the laser intensity image can be considered as a second two-dimensional grayscale image. It should be noted that the first two-dimensional grayscale image and the second two-dimensional grayscale image can have different resolutions; for example, the first two-dimensional grayscale image from the image sensor has a higher resolution, while the second two-dimensional grayscale image from the lidar sensor has a lower resolution.

[0034] Step S2: Register the first two-dimensional grayscale image with the second two-dimensional grayscale image to obtain registration information; Step S3: Based on the registration information, transform the first two-dimensional color image to obtain a second two-dimensional color image; Step S4: Assign color information to the corresponding pixels in the depth image based on the second two-dimensional color image to generate a color depth image.

[0035] Through the technical solution of this embodiment, since the first two-dimensional grayscale image from the image sensor and the second two-dimensional grayscale image from the lidar sensor have the same image properties and strong feature consistency, registering the first two-dimensional grayscale image with the second two-dimensional grayscale image can obtain more accurate registration information. Then, based on this registration information, the first two-dimensional color image from the image sensor is transformed to obtain a second two-dimensional color image that is more closely matched to the depth image originating from the second two-dimensional grayscale image. Using the color information of the transformed second two-dimensional color image to assign color to the corresponding pixels on the depth image can greatly reduce color misalignment in the generated color depth image.

[0036] Furthermore, in an optional embodiment, combining Figure 2The lidar sensor 20 includes a transmitting module 21 for emitting sensing light signals into the current field of view, and a receiving module 22 for receiving echo signals returned from the current field of view. In step S1, acquiring the second two-dimensional grayscale image of the current field of view includes: detecting the echo signals and outputting corresponding light-sensing signals, and acquiring the second two-dimensional grayscale image of the current field of view based on the number of light-sensing signals counted within a preset statistical time period. In this embodiment, the second two-dimensional grayscale image is generated by a detection module 30 disposed outside the lidar sensor 20. It should be understood that in other embodiments, the detection module 30 may also be built into the lidar sensor 20, so that the lidar sensor 20 can directly output the second two-dimensional grayscale image.

[0037] Further, in an optional embodiment, in step S1, acquiring the depth image of the current field of view includes: calculating the time interval between the emission time of the sensing light signal and the reception time of the echo signal, and acquiring the depth image of the current field of view based on the time interval. In this embodiment, the depth image is generated by a detection module 30 disposed outside the lidar sensor 20. It should be understood that in other embodiments, the detection module 30 may also be built into the lidar sensor 20, so that the lidar sensor 20 can directly output the depth image.

[0038] In some embodiments, the detection module 30 uses a single-photon avalanche diode to detect the echo signal, and generates the depth image and the second two-dimensional grayscale image by different processing and analysis of the output light-sensing signal. For example, the detection module 30 performs statistical analysis on the output time of the light-sensing signal according to the TCSPC principle to obtain a histogram, and obtains the depth information of the region corresponding to the sensing pixel by analyzing the histogram, thereby generating a depth image; the detection module 30 obtains the grayscale information of the region corresponding to the sensing pixel by counting the number of light-sensing signals within a preset statistical period, thereby generating the second two-dimensional grayscale image of the current field of view.

[0039] Furthermore, in an optional embodiment, between step S1 and step S2, the following step is further included: The first two-dimensional color image, the depth image, the first two-dimensional grayscale image, and the second two-dimensional grayscale image are preprocessed, wherein the preprocessing includes at least one of the following: resolution adjustment processing, field of view adjustment processing, distortion correction processing, affine transformation processing, and timestamp synchronization processing.

[0040] Regarding resolution adjustment, it should be noted that since the resolutions of the image sensor and the LiDAR sensor are inconsistent, the resolutions of the first 2D color image, the depth image, the first 2D grayscale image, and the second 2D grayscale image can be adjusted to make them as consistent as possible. For example, when the resolution of the image sensor is higher than that of the LiDAR sensor, the first 2D color image and the first 2D grayscale image can be downsampled to match their resolutions with those of the depth image and the second 2D grayscale image. This reduces the deviation caused by scale differences during subsequent registration.

[0041] Regarding the field of view adjustment process, it should be noted that since the field of view of the image sensor and the lidar sensor are not the same, the field of view of the first two-dimensional color image, the depth image, the first two-dimensional grayscale image and the second two-dimensional grayscale image can be adjusted to make them as consistent as possible, thereby reducing the deviation caused by scale differences during subsequent registration.

[0042] Regarding distortion correction, it should be noted that image distortion is caused by the inherent parameters of the image sensor and LiDAR sensor. Therefore, distortion correction and coordinate alignment are required for the first two-dimensional color image, depth image, first two-dimensional grayscale image, and second two-dimensional grayscale image to ensure initial geometric alignment between the images from different sensors, providing a unified benchmark for subsequent matching. For example, distortion correction processing is performed on the first two-dimensional color image and the first two-dimensional grayscale image based on the image sensor's calibration parameters (lens factory calibration parameters, such as focal length, optical center position, and distortion coefficients).

[0043] Regarding affine transformation processing, it should be noted that, due to the difference in the sensing surfaces of the image sensor and the lidar sensor, an affine transformation from the image sensor plane to the lidar sensor plane needs to be applied to the first two-dimensional color image and the first two-dimensional grayscale image in order to correct the difference in the sensing surfaces of the image sensor and the lidar sensor.

[0044] Regarding timestamp synchronization, it should be noted that due to differences in image acquisition time and device orientation between image sensors and LiDAR sensors, it is necessary to utilize these time differences and device orientation differences to perform time difference correction on asynchronously acquired image data. For example, interpolation is used to calculate the relative sensor orientation at the moment of exposure / scanning to compensate for pose differences caused by shutter rolling or scanning delays. In dynamic scenes, this step can significantly reduce image misalignment caused by motion.

[0045] It should also be noted that the first two-dimensional color image, depth image, first two-dimensional grayscale image, and second two-dimensional grayscale image after the above preprocessing can be used as stable inputs for subsequent image registration.

[0046] Furthermore, in an optional embodiment, the registration information includes: a global homography transformation matrix and a local deformation field. Combined with... Figure 2 Step S2 includes: Step S21: Perform image matching on the first two-dimensional grayscale image and the second two-dimensional grayscale image to obtain the global homography transformation matrix and the set of interior point pairs; In this step, the global homography transformation matrix is ​​used to describe the projection correspondence between two grayscale images.

[0047] Step S22: Filter the set of interior point pairs to obtain a set of soft anchor points, and determine the local deformation field based on the set of soft anchor points.

[0048] In this embodiment, when registering the first two-dimensional grayscale image and the second two-dimensional grayscale image, a two-level registration framework of global + local is adopted, which can improve the accuracy and stability of cross-modal registration and obtain stable and detailed consistent alignment results even in dynamic scenes.

[0049] Further, in an optional embodiment, step S21 includes: Step S211: Extract feature points from the second two-dimensional grayscale image and the first two-dimensional grayscale image, and calculate the descriptor of the feature points; In this step, for example, robust feature points can be extracted from the preprocessed first two-dimensional grayscale image and the second two-dimensional grayscale image using feature extraction methods such as SIFT, ORB, and AKAZE. Then, descriptors of the feature points are calculated. The descriptors are feature vectors used to describe the region surrounding the feature points.

[0050] Step S212: By matching the descriptors of the feature points, determine the corresponding point set of the second two-dimensional grayscale image and the first two-dimensional grayscale image; In this step, for example, FLANN or brute-force matching algorithms can be used to match the descriptors of the feature points. This can filter out high-quality, reliable corresponding point pairs from a large number of feature points, laying a solid foundation for subsequent geometric calculations. These corresponding points (i.e., matching points) are points in the same or similar positions in the first and second two-dimensional grayscale images. Therefore, the set of points obtained through matching is the candidate set of corresponding points between the two grayscale images.

[0051] Step S213: Estimate the global homography transformation matrix based on the corresponding point set, and generate an inner point pair set based on the point pairs in the corresponding point set that conform to the global homography transformation matrix.

[0052] In this step, it should be noted that the purpose of image registration is to find the identical parts in two images or to calculate the transformation relationship between them (such as translation, rotation, scaling, etc.). The set of corresponding points found is the basis of image registration. Using these corresponding points, the geometric transformation between the two grayscale images is further estimated, i.e., the global homography transformation matrix.

[0053] Furthermore, between steps S212 and S213, the following is also included: Step S214: Optimize the set of corresponding points determined in step S212 to obtain an optimized set of corresponding points; Step S213 includes: Based on the optimized set of corresponding points, estimate the global homography transformation matrix, and divide the optimized set of corresponding points into a set of interior point pairs and a set of exterior point pairs.

[0054] In this embodiment, the set of corresponding points determined in step S212 is the initial matching point pair. In order to improve the accuracy of registration, the set of corresponding points needs to be optimized. Step S213 uses the optimized set of corresponding points to estimate the global homography transformation matrix.

[0055] Further, in an optional embodiment, step S214 includes: The ratio test method is used to remove errors from the set of corresponding points determined in step S212; A symmetric verification method is used to perform symmetric verification on the set of corresponding points after erroneous removal in order to obtain an optimized set of corresponding points.

[0056] In this embodiment, since correct matches are more significant than incorrect matches, Lowe's ratio test can be used first to eliminate incorrect matches. When performing symmetric verification on the corresponding point set, a first corresponding point set can be obtained by matching the first two-dimensional grayscale image to the second two-dimensional grayscale image, and then a second corresponding point set can be obtained by matching the second two-dimensional grayscale image to the first two-dimensional grayscale image, retaining only the consistent corresponding points between the two. Through the corresponding point set optimization method of this embodiment, a more reliable candidate corresponding point set can be obtained.

[0057] Further, in an optional embodiment, step S213 includes: employing a robust estimation algorithm to estimate the global homography transformation matrix based on the corresponding point set, and determining the set of interior point pairs. In this embodiment, a robust estimation algorithm, such as RANSAC or MAGSAC++, is used to estimate the global two-dimensional homography transformation matrix. This algorithm can not only automatically eliminate mismatches during iteration, but also divide the corresponding points into inliers and outliers. Inliers are point pairs that conform to the estimated homography transformation model, while outliers are point pairs that do not conform to the model. Moreover, the optimal global homography transformation matrix is ​​calculated using only the set of inlier pairs.

[0058] Furthermore, after step S213, the method further includes: Step S215: Calculate the interior point rate and / or residual distribution based on the set of interior point pairs and the set of corresponding points; In this step, the inlier ratio can be calculated by dividing the number of inliers by the total number of matched point pairs. A higher inlier ratio indicates that more point pairs conform to the estimated model, the better the matching quality, and the more reliable the estimated homography transformation. A lower inlier ratio indicates that the model may be incorrect or the scene is not suitable for homography transformation.

[0059] In this step, the residual distribution refers to the distribution of reprojection errors for each matched point pair after homography. The reprojection error is the Euclidean distance between a point in one grayscale image projected onto the coordinate system of another grayscale image and its corresponding point in the other grayscale image. The residual distribution characterizes the accuracy of each interior point match from a local perspective. A smaller residual indicates better alignment of the matched point pairs. Furthermore, if the residual distribution is mostly concentrated in very small values ​​(e.g., less than 1 pixel), the registration is very accurate; if the residual distribution is scattered, there is a large alignment error. Therefore, small and concentrated residuals indicate high-precision registration; large or scattered residuals indicate poor registration quality.

[0060] Step S216: Evaluate the global homography transformation matrix based on the inlier rate and / or the residual distribution; In this step, the quality of registration can be comprehensively evaluated by combining the in-point rate and the residual distribution. The in-point rate represents how many matches are reliable, and the residual distribution represents the accuracy of these reliable matches.

[0061] In step S217, if the evaluation result does not meet the preset conditions, step S211 is re-executed by adjusting the feature density and / or relaxing the matching threshold.

[0062] Through the above embodiments, step S21 can obtain the global homography transformation matrix and the corresponding set of interior point pairs (including the coordinates of the matching points and their error and other quality statistics). These results provide reliable initial solutions and seed control points for subsequent local deformation.

[0063] Further, in an optional embodiment, in step S22, the step of filtering the set of interior point pairs to obtain a set of soft anchor points includes: removing interior point pairs from the set of interior point pairs whose reprojection errors are greater than a set value, thereby generating a set of soft anchor points. In this embodiment, for the M interior point pairs (pi, qi) obtained in step S21, abnormal matching pairs with large reprojection errors can be removed. Specifically, the image transformed from the first two-dimensional grayscale image using a global homography transformation matrix is ​​projected onto the space of the second two-dimensional grayscale image. Based on the projection result, the matched points are the soft anchor points.

[0064] Further, in an optional embodiment, in step S22, determining the local deformation field based on the set of soft anchor points includes: Step S221: Construct a parameterized local deformation field; In this step, it is first explained that the local deformation field is a continuous two-dimensional deformation field, which can capture complex local deformations and can be used for non-rigid image registration. A continuous two-dimensional deformation field refers to a displacement field that changes continuously in two-dimensional space. That is, for each point in the image, there is a displacement vector (including the x and y directions) to represent the deformation of the point from one image to another. Therefore, it is usually represented by a two-dimensional array (or two two-dimensional arrays, representing the displacements in the x and y directions respectively), and the size of the array is the same as that of the image. Therefore, the local deformation field has the following characteristics: (1) Continuity: The deformation field is a continuous function, which can be realized by interpolation (such as linear interpolation, B-spline, etc.), so that the displacements of adjacent points are smoothly transitioned, avoiding the occurrence of breaks or abrupt changes; (2) Denseness: The deformation field defines a displacement for each point, not just feature points, and can deform the entire image. Based on the above characteristics, the registered image can be generated by resampling the source image through the local deformation field (such as bilinear interpolation).

[0065] In this step, a parameterized local deformation field can be constructed in various ways, such as based on the deformation of physical models (e.g., elasticity models, fluid models, etc.); based on the deformation of splines (e.g., thin-plate splines (TPS), B-splines, etc.); based on optical flow fields (e.g., dense optical flow fields); or based on deep learning models (neural network models). Taking B-spline interpolation as an example, the functional form of the local deformation field is: ,in, For local deformation field, This represents the grid point displacement, i.e., the parameter to be optimized. The basis function is B-spline, with a grid spacing of 32–64 pixels, and a multi-scale strategy is used to optimize it step by step from coarse to fine.

[0066] Step S222: Determine the data constraint information of the local deformation field based on the set of soft anchor points; In this step, data constraint information makes , where (pi,qi) are soft anchor point pairs.

[0067] Step S223: Determine the regularization constraint information of the local deformation field by regularizing the local deformation field; In this step, in order to ensure the smoothness of deformation and local stiffness, regular constraint information is introduced. The regular constraint information is used to constrain the physical rationality of the local deformation field and prevent overfitting.

[0068] Step S224: By analyzing the structure of the first two-dimensional grayscale image and the second two-dimensional grayscale image, the structural constraint information of the local deformation field is determined; In this step, structural constraint information is introduced to ensure the structural consistency of deformation. The structural constraint information is used to constrain the consistency of geometric structures such as edges, line segments, corners, and abrupt changes in normals, without requiring a pixel-by-pixel correspondence.

[0069] Step S225: Optimize the parameters of the local deformation field based on the data constraint information, the regularization constraint information, and the structural constraint information to determine the optimal parameters of the local deformation field.

[0070] In this embodiment, the parameters of the local deformation field are optimized based on data constraint information, regularization constraint information and structural constraint information. This ensures that the local deformation field can accurately and stably register the image, avoid overfitting, and improve robustness, especially under conditions such as dynamic scenes and weak textures.

[0071] Further, in an optional embodiment, step S222 includes: Step S2221: Calculate the quality confidence of each soft anchor point in the set of soft anchor points; In this step, for example, the quality confidence ci of each soft anchor can be calculated based on the ratio test score, RANSAC / MAGSAC++ residuals, and local texture energy, where 0 ≤ ci ≤ 1.

[0072] Step S2222: Determine the soft-matching term of the local deformation field based on the quality confidence level; In this step, a robust soft-matching term is introduced for each soft anchor point to reflect the degree of registration of the soft anchor point before and after deformation. For example, the soft-matching term is: in, For soft matches, M represents the number of soft anchors. The quality confidence level is given by α, where α is the soft-match coefficient. Here is the Huber or Tukey loss calculation function, and (pi,qi) is a soft anchor pair. To map the coordinates of pi using a local deformation field, To match the uncertainty coefficient.

[0073] Step S2223: Determine the soft field term of the local deformation field based on the mass confidence level; In this step, to alleviate the sparsity problem of soft anchors, a nuclear diffusion term is introduced to extend the influence range of the soft anchors to their surrounding neighborhood. Therefore, a soft field term is introduced, for example: in, For soft field terms, For the neighborhood of pi, For example, a kernel function, such as a B-spline kernel. To influence the radius, x represents the pixel within the neighborhood of the soft anchor point pi. To map the x-coordinates using a local deformation field, This is the tolerance standard.

[0074] Step S2224: Determine the data constraint information of the local deformation field based on the soft matching term and the soft field term.

[0075] In this step, data constraint information is determined by summing the soft matching term and the soft field term.

[0076] Further, in an optional embodiment, step S223 includes: Step S2231: Determine the smoothing term of the local deformation field by determining the degree of curvature of the local deformation field; In this step, a smoothing term is introduced to ensure the smoothness of the deformation. This smoothing term is, for example: in, For smoothing terms, The second derivative of the local deformation field. It is the square of the Frobenius norm.

[0077] Step S2232: Determine the stiffness term of the local deformation field by determining the degree to which the local deformation deviates from the rigid body transformation; In this step, to ensure local rigidity of the deformation, a rigidity term is introduced, which is, for example: in, It is a rigid term. Let be the Jacobian matrix of the local deformation field. The nearest rotation matrix is ​​the extreme decomposition of the Jacobian matrix; the difference between the two characterizes the degree of deviation from a rigid body (tension / shear). The smaller the value, the closer the local deformation is to rigid body deformation.

[0078] Step S2233: Determine the regular constraint information of the local deformation field based on the smoothing term and the rigidity term.

[0079] Furthermore, step S223 also includes: Step S2234: Determine the topological barrier term of the local deformation field by determining the rate of change of the local area before and after deformation; In this step, to suppress folding and flipping, a topology barrier item can also be set, such as: And step S2233 includes: The regular constraint information of the local deformation field is determined based on the smoothing term, the rigidity term, and the topological barrier term.

[0080] Further, in an optional embodiment, step S224 includes: Step S2241: Extract the edge set, line segment set, and corner point set from the first two-dimensional grayscale image and the second two-dimensional grayscale image, respectively; Step S2242: Calculate the edge field, line segment field, and corner field based on the extracted edge set, line segment set, and corner point set; In this step, for the binary structure (edges, line segments, corners) extracted from the reference mode (first two-dimensional grayscale image), its Euclidean distance transform is calculated to obtain the edge field D. edge Line segment field D line Corner Field D corner For the binary structures (edges, line segments, and corners) extracted from the second-dimensional grayscale image, their Euclidean distance transforms are calculated to obtain the edge field ε. d Line segment field L d Corner Field C d .

[0081] Step S2243: Calculate the edge soft penalty term, the line segment soft penalty term, and the corner soft penalty term based on the edge field, the line segment field, and the corner field, respectively. In this step, the soft penalty terms applied to the structural point x of the second two-dimensional grayscale image and then to the first two-dimensional grayscale image are as follows: in, This is a borderline soft penalty item. This is a soft penalty term for line segments. For corner point soft penalty items, , , These are the weights for edges, line segments, and corner points, respectively. It is either the Charbonnier or Soft-L1 algorithm.

[0082] Step S2244: Determine the main direction consistency term based on the main directions of the local regions in the first two-dimensional grayscale image and the second two-dimensional grayscale image; In this step, to suppress the distortion of the scene's principal direction caused by deformation, a principal direction consistency term is introduced. This principal direction consistency term is, for example: in, Consistent terms in the main direction The weights for the consistency between the normal mutation and the principal direction are denoted as . Let x be the local principal direction in the second two-dimensional grayscale image. The x-direction is the local principal direction after x-deformation in the first two-dimensional grayscale image. This is for calculating cosine similarity.

[0083] Step S2245: Determine the structural constraint information of the local deformation field based on the edge soft penalty term, the line segment soft penalty term, the corner soft penalty term, and the main direction consistency term.

[0084] Further, in an optional embodiment, step S225 includes: Step S2251: Construct the overall objective function of the local deformation field based on the data constraint information, the regularization constraint information, and the structural constraint information; In this step, for example, the constructed overall objective function is: in, For data constraints, For regularization constraints, For structural constraints, , , These are the weights for smoothness, rigidity, and topological barrier, respectively.

[0085] Step S2252: A multi-scale optimization strategy is adopted to optimize the parameters of the local deformation field in order to determine the optimal parameters of the local deformation field.

[0086] Further, step S2252 includes: Based on the soft field term and the smoothing term, the parameters of the local deformation field are coarsely optimized; Based on the edge soft penalty term, the line segment soft penalty term, the corner soft penalty term, and the rigidity term, the parameters of the local deformation field are optimized at a mid-level. Based on the soft matching term, the parameters of the local deformation field are optimized in a fine-grained manner.

[0087] In this embodiment, when employing a multi-scale optimization strategy to progressively refine the deformation field, for example, an image pyramid is first constructed, with B-spline grid spacing set from coarse to fine (e.g., 96→64→32 pixels). Then, the coarse layer... Primarily for fitting large-scale deformations; middle layer introduces 4. Correcting the structure and suppressing non-physical distortion; 5. Adding fine layers Complete the refinement (robust loss suppresses outliers).

[0088] Further, in step S2252, the parameters of the local deformation field are optimized in the following manner: The Gauss-Newton or L-BFGS algorithm is used to optimize the parameters of the local deformation field. For example, the Gauss-Newton or L-BFGS algorithm can be used to update the parameters of the local deformation field at each layer, and the upper layer solution is upsampled as initialization. Further, in an optional embodiment, combined with... Figure 2 Step S3 includes: Step S31: Perform a global transformation on the first two-dimensional color image according to the global homography transformation matrix to obtain a third two-dimensional color image; Step S32: Based on the local deformation field, perform local deformation on the third two-dimensional color image to obtain a second two-dimensional color image.

[0089] In this embodiment, the registration information obtained through step S2 includes a global homography transformation matrix and a local deformation field. Therefore, when transforming the first two-dimensional color image according to this registration information in step S3, a global transformation is first performed based on the global homography transformation matrix, and then a local deformation is performed on the globally transformed image based on the local deformation field. This two-level registration framework, which combines global homography with local continuous deformation, can ensure the stability of the overall registration and avoid the destruction of the entire image by local misalignment in dynamic or weak texture conditions.

[0090] Furthermore, in an optional embodiment, it further includes: Step S5: Determine the registration confidence map based on the set of interior point pairs; Regarding this step, it should be noted that this step can occur after step S22 or after step S32. Furthermore, step S5 specifically includes: Step S51: Calculate the matching density, local residual, and texture energy of each local region based on the set of interior point pairs. In this step, Match Density is the number of feature points that can be successfully matched within a local region. In practice, the image can be divided into multiple grids (e.g., each grid is 16×16 pixels), the number of interior points within each grid is counted, and then a Gaussian filter is used for smoothing to obtain a continuous power density map. Local Residual is the reprojection error of the matched feature points within a local region. In practice, for each grid, the average reprojection error of all interior points within that grid is calculated. The reprojection error is calculated by transforming the points in the first 2D grayscale image to the image to be registered through global homography and local deformation field transformation, and then calculating the Euclidean distance between the points and the corresponding matching points in the second 2D grayscale image. Texture Energy is the richness of texture in a local region, which can be measured using image gradient energy or variance.

[0091] It should be understood that after calculating the matching density, local residuals, and texture energy, they should be normalized separately. For example, the power density is divided by the maximum number of matches in the neighborhood; the local residuals are mapped to values ​​in the range of 0 to 1 through a threshold; and the texture energy is normalized by contrast.

[0092] Step S52: Calculate the confidence score for each local region based on the matching density and / or local residual and / or texture energy and / or soft field term consistency of each local region. In this step, it should be noted that higher matching density corresponds to higher confidence; lower local residuals also result in higher confidence; higher texture energy indicates higher reliability of feature matching and thus higher confidence; the consistency of the soft field term serves as an indirect evaluation indicator, reflecting the accuracy of registration. Smaller deviations indicate greater effectiveness of soft constraints in region registration, leading to higher confidence. When calculating the confidence score, the four indicators can be combined into a single confidence score. For example, a weighted average (with weights set empirically for each component) or multiplication method can be used.

[0093] Step S233: Determine the registration confidence map based on the confidence score of each local region. Further, step S32 includes: Step S321: Based on the local deformation field, perform local deformation on each local region of the third two-dimensional color image; Step S322: Based on the registration confidence map, local areas with confidence scores less than a set threshold are identified as backtracking regions. For backtracking regions, the results of global homography are used to replace the results of local deformation to obtain a second two-dimensional color image.

[0094] In this embodiment, when transforming the first two-dimensional color image according to the registration information in step S3, a global transformation is first performed based on the global homography transformation matrix, and then a local deformation is performed on the globally transformed image based on the local deformation field. Then, the registration confidence map is combined; for example, a confidence threshold (e.g., 0.3) can be preset. In regions where the confidence score is higher than the confidence threshold, the registration result of the local deformation field is used; in regions where the confidence score is lower than the confidence threshold, the registration result of the global homography transformation is used. This reduces noise drag and error propagation. In other words, if local non-rigid registration in a certain region is unreliable, it is preferable to maintain rigid global alignment to avoid severe mismatch caused by erroneous deformation.

[0095] Furthermore, in an optional embodiment, after completing the stability optimization of geometric registration, in order to ensure the correct propagation of color at the object edges, the coloring stage is entered through step S4, such as... Figure 2 As shown. Step S4 of this embodiment includes: Step S41: Detect the geometric boundary region of the field of view on the depth image; In this step, the edge band of the depth image is extracted by detecting the geometric boundary region of the field of view, i.e., the edge band E. Specifically, the edge intensity of the depth image is first calculated, for example, by using the Sobel operator to obtain the gradient magnitude map, and then the initial set of depth edge pixels (representing the boundary of sudden depth changes) is obtained through thresholding. Then, the bandwidth of this edge pixel set is expanded. Based on both sides of the edge normal, a certain pixel range around the edge (the bandwidth width can be adaptively determined according to the depth noise level) is marked as the edge influence area, forming the edge band E. The edge band E covers the area near the object contour and is a sensitive band that needs to be carefully processed for color assignment to avoid the "color bleeding" phenomenon that occurs when the color crosses the boundary of the real object.

[0096] Step S42: Determine the color assignment weight map based on the geometric boundary region, the registration confidence map, and the echo quality of the lidar sensor; In this step, a color assignment weight is defined to guide the intensity of color propagation, and this color assignment weight is determined by the following factors: geometric boundary region, registration confidence map, and echo quality of the LiDAR sensor. Therefore, a comprehensive color assignment weight can be formed by fusing multiple factors.

[0097] Step S43: For each pixel to be colored in the depth image, perform the following steps to generate a color depth image: If its color assignment weight is greater than 0, then the color value of the corresponding pixel in the second two-dimensional color image is used as its initial color value. If its color assignment weight is 0, then its color value is recorded as empty.

[0098] Further, in an optional embodiment, step S43 specifically includes: For pixels whose color assignment weight is greater than or equal to the preset color assignment weight threshold, their initial color value is directly used as the final color value. For pixels whose color assignment weight is less than the preset color assignment weight threshold, the initial color value of the pixel to be assigned is fused with the initial color values ​​of other pixels in the neighboring area based on their respective color assignment weights, and the fused color value is used as the final color value. For pixels with missing color values, the initial color value of one other pixel in the neighboring region with a color assignment weight greater than the preset color assignment weight threshold is used, or the initial color values ​​of at least two other pixels in the neighboring region are fused based on their respective color assignment weights, and the fused color value is used as the final color value.

[0099] In this embodiment, after obtaining the color weight map, the preliminary mapped color results are fused and filtered to preserve structure, eliminate visible seams and protect the depth structure. Specifically, firstly, the color image is projected onto the coordinate system of the depth image after undergoing global homography and local deformation to obtain the initial color map, i.e., the second two-dimensional color image.

[0100] For each pixel x in the depth image, if there is a corresponding pixel in the second two-dimensional color image, its color value can be obtained; if there is no corresponding pixel (e.g., in an echo-free region, the weight is 0), then the color value of pixel x is recorded as missing. Simultaneously, the color is modulated according to the color assignment weight. Specifically, for pixels with high color assignment weights, the initial color is directly used; for pixels with low color assignment weights, their color contribution should be weakened. For example, the color assignment weights can be used as interpolation coefficients to fuse the initial color with the neighboring colors (the color values ​​of several surrounding pixels). This avoids using unreliable colors.

[0101] Further, in an optional embodiment, step S42 includes: Step S421: Based on the geometric boundary region, the registration confidence map, and the echo quality of the lidar sensor, determine the registration confidence weight, edge suppression weight, and echo quality weight for each pixel, respectively. In this step, the registration confidence weights can be directly derived from the registration confidence map output by S3. It should be noted that, to facilitate integration with other weights, the registration confidence weights need to be normalized to 0-1 with a consistent direction (high indicates confidence). Specifically, linear stretching or piecewise linear functions can be used to map the original confidence values ​​to weight values. For example, the lowest confidence region can be mapped to 0, the highest to 1, with a transition function in between. After this processing, the registration confidence weights represent the geometric registration reliability of that pixel. High-confidence pixels can be confidently assigned colors, while low-confidence pixels should have their colors assigned less weight.

[0102] Edge suppression weights are used to prevent color propagation beyond object boundaries. Pixels within the edge band E are assigned reduced weight values. For example, pixels within edge band E can have an edge suppression weight of 0 (completely disregarding color propagation), while pixels outside edge band E can have an edge suppression weight of 1. Another approach is to assign weights based on distance from the edge, with closer pixels receiving lower weights; a Gaussian function can be used. ,in, This represents the distance from pixel x to the nearest edge. To control the width of the edge band, normalization is applied so that the color weight is close to 0 or low in the object boundary and its neighborhood E, while it returns to 1 in areas far from the boundary. This ensures that the color automatically decays near the depth boundary, thereby suppressing the phenomenon of color propagation across the object boundary.

[0103] The echo quality weight reflects the quality and reliability of the depth image itself. Some pixels in a LiDAR sensor may not return valid depth (no echo) or have a low signal-to-noise ratio due to material, distance, or other reasons. For no-echo areas (such as points where infrared ranging fails), a weight of 0 is directly assigned, indicating that these pixels should not be forcibly assigned color to avoid interfering with the color consistency of their neighborhood. For pixels with echoes but poor quality, continuous weights can be defined based on echo intensity or signal-to-noise ratio. For example, the LiDAR sensor's reflection intensity can be normalized to 0-1 as a weight; lower intensity results in a smaller echo quality weight. Multiple echo cases can also be considered: if a point is a sub-echo (subsurface) or the measurement is unstable, its echo quality weight can be reduced. By normalizing the mapping, we obtain the echo quality weights representing the reliability of each depth pixel, with 1 for high quality and 0 for poor quality.

[0104] Step S422: Normalize the registration confidence weight, the edge suppression weight, and the echo quality weight respectively, and multiply the normalized registration confidence weight, edge suppression weight, and echo quality weight to obtain the color assignment weight for each pixel.

[0105] In this step, after obtaining the three weights mentioned above and normalizing them, the final color assignment weights can be obtained according to a certain combination strategy. For example, a product strategy can be used to combine the three weights together, i.e.: ,in, Assign weights to colors, To register confidence weights, To register confidence weights, This refers to echo quality weights. The multiplicative combination means applying a gated approach to color assignment: if any factor indicates unreliability (corresponding to a weight close to 0), the final color weight will also be very low, effectively suppressing the influence of color propagation on that pixel. This method ensures that even in areas with low confidence, edges, or low measurement quality, if only a single factor is unfavorable for color assignment, it will be treated conservatively, avoiding erroneous color transmission. If all factors are good (all close to 1), the color assignment weight is close to 1, indicating that the color can be safely assigned to that pixel. The weights multiplied usually naturally fall within the range [0,1], requiring no further normalization. However, to prevent excessively small values ​​from causing the color to be "too dark," the final weight map can be slightly smoothed or a lower threshold can be set (e.g., a minimum limit of 0.1) to retain a small amount of basic color. In contrast, if a weighted sum combination is used, a problem may arise where an unfavorable factor is masked by compensation from other factors, making the multiplicative strategy less intuitive and effective.

[0106] Furthermore, in an optional embodiment, step S4 further includes: Step S44: Filter the color depth image.

[0107] In this embodiment, color discontinuities in the color depth image caused by time differences or registration errors can be smoothed through filtering. For example, algorithms such as joint bilateral filtering (using the depth image as a guide image to filter the color image) or guided filtering (using depth as the guide image and color as the filter input) can be used. These filters utilize depth edge information, weighting in both the spatial and color domains, and performing averaging only between pixels with similar depths (i.e., belonging to the same object plane) and similar colors, thus achieving edge-preserving noise reduction.

[0108] It should also be noted that since the edge regions are assigned extremely low weights in the color assignment, essentially meaning the color is empty or very weak at the boundaries, the bilateral / guided filtering will detect depth discontinuities and color gaps when crossing edges, thus automatically preventing color from spreading from one side to the other. Furthermore, explicit constraints can be imposed on the filtering implementation. For example, for pixels within the edge band E, color fusion is only allowed with pixels on the same side (within the same object), and not with pixels from the other side. This process aligns the color boundaries at object edges with the depth boundaries, preventing "color overflow." In non-edge regions, since the color assignment weights are close to 1, the filtering is essentially equivalent to averaging the colors within the same plane, effectively eliminating color gaps caused by multi-camera or registration micro-errors, resulting in a uniform and continuous color distribution. Simultaneously, the filtering preserves depth details and does not cause blurring by crossing the real structure. After filtering optimization, a color distribution that preserves structure and has no obvious seams is finally obtained.

[0109] Furthermore, in an optional embodiment, after step S4, the method further includes: Each pixel of the color depth image is back-projected into a three-dimensional space, and its three-dimensional coordinates are determined to generate a 3D image; The color value of each pixel in the color depth image is assigned to the corresponding pixel in the 3D image to generate a color point cloud map.

[0110] In this embodiment, using the intrinsic and extrinsic parameters of the depth camera, each pixel (u, v, Z) of the color depth image is back-projected into three-dimensional point coordinates (X, Y, Z), and the color of that pixel is assigned to the 3D point, resulting in a color point cloud. The specific calculations employ a typical camera model: , , among which, (c x ,c y (f) is the principal point of the camera. x ,f y ( ) represents the focal length in pixels. This calculation yields a point cloud that corresponds one-to-one with the pixels in the depth image, with each point possessing a color attribute. Since the color depth image ensures color and depth consistency, the corresponding point cloud color will faithfully match the object's surface.

[0111] Furthermore, in an optional embodiment, the fusion processing method further includes: The color point cloud image is subjected to additional processing, wherein the additional processing includes at least one of the following: outlier noise removal processing; color consistency optimization processing.

[0112] In this embodiment, the generated color point cloud can undergo additional processing as needed by the application. This includes removing outliers and noise, re-estimating normal vectors based on local neighborhoods for subsequent high-level perception, filling depth holes caused by echo-free conditions (which can be filled by interpolation of surrounding points), and further optimizing color consistency (e.g., ensuring smooth color transitions in overlapping areas for point cloud segments merged at different time differences). These steps improve the visual quality and structural integrity of the point cloud. If the application has high speed requirements, this step can be omitted or performed offline.

[0113] Finally, when the fusion processing method of the present invention is applied, two result data, a color depth image and a color point cloud, can be generated. They are strictly aligned in geometry and color and can be directly used for downstream tasks such as 3D reconstruction and environmental perception, realizing the integrated and unified output of 2D and 3D results.

[0114] Figure 3 This is a logical structure diagram of a computer device according to an embodiment of the present invention. The computer device in this embodiment includes an image sensor 10, a lidar sensor 20, a processor 50, and a memory 40. The memory 40 stores a computer program, and the processor 50 implements the steps of the fusion processing method described above when executing the computer program.

[0115] Those skilled in the art will understand that Figure 3 The examples shown are merely illustrations of computer devices and do not constitute a limitation on computer devices. In practice, computer devices may include more or fewer components than those shown, or combinations of certain components, or different components. For example, they may also include input / output modules, network access modules, etc.

[0116] Processor 50 may include one or more processing units, such as: a Central Processing Unit (CPU), an Application Processor (AP), a Modem Processor, a Graphics Processing Unit (GPU), an Image Signal Processor (ISP), a controller, memory, a video codec, a Digital Signal Processor (DSP), a Baseband Processor, and / or a Neural-Network Processing Unit (NPU), other general-purpose processors, Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor. Different processing units may be independent devices or integrated into one or more processors.

[0117] The controller can serve as the nerve center and command center of a computer device. Based on the instruction opcode and timing signals, the controller generates operation control signals to control the fetching and execution of instructions.

[0118] In some embodiments, memory 40 may be an internal storage unit of a computer device, such as a hard disk or RAM. In other embodiments, memory 40 may be an external storage device of a computer device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Optionally, memory 40 may include both internal and external storage units of the computer device. Memory 40 is used to store operating systems, applications, bootloaders, data, and other programs, such as the program code of the computer programs. Memory 40 may also be used to temporarily store data that has been output or will be output.

[0119] The processor 50 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 50 is a cache memory. This memory can store instructions or data that the processor 50 has just used or that are used repeatedly. If the processor 50 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 50, and thus improves the efficiency of the system.

[0120] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0121] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0122] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of the claims of the present invention.

Claims

1. A method for fusing color images and depth images, characterized in that, include: Step S1: Obtain a first two-dimensional color image and a first two-dimensional grayscale image of the current field of view from the image sensor, and obtain a depth image and a second two-dimensional grayscale image of the current field of view based on the output of the lidar sensor. Step S2: Register the first two-dimensional grayscale image with the second two-dimensional grayscale image to obtain registration information; Step S3: Based on the registration information, transform the first two-dimensional color image to obtain a second two-dimensional color image; Step S4: Assign color information to the corresponding pixels in the depth image based on the second two-dimensional color image to generate a color depth image.

2. The fusion processing method according to claim 1, characterized in that, The lidar sensor includes a transmitting module for transmitting sensing light signals into the current field of view, and a receiving module for receiving echo signals returned from the current field of view. In step S1, acquiring the second two-dimensional grayscale image of the current field of view includes: The echo signal is detected and a corresponding light-sensing signal is output. The second two-dimensional grayscale image of the current field of view is obtained based on the number of light-sensing signals counted within a preset statistical time period.

3. The fusion processing method according to claim 1, characterized in that, Between step S1 and step S2, the following is also included: The first two-dimensional color image, the depth image, the first two-dimensional grayscale image, and the second two-dimensional grayscale image are preprocessed, wherein the preprocessing includes at least one of the following: resolution adjustment processing, field of view adjustment processing, distortion correction processing, affine transformation processing, and timestamp synchronization processing.

4. The fusion processing method according to any one of claims 1-3, characterized in that, The registration information includes: a global homography transformation matrix and a local deformation field; step S2 includes: Step S21: Perform image matching on the first two-dimensional grayscale image and the second two-dimensional grayscale image to obtain the global homography transformation matrix and the set of interior point pairs; Step S22: Filter the set of interior point pairs to obtain a set of soft anchor points, and determine the local deformation field based on the set of soft anchor points.

5. The fusion processing method according to claim 4, characterized in that, Step S21 includes: Step S211: Extract feature points from the second two-dimensional grayscale image and the first two-dimensional grayscale image, and calculate the descriptors of the feature points; Step S212: By matching the descriptors of the feature points, determine the corresponding point set of the second two-dimensional grayscale image and the first two-dimensional grayscale image; Step S213: Estimate the global homography transformation matrix based on the corresponding point set, and generate an interior point pair set based on the point pairs in the corresponding point set that conform to the global homography transformation matrix.

6. The fusion processing method according to claim 5, characterized in that, Following step S213, the method further includes: Step S215: Calculate the interior point rate and / or residual distribution based on the set of interior point pairs and the set of corresponding points; Step S216: Evaluate the global homography transformation matrix based on the inlier rate and / or the residual distribution; In step S217, if the evaluation result does not meet the preset conditions, step S211 is re-executed by adjusting the feature density and / or the matching threshold.

7. The fusion processing method according to claim 4, characterized in that, In step S22, obtaining a set of soft anchor points by filtering the set of interior point pairs includes: The set of interior point pairs is used to remove interior point pairs whose reprojection error is greater than a set value in order to generate a set of soft anchor points.

8. The fusion processing method according to claim 7, characterized in that, In step S22, determining the local deformation field based on the set of soft anchor points includes: Step S221: Construct a parameterized local deformation field; Step S222: Determine the data constraint information of the local deformation field based on the set of soft anchor points; Step S223: Determine the regularization constraint information of the local deformation field by regularizing the local deformation field; Step S224: By analyzing the structure of the first two-dimensional grayscale image and the second two-dimensional grayscale image, the structural constraint information of the local deformation field is determined; Step S225: Optimize the parameters of the local deformation field based on the data constraint information, the regularization constraint information, and the structural constraint information to determine the optimal parameters of the local deformation field.

9. The fusion processing method according to claim 8, characterized in that, Step S222 includes: Step S2221: Calculate the quality confidence of each soft anchor point in the set of soft anchor points; Step S2222: Determine the soft-matching term of the local deformation field based on the quality confidence level; Step S2223: Determine the soft field term of the local deformation field based on the mass confidence level; Step S2224: Determine the data constraint information of the local deformation field based on the soft matching term and the soft field term.

10. The fusion processing method according to claim 9, characterized in that, Step S223 includes: Step S2231: Determine the smoothing term of the local deformation field by determining the degree of curvature of the local deformation field; Step S2232: Determine the stiffness term of the local deformation field by determining the degree to which the local deformation deviates from the rigid body transformation; Step S2233: Determine the regular constraint information of the local deformation field based on the smoothing term and the rigidity term.

11. The fusion processing method according to claim 10, characterized in that, Step S223 further includes: Step S2234: Determine the topological barrier term of the local deformation field by determining the rate of change of the local area before and after deformation; Step S2233 includes: The regular constraint information of the local deformation field is determined based on the smoothing term, the rigidity term, and the topological barrier term.

12. The fusion processing method according to claim 10, characterized in that, Step S224 includes: Step S2241: Extract the edge set, line segment set, and corner point set from the first two-dimensional grayscale image and the second two-dimensional grayscale image, respectively; Step S2242: Calculate the edge field, line segment field, and corner field based on the extracted edge set, line segment set, and corner point set; Step S2243: Calculate the edge soft penalty term, the line segment soft penalty term, and the corner soft penalty term based on the edge field, the line segment field, and the corner field, respectively. Step S2244: Determine the main direction consistency term based on the main directions of the local regions in the first two-dimensional grayscale image and the second two-dimensional grayscale image; Step S2245: Determine the structural constraint information of the local deformation field based on the edge soft penalty term, the line segment soft penalty term, the corner soft penalty term, and the main direction consistency term.

13. The fusion processing method according to claim 12, characterized in that, Step S225 includes: Step S2251: Construct the overall objective function of the local deformation field based on the data constraint information, the regularization constraint information, and the structural constraint information; Step S2252: A multi-scale optimization strategy is adopted to optimize the parameters of the local deformation field in order to determine the optimal parameters of the local deformation field.

14. The fusion processing method according to claim 13, characterized in that, Step S2252 includes: Based on the soft field term and the smoothing term, the parameters of the local deformation field are coarsely optimized; Based on the edge soft penalty term, the line segment soft penalty term, the corner soft penalty term, and the rigidity term, the parameters of the local deformation field are optimized at a mid-level. Based on the soft matching term, the parameters of the local deformation field are optimized in a fine-grained manner.

15. The fusion processing method according to claim 14, characterized in that, In step S2252, the parameters of the local deformation field are optimized according to the following method: The parameters of the local deformation field are optimized using the Gauss-Newton or L-BFGS algorithm.

16. The fusion processing method according to claim 4, characterized in that, Step S3 includes: Step S31: Perform a global transformation on the first two-dimensional color image according to the global homography transformation matrix to obtain a third two-dimensional color image; Step S32: Based on the local deformation field, perform local deformation on the third two-dimensional color image to obtain a second two-dimensional color image.

17. The fusion processing method according to claim 16, characterized in that, Also includes: Step S5: Determine the registration confidence map based on the set of interior point pairs; Step S5 includes: Step S51: Calculate the matching density, local residual, and texture energy of each local region based on the set of interior point pairs. Step S52: Calculate the confidence score for each local region based on the matching density and / or local residual and / or texture energy and / or soft field term consistency of each local region. Step S53: Determine the registration confidence map based on the confidence score of each local region.

18. The fusion processing method according to claim 17, characterized in that, Step S32 includes: Step S321: Based on the local deformation field, perform local deformation on each local region of the third two-dimensional color image; Step S322: Based on the registration confidence map, local areas with confidence scores less than a set threshold are identified as backtracking regions. For backtracking regions, the results of global homography are used to replace the results of local deformation to obtain a second two-dimensional color image.

19. The fusion processing method according to claim 17, characterized in that, Step S4 includes: Step S41: Detect the geometric boundary region of the scene on the depth image; Step S42: Determine the color assignment weight map based on the geometric boundary region, the registration confidence map, and the echo quality of the lidar sensor; Step S43: For each pixel to be colored in the depth image, perform the following steps to generate a color depth image: If its color assignment weight is greater than 0, then the color value of the corresponding pixel in the second two-dimensional color image is used as its initial color value. If its color assignment weight is 0, then its color value is recorded as empty.

20. The fusion processing method according to claim 19, characterized in that, Step S43 includes: For pixels whose color assignment weight is greater than or equal to the preset color assignment weight threshold, their initial color value is directly used as the final color value. For pixels whose color assignment weight is less than the preset color assignment weight threshold, the initial color value of the pixel to be assigned is fused with the initial color values ​​of other pixels in the neighboring area based on their respective color assignment weights, and the fused color value is used as the final color value. For pixels with missing color values, the initial color value of one other pixel in the neighboring region with a color assignment weight greater than the preset color assignment weight threshold is used, or the initial color values ​​of at least two other pixels in the neighboring region are fused based on their respective color assignment weights, and the fused color value is used as the final color value.

21. The fusion processing method according to claim 19, characterized in that, Step S42 includes: Step S421: Based on the geometric boundary region, the registration confidence map, and the echo quality of the lidar sensor, determine the registration confidence weight, edge suppression weight, and echo quality weight for each pixel, respectively. Step S422: Normalize the registration confidence weight, the edge suppression weight, and the echo quality weight respectively, and multiply the normalized registration confidence weight, edge suppression weight, and echo quality weight to obtain the color assignment weight for each pixel.

22. The fusion processing method according to claim 19, characterized in that, Step S4 further includes: Step S44: Filter the color depth image.

23. The fusion processing method according to claim 1, characterized in that, Following step S4, the method further includes: Each pixel of the color depth image is back-projected into a three-dimensional space, and its three-dimensional coordinates are determined to generate a 3D image; The color value of each pixel in the color depth image is assigned to the corresponding pixel in the 3D image to generate a color point cloud map.

24. The fusion processing method according to claim 23, characterized in that, Also includes: The color point cloud image is subjected to additional processing, wherein the additional processing includes at least one of the following: outlier noise removal processing; Color consistency optimization processing.

25. A computer device comprising an image sensor, a lidar sensor, a processor, and a memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the fusion processing method according to any one of claims 1-24.