Visible light image and depth image fusion method and device and computer readable storage medium

By combining camera parameters to calculate the image conversion matrix and dedistortion mapping coordinates, converting and filling the image pixel coordinates, the complex and cost-effective fusion of visible light images and depth images in the prior art is solved, and a fast, low-cost and efficient image fusion effect is achieved.

CN120047321APending Publication Date: 2025-05-27CHENGDU BOYN TIANFU SOFTWARE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411990786.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art requires a complex image learning training process in the fusion of visible light images and depth images, with high hardware computing power, high usage cost, and low fusion computing efficiency.

Method used

By combining the pre-calibrated camera parameters of the visible light camera and the depth camera, the image conversion matrix and dedistortion mapping coordinates are calculated, the depth distance data of the original depth image is converted into depth coordinates, and the visible pixel coordinates are calculated based on the image conversion matrix, distortion processing and bilinear interpolation filling are performed, and the target fusion image is finally generated.

Benefits of technology

Fast and low-cost fusion of visible light images and depth images is achieved, improving image fusion efficiency, reducing usage costs, and retaining the resolution and detailed information of the original image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047321A_ABST
    Figure CN120047321A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a fusion method and device of a visible light image and a depth image, and a computer readable storage medium. The method comprises the following steps: obtaining an original visible light image and an original depth image which are respectively obtained by shooting the same area by a relatively fixed visible light camera and a depth camera; calculating to obtain an image conversion matrix of the visible light camera and the depth camera and a distortion-removing mapping coordinate of the original depth image; calculating corresponding visible light pixel coordinates by combining the depth coordinates of the original depth image with the distortion-removing mapping coordinates; distortion processing is carried out on the visible light pixel coordinates to obtain distorted visible light pixel coordinates; filling a depth value corresponding to the distorted visible light pixel coordinate into a corresponding pixel point in the original visible light image to generate an initial fusion image; and carrying out bilinear interpolation filling on pixel points lacking depth values in the initial fusion image, and finally obtaining a target fusion image. According to the embodiment, fusion of the visible light image and the depth image can be quickly realized with low cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of computer image processing, and in particular, to a method, device, and computer-readable storage medium for fusing visible light images and depth images. Background Art

[0002] With the popularization of technologies such as virtual reality, intelligent cockpits, and autonomous driving, and the improvement of chip computing power, the demand for multi-dimensional sensors is increasing day by day. Among them, fused images with unified resolution and field of view obtained through fusion algorithms have wide applications in the above fields. Currently, fused images obtained by fusing visible light images and depth images are the most widely used.

[0003] An existing method for fusing visible light images and depth images mainly involves collecting a large amount of raw data based on current fixed hardware, constructing a corresponding deep learning model, and then enabling the model to learn the internal correspondence between the depth map and the visible light image through multiple rounds of training. Finally, an algorithm model obtained by deep learning is used for image fusion.

[0004] However, the inventors found during specific implementation that the above method requires a complex image learning and training process, has high hardware computing power requirements, high usage costs, and relatively low fusion calculation efficiency. Summary of the Invention

[0005] The technical problem to be solved by the embodiments of the present invention is to provide a method for fusing visible light images and depth images, which can quickly and low-costly fuse visible light images and depth images.

[0006] The further technical problem to be solved by the embodiments of the present invention is to provide a device for fusing visible light images and depth images, which can quickly and low-costly fuse visible light images and depth images.

[0007] The further technical problem to be solved by the embodiments of the present invention is to provide a computer-readable storage medium for storing a computer program that can quickly and low-costly fuse visible light images and depth images.

[0008] To solve the above technical problems, the embodiments of the present invention first provide the following technical solution: A method for fusing visible light images and depth images, comprising the following steps: Obtain the original visible light image and the original depth image respectively obtained by a relatively fixed visible light camera and a depth camera shooting the same area; Calculate the image conversion matrix based on the first internal parameter and the first external parameter pre-calibrated by the depth camera, the second internal parameter and the second external parameter pre-calibrated by the visible light camera, and the external parameter pre-calibrated between the visible light camera and the depth camera, and calculate the undistorted mapping coordinates based on the first internal parameter and the first distortion coefficient pre-calibrated by the depth camera; Convert the depth distance data of the original depth image into depth coordinates, and calculate the corresponding visible light pixel coordinates by combining the undistorted mapping coordinates and the image conversion matrix; Perform distortion processing on the visible light pixel coordinates according to the second distortion coefficient pre-calibrated by the visible light camera to obtain distorted visible light pixel coordinates; Check whether the distorted visible light pixel coordinates are within the valid area. For the distorted visible light pixel coordinates located within the valid area, fill the depth value corresponding to the distorted visible light pixel coordinates into the corresponding pixel points in the original visible light image to generate an initial fused image; and Combine the relationship of the pixels around the corresponding coordinates of the original visible light image, and perform bilinear interpolation filling on the pixel points lacking depth values in the initial fused image to finally obtain the target fused image.

[0009] Further, the specific steps of converting the depth distance data of the original depth image into depth coordinates and calculating the corresponding visible light pixel coordinates by combining the undistorted mapping coordinates and the image conversion matrix include: Convert the depth distance data corresponding to each pixel point of the original depth image into depth coordinates; Combine the depth coordinates of each pixel point in the original depth image with the undistorted mapping coordinates one by one to form combined depth coordinates; and Convert the combined depth coordinates into visible light pixel coordinates based on the image conversion matrix.

[0010] Further, by checking whether there is a corresponding pixel point in the original visible light image for each distorted visible light pixel coordinate, if it exists, it is determined that the corresponding distorted visible light pixel coordinate is within the valid area.

[0011] Further, the specific meaning of converting the depth distance data corresponding to each pixel point of the original depth image into depth coordinates is to convert the depth distance data corresponding to each pixel point in the original depth image into depth coordinates with a unified unit.

[0012] Further, the Zhang-Zhengyou calibration algorithm model is used to pre-calculate the first internal parameter of the depth camera, the pre-calibrated second internal parameter and second external parameter of the visible light camera, and the pre-calibrated external parameter between the visible light camera and the depth camera through image calibration of the visible light camera and the depth camera.

[0013] On the other hand, to solve the above technical problems, an embodiment of the present invention further provides the following technical solution: A fusion device for visible light images and depth images, which is respectively connected to a relatively fixed visible light camera and a depth camera. The device includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the fusion method of visible light images and depth images as described in any one of the above.

[0014] On another aspect, to solve the above technical problems, an embodiment of the present invention further provides the following technical solution: A computer-readable storage medium, the computer-readable storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the fusion method of visible light images and depth images as described in any one of the above.

[0015] After adopting the above technical solutions, the embodiments of the present invention have at least the following beneficial effects: In the embodiments of the present invention, by first combining relevant camera parameters (one or more of internal parameters, external parameters, and distortion coefficients, etc.) pre-calibrated by the visible light camera and the depth camera, the image conversion matrix and the undistorted mapping coordinates between the visible light camera and the depth camera are respectively calculated. Thus, after converting the depth distance data of the original depth image into depth coordinates, the visible light pixel coordinates can be calculated by combining the undistorted mapping coordinates and the image conversion matrix. To ensure that the visible light pixel coordinates can correspond to the pixel coordinates of each pixel point in the original visible light image that has not undergone distortion correction, the visible light pixel coordinates are then distorted to obtain distorted visible light pixel coordinates. Then, based on the correspondence between the distorted visible light pixel coordinates and the pixel coordinates of each pixel point in the original visible light image, the depth value in each distorted visible light pixel coordinate is filled into the corresponding pixel point in the original visible light image to generate an initial fusion image. Finally, by combining the relationship of the pixels around the corresponding coordinates of the original visible light image, the pixel points lacking depth values in the initial fusion image are filled by bilinear interpolation, and the target fusion image can be obtained. The depth values of each pixel point in the target fusion image are highly credible, and the target fusion image retains the original resolution of the original visible light image, retains more image details and texture information, and also avoids the problem of information loss caused by introducing a large number of holes; in addition, the image fusion efficiency is higher, and there is no need for complex image learning and training, reducing the use cost. Description of the Drawings

[0016] Figure 1 This is a flowchart of the steps of an optional embodiment of the method for fusing visible light images and depth images according to the present invention.

[0017] Figure 2 This is a specific flowchart of step S3 in an optional embodiment of the method for fusing visible light images and depth images according to the present invention.

[0018] Figure 3 This is a schematic block diagram of an optional embodiment of the device for fusing visible light images and depth images according to the present invention.

[0019] Figure 4 This is a functional module diagram of an optional embodiment of the device for fusing visible light images and depth images according to the present invention. Detailed implementation manners

[0020] The following further elaborates on the present application in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following illustrative embodiments and explanations are only used to explain the present invention and are not intended to limit the present invention. Moreover, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0021] As Figure 1 shown, an optional embodiment of the present invention provides a method for fusing visible light images and depth images, including the following steps: S1: Obtain the original visible light image and the original depth image respectively acquired by a relatively fixed visible light camera 1 and a depth camera 2 shooting the same area; S2: Calculate and obtain an image transformation matrix by combining the first internal parameters and the first external parameters pre-calibrated by the depth camera 2, the second internal parameters and the second external parameters pre-calibrated by the visible light camera 1, and the external parameters pre-calibrated between the visible light camera 1 and the depth camera 2, and calculate and obtain a de-distortion mapping coordinate based on the first internal parameters and the first distortion coefficient pre-calibrated by the depth camera 2; S3: Convert the depth distance data of the original depth image into depth coordinates and calculate the corresponding visible light pixel coordinates by combining the de-distortion mapping coordinate and the image transformation matrix; S4: Perform distortion processing on the visible light pixel coordinates according to the second distortion coefficient pre-calibrated by the visible light camera 1 to obtain distorted visible light pixel coordinates; S5: Check whether the distorted visible light pixel coordinates are within the valid area. For the distorted visible light pixel coordinates within the valid area, fill the depth value corresponding to the distorted visible light pixel coordinates into the corresponding pixel points in the original visible light image to generate an initial fused image; and S6: Combine the relationships of the pixels around the corresponding coordinates of the original visible light image, perform bilinear interpolation filling on the pixel points lacking depth values in the initial fusion image, and finally obtain the target fusion image.

[0022] In the embodiment of the present invention, by first combining the relevant camera parameters (one or more of internal parameters, external parameters, distortion coefficients, etc.) pre-calibrated by the visible light camera 1 and the depth camera 2, the image conversion matrix and the undistorted mapping coordinates between the visible light camera 1 and the depth camera 2 are respectively calculated. Thus, after converting the depth distance data of the original depth image into depth coordinates, the visible light pixel coordinates can be calculated by combining the undistorted mapping coordinates and the image conversion matrix. To ensure that the visible light pixel coordinates can correspond to the pixel coordinates of each pixel point in the original visible light image that has not undergone distortion correction, the visible light pixel coordinates are then distorted to obtain distorted visible light pixel coordinates. Then, based on the correspondence between the distorted visible light pixel coordinates and the pixel coordinates of each pixel point in the original visible light image, the depth value in each distorted visible light pixel coordinate is filled into the corresponding pixel point in the original visible light image to generate the initial fusion image. Finally, by combining the relationships of the pixels around the corresponding coordinates of the original visible light image, bilinear interpolation filling is performed on the pixel points lacking depth values in the initial fusion image, and the target fusion image can be obtained. The credibility of the depth values of each pixel point in the target fusion image is high, and the target fusion image retains the original resolution of the original visible light image, retains more image details and texture information, and also avoids the problem of information loss caused by introducing a large number of holes; in addition, the image fusion efficiency is higher, and there is no need for complex image learning and training, reducing the usage cost.

[0023] Specifically, in step S2, the pixel coordinates of each pixel point in the original depth image can be expressed as {(u 1 ', v 1 ')...(u n ', v n ')}, and the undistorted mapping coordinates can be expressed as {(u 1 , v 1 )...(u n , v n )}; in step S3, the depth coordinate is the z coordinate, and the combined depth coordinate is expressed as: {(u 1 , v 1 , z 1 )...(u n , v n , z n )}.

[0024] In an alternative embodiment of the present invention, as Figure 2 shown, step S3 specifically includes: S31: Convert the depth distance data corresponding to each pixel point of the original depth image into depth coordinates; S32: Combine the depth coordinates of each pixel point in the original depth image with the undistorted mapping coordinates one by one to form combined depth coordinates; and S33: Convert the combined depth coordinates into visible light pixel coordinates based on the image transformation matrix.

[0025] In this embodiment, by first combining the depth coordinates of each pixel point in the original depth image with the undistorted mapping coordinates one by one to form combined depth coordinates, and then performing a one-time coordinate transformation on the combined depth coordinates based on the image transformation matrix, the data processing efficiency is higher.

[0026] In an alternative embodiment of the present invention, by checking whether there is a corresponding pixel point in the original visible light image for each of the distorted visible light pixel coordinates, if there is, it is determined that the corresponding distorted visible light pixel coordinate is within the valid area. Since there will be a certain difference in the shooting fields of view of the visible light camera 1 and the depth camera 2, the images captured by the two cannot completely overlap. Some distorted visible light pixel coordinates generated based on the depth image may be outside the original visible light image. Therefore, in this embodiment, corresponding inspection and judgment are usually performed first to filter out the distorted visible light pixel coordinates within the valid area, thereby reducing the data processing volume and improving the processing efficiency.

[0027] In an alternative embodiment of the present invention, the above step S31 specifically refers to converting the depth distance data corresponding to each pixel point in the original depth image into depth coordinates in a unified unit. In this embodiment, the pixel value of each pixel point in the original depth image represents the distance from the depth sensor of the depth camera to the corresponding point in the actual scene, generally represented in 16-bit or 32-bit pixel units. After unifying it into the same unit (for example: meters, centimeters, etc.), it is convenient for subsequent calculation and processing.

[0028] In an alternative embodiment of the present invention, the Zhang Zhengyou calibration algorithm model is used to pre-calculate the first internal parameter of the depth camera 2, the second internal parameter and the second external parameter of the pre-calibrated visible light camera 1, and the external parameter pre-calibrated between the visible light camera 1 and the depth camera 2 through image calibration of the visible light camera 1 and the depth camera 2. In this embodiment, the Zhang Zhengyou calibration algorithm model is a common camera parameter calibration algorithm model, which can quickly calibrate the above camera parameters.

[0029] Specifically, based on the camera imaging principle, the first internal parameter of the depth camera 2 is expressed as: (Formula 1) where, (xL , y L , z L ) represents a point in the camera coordinate system, (u L , v L ) represents a point in the image coordinate system.

[0030] The conversion relationship from the image coordinate system of the depth camera 2 to the camera coordinate system is expressed as: (Formula 2) The conversion from the image coordinate system of the visible light camera 1 to the camera coordinate system is expressed as: (Formula 3) Among them, M is the transformation matrix.

[0031] Substituting Formula 1 and Formula 2 into Formula 3, the image transformation matrix K can be obtained as: (Formula 4) The obtained image transformation matrix K above can be used for subsequent image data processing.

[0032] On the other hand, as Figure 3 shown, the embodiment of the present invention also provides a fusion device 3 for visible light images and depth images, which is respectively connected to the relatively fixed visible light camera 1 and depth camera 2. The device 3 includes a processor 30, a memory 32, and a computer program stored in the memory 32 and configured to be executed by the processor 30. When the processor 30 executes the computer program, it implements the fusion method of visible light images and depth images as described in any of the above embodiments.

[0033] Exemplarily, the computer program can be divided into one or more modules / units. The one or more modules / units are stored in the memory 32 and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program in the fusion device 3 of visible light images and depth images. For example, the computer program can be divided into Figure 4 the functional modules in the fusion device 3 of visible light images and depth images. Among them, the image acquisition module 41, the coordinate calculation module 42, the coordinate conversion module 43, the coordinate distortion processing module 44, the depth value filling module 45, and the depth value filling module 46 respectively execute the above steps S1 - step S6.

[0034] The visible light image and depth image fusion device 3 can be a computing device such as a desktop computer, notebook, palm computer, or cloud server. The visible light image and depth image fusion device 3 may include, but is not limited to, a processor 30 and a memory 32. Those skilled in the art can understand that the schematic diagram is only an example of the visible light image and depth image fusion device 3, and does not constitute a limitation on the visible light image and depth image fusion device 3. It may include more or fewer components than shown, or combine certain components, or different components. For example, the visible light image and depth image fusion device 3 may also include input / output devices, network access devices, buses, etc.

[0035] The processor 30 can be a central processing unit (CPU), or it can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The processor 30 is the control center of the visible light image and depth image fusion device 3, and connects all parts of the visible light image and depth image fusion device 3 through various interfaces and lines.

[0036] The memory 32 can be used to store the computer programs and / or modules. The processor 30 realizes various functions of the visible light image and depth image fusion device 3 by running or executing the computer programs and / or modules stored in the memory 32, and by calling the data stored in the memory 32. The memory 32 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a graphic recognition function, a graphic overlay function, etc.); the data storage area can store data created according to the use of the control device (such as graphic data, etc.). In addition, the memory 32 can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.

[0037] If the functions described in the embodiments of the present invention are implemented in the form of software function modules or units and sold or used as independent products, they can be stored in a storage medium readable by a computing device. Based on such an understanding, to implement all or part of the processes in the above-described method embodiments of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor 30, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0038] In another aspect, the embodiments of the present invention further provide a computer-readable storage medium. The computer-readable storage medium includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the fusion method of visible light images and depth images as described in any of the above embodiments.

[0039] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.

[0040] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the present invention's claims. All of these fall within the protection scope of the present invention.

Claims

1. A method for fusing a visible light image and a depth image, characterized in that: The method comprises the following steps: Obtaining an original visible light image and an original depth image respectively obtained by shooting the same area by a relatively fixed visible light camera and a depth camera; An image conversion matrix is ​​calculated by combining a first intrinsic parameter and a first extrinsic parameter pre-calibrated by the depth camera, a second intrinsic parameter and a second extrinsic parameter pre-calibrated by the visible light camera, and an external parameter pre-calibrated between the visible light camera and the depth camera, and a dedistorted mapping coordinate is calculated based on the first intrinsic parameter and a first distortion coefficient pre-calibrated by the depth camera; Convert the depth distance data of the original depth image into depth coordinates and calculate the corresponding visible light pixel coordinates by combining the dedistortion mapping coordinates and the image conversion matrix; Performing distortion processing on the visible light pixel coordinates according to a second distortion coefficient pre-calibrated by the visible light camera to obtain distorted visible light pixel coordinates; Checking whether the distorted visible light pixel coordinates are within a valid area, and for the distorted visible light pixel coordinates within the valid area, filling the depth values ​​corresponding to the distorted visible light pixel coordinates into corresponding pixel points in the original visible light image to generate an initial fused image; and Combined with the relationship between the pixels around the corresponding coordinates of the original visible light image, bilinear interpolation is performed to fill the pixels lacking depth values ​​in the initial fused image, and finally a target fused image is obtained.

2. The method for fusing a visible light image and a depth image according to claim 1, wherein: The converting the depth distance data of the original depth image into depth coordinates and calculating the corresponding visible light pixel coordinates in combination with the dedistortion mapping coordinates and the image conversion matrix specifically includes: Convert the depth distance data corresponding to each pixel point of the original depth image into depth coordinates; Combining the depth coordinate of each pixel in the original depth image with the dedistortion mapping coordinate in a one-to-one correspondence to form a combined depth coordinate; and The combined depth coordinates are converted to visible light pixel coordinates based on the image conversion matrix.

3. The method for fusing a visible light image and a depth image according to claim 1, wherein: Check whether each of the distorted visible light pixel coordinates has a corresponding pixel point in the original visible light image, and if so, determine whether the corresponding distorted visible light pixel coordinate is located in the valid area.

4. The method for fusing a visible light image and a depth image according to claim 2, wherein: The converting the depth distance data corresponding to each pixel point of the original depth image into depth coordinates specifically refers to converting the depth distance data corresponding to each pixel point in the original depth image into depth coordinates of a unified unit.

5. The method for fusing a visible light image and a depth image according to claim 1, wherein: The Zhang Zhengyou calibration algorithm model is used to pre-calculate the visible light camera and the depth camera through image calibration to obtain a first intrinsic parameter of the depth camera, a second intrinsic parameter and a second extrinsic parameter pre-calibrated by the visible light camera, and an external parameter pre-calibrated between the visible light camera and the depth camera.

6. A visible light image and depth image fusion device, connected to a relatively fixed visible light camera and a depth camera respectively, characterized in that: The device includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, the method for fusing a visible light image and a depth image as described in any one of claims 1 to 5 is implemented.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the method for fusing a visible light image and a depth image as described in any one of claims 1 to 5.