Single photon and visible light image fusion method and system based on multi-scale Markov random field model
By fusing single-photon and visible light images using a multi-scale Markov random field model, the problem of low point cloud resolution in single-photon lidar is solved, improving the resolution of long-range imaging and target recognition capabilities.
Patent Information
- Application Number
- CN202511618798.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-06
- Publication Date
- 2026-02-10
AI Technical Summary
Single-photon lidar has low point cloud resolution and lacks edge detail information. The resolution drops significantly, especially when imaging at long distances, which affects the target recognition capability.
A multi-scale Markov random field model is adopted, and images are fused by array single-photon lidar and visible light high-speed camera. Taking advantage of the high resolution of visible light intensity images, the multi-scale Markov random field model is combined to perform image fusion and optimize pixel values to improve resolution.
It effectively improves the image resolution and target recognition capability of single-photon lidar, enhances the image fusion effect in complex scenes, and avoids global structural deviation and local overfitting at a single scale.
Smart Images

Figure CN121504741A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of multi-modal image fusion, and particularly relates to a single-photon and visible light image fusion method and system based on a multi-scale Markov random field model. BACKGROUND
[0002] The array single-photon laser radar based on the photon time-of-flight method and the time-correlated single-photon counting technology can record the distance information and the reflection intensity information of a scene target, but the limitation of the number of pixels makes the three-dimensional point cloud resolution detected by the single-photon laser radar low, and the target texture detail information is lacking, especially as the detection distance increases, the quality of the point cloud map accumulated by the system becomes worse and worse, for example, for a common 64*64 array detector, the size of a pixel is usually 50um, assuming that a receiving lens with a focal length of 100mm is used , in an ideal state, the lateral resolution thereof is 3.33mm when imaging at a distance of 100m, and the lateral resolution thereof is 33.33mm when imaging at a distance of 1km, in actual application, the system is affected by environmental factors, resulting in sparse point cloud, increased noise, reduced edge definition of the target, and reduced effective lateral resolution, which seriously limits the sensing and recognition ability of the single-photon laser radar on scene information. Therefore, the most essential method to solve the low point cloud resolution of the array single-photon laser radar is to improve the number of pixels of the hardware, but the manufacturing of a high-density SPAD array has many limitations, such as small-area integration of the CMOS readout circuit, suppression of interference (crosstalk) between pixels, and the like, in summary, the method of improving the point cloud resolution from the hardware aspect has a very high manufacturing process cost.
[0003] The point cloud map information of the single-photon laser radar is sparse, so the method of using the point cloud map information itself to obtain high resolution is greatly affected by noise, the edge contour details are difficult to recover, and the effect is unstable. SUMMARY
[0004] In view of the above technical problems, the application provides a single-photon and visible light image fusion method and system based on a multi-scale Markov random field model, to solve the problems of sparse point cloud, low resolution, and lack of edge detail information of the single-photon laser radar in the prior art.
[0005] The first aspect of the application discloses a single-photon and visible light image fusion method based on a multi-scale Markov random field model, the method comprising: S1, detecting a target at different distances by an array single-photon laser radar system to obtain a return signal; S2, pre-processing the return signal to obtain a single-photon point cloud image after frame number accumulation; S3 utilizes a dichroic mirror to perform coaxial processing on the single-photon detectors in the visible light high-speed camera and array single-photon lidar system in order to achieve image registration. S4, acquires a visible light intensity image of the same target using a high-speed visible light camera; S5 unfolds visible light in the image at multiple resolution scales, constructs an independent Markov random field model at each scale, and uses the distance, contour texture information and coarse-scale fusion results of the single-photon point cloud image and the visible light intensity image as constraints and priors. Multi-scale joint optimization is performed by minimizing the energy function to obtain a high-resolution image.
[0006] Optionally, in step S1, the pulse width and pulse repetition frequency of the laser in the array single-photon lidar system are determined according to the required ranging distance.
[0007] Optionally, step S2 specifically includes: Time-correlated single-photon counting is performed on the arrival timestamps of photons detected multiple times in each pixel unit of the single-photon detector to generate a photon count-time delay histogram; background noise suppression and signal peak detection are performed on the photon count-time delay histogram to extract the flight time corresponding to the target echo signal; pixel distance values are calculated based on the flight time, and the final low-resolution single-photon depth image is generated by multi-frame accumulation and fusion.
[0008] Optionally, step S5 specifically includes: S51, Gaussian pyramid downsampling is performed on the visible light image to generate multi-resolution scale levels; S52, at each resolution scale, models each pixel of the single-photon image as a node of a Markov random field and defines neighborhood relations to model spatial dependencies. S53 defines the energy function at each layer resolution scale; the energy function includes a data term and a smoothing term. S54 defines a multi-scale global energy function, optimizes the energy function at each resolution scale and the multi-scale global energy function, iteratively updates the pixel values of the image, and gradually merges the information in the two images. S55 performs post-processing on the fused image and evaluates it using image quality metrics.
[0009] Optionally, before step S51, the following steps are also included: Image alignment and pixel normalization are performed on single-photon point cloud images and visible light intensity images.
[0010] Optionally, in step S55, the post-processing steps include non-local mean denoising and edge sharpening to improve image quality.
[0011] Optionally, in step S55, the image quality evaluation metrics include: peak signal-to-noise ratio, structural similarity, and learned perceptual image patch similarity.
[0012] A second aspect of this invention discloses a single-photon and visible light image fusion system based on a multi-scale Markov random field model, the system comprising: The first processing module is configured to detect targets at different distances using an array single-photon lidar system and acquire echo signals. The second processing module is configured to preprocess the echo signal to obtain a single-photon point cloud image after frame accumulation. The third processing module is configured to use a dichroic mirror to perform coaxial processing on the single-photon detector in the visible light high-speed camera and the array single-photon lidar system in order to achieve image registration. The fourth processing module is configured to acquire visible light intensity images of the same target through a high-speed visible light camera. The fifth processing module is configured to unfold the visible light image at multiple resolution scales, construct an independent Markov random field model at each scale, and use the distance, contour texture information of the single-photon point cloud image and the visible light intensity image, as well as the fusion result at the coarse scale, as constraints and priors. It then performs multi-scale joint optimization by minimizing the energy function to fuse and obtain a high-resolution image.
[0013] A third aspect of this invention discloses an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps in the single-photon and visible light image fusion method based on a multi-scale Markov random field model described in the first aspect of this invention.
[0014] A fourth aspect of this invention discloses a computer-readable storage medium. The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the single-photon and visible light image fusion method based on a multi-scale Markov random field model described in the first aspect of this invention.
[0015] In summary, the proposed solution of this invention has the following technical effects: In the process of ranging and imaging of long-distance targets, this invention uses a multi-sensor fusion method, taking advantage of the high resolution of visible light intensity images, to effectively solve the problem of low point cloud density and low image resolution of single-photon lidar, and effectively enhances the perception capability of single-photon lidar for scene target information; the fusion method based on a multi-scale Markov random field model enhances structural consistency and anti-interference capability, and effectively avoids global structural deviation or local overfitting under a single scale, making it suitable for image fusion in multimodal and complex scenes. Attached Figure Description
[0016] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of the process of the single-photon and visible light image fusion method based on the Markov random field model according to an embodiment of the present invention; Figure 2 This is a structural diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] Markov random field models can model the spatial consistency between image pixels or regions in image fusion. By combining complementary information from two source images, they can enhance the structural details of single-photon images and improve their resolution. However, single-scale Markov random field models lack the ability to model the multi-level structure of images, making it difficult to balance details and global structure in the fusion results.
[0020] Based on the above analysis, it is extremely important to propose a suitable method to improve the point cloud resolution of single-photon lidar in order to achieve long-range high-resolution ranging and imaging of single-photon lidar.
[0021] The first aspect of this invention discloses a method for fusing single-photon and visible light images based on a multi-scale Markov random field model, the method comprising: S1 detects targets at different distances using an array single-photon lidar system and acquires echo signals; Optionally, in step S1, the pulse width and pulse repetition frequency of the laser in the array single-photon lidar system are determined according to the required ranging distance.
[0022] When the pulse width of the laser is The pulse repetition frequency is At that time, the maximum unambiguous range of a single-photon lidar is: Depth resolution is It can detect targets at different distances and obtain echo signals; S2, preprocess the echo signal to obtain a single-photon point cloud image after frame accumulation; When the time difference between the transmitted signal and the received echo is At that time, the distance to the target is Based on the time-of-flight method and time-correlated single-photon counting, the echo signal is preprocessed to obtain a low-resolution single-photon image after frame accumulation. Optionally, step S2 specifically includes: Time-correlated single-photon counting is performed on the arrival timestamps of photons detected multiple times in each pixel unit of the single-photon detector to generate a photon count-time delay histogram; background noise suppression and signal peak detection are performed on the photon count-time delay histogram to extract the flight time corresponding to the target echo signal; pixel distance values are calculated based on the flight time, and the final low-resolution single-photon depth image is generated by multi-frame accumulation and fusion.
[0023] S3 utilizes a dichroic mirror to perform coaxial processing on the single-photon detectors in the visible light high-speed camera and array single-photon lidar system in order to achieve image registration. S4, acquires a visible light intensity image of the same target using a high-speed visible light camera; S5 unfolds visible light in the image at multiple resolution scales, constructs an independent Markov random field model at each scale, and uses the distance, contour texture information and coarse-scale fusion results of the single-photon point cloud image and the visible light intensity image as constraints and priors. Multi-scale joint optimization is performed by minimizing the energy function to obtain a high-resolution image.
[0024] Optionally, step S5 specifically includes: S51 performs Gaussian pyramid downsampling on visible light in the image to generate multi-resolution scale levels; Optionally, before step S51, the following steps are also included: Image alignment and pixel normalization are performed on single-photon point cloud images and visible light intensity images to better adapt to the subsequent fusion process.
[0025] If a single-photon point cloud image (low-resolution image) contains little information after multi-scale decomposition, the method of adding filters can be used to try to recover the details in the low-resolution image; S52, at each resolution scale, models each pixel of the single-photon image as a node of a Markov random field and defines neighborhood relations to model spatial dependencies. S53 defines the energy function at each resolution scale; the energy function is essentially to minimize the difference from the real high-resolution image, and the energy function includes a data term and a regularization term. S54 defines a multi-scale global energy function, optimizes the energy function at each resolution scale and the multi-scale global energy function, iteratively updates the pixel values of the image, and gradually merges the information in the two images. When optimizing the energy function, it is necessary to consider both the data fidelity of the low-resolution image to ensure that the fusion result matches the observation value, and to adjust the guiding intensity of the high-resolution image to avoid overfitting.
[0026] S55 performs post-processing on the fused image and evaluates it using image quality metrics.
[0027] Optionally, in step S55, the post-processing steps include non-local mean denoising and edge sharpening to improve image quality.
[0028] Optionally, in step S55, the image quality evaluation metrics include: peak signal-to-noise ratio, structural similarity, and learned perceptual image patch similarity.
[0029] Example of fusing single-photon lidar point cloud with visible light image: The purpose of this invention is to provide a method for fusing single-photon and visible light images based on a multi-scale Markov random field model. High-speed cameras capture high-resolution two-dimensional intensity maps of visible light, which can record color and texture details of the scene being captured. However, they lack depth information and are sensitive to light intensity, resulting in poor imaging quality under extreme conditions such as heavy fog and rainstorms. This invention utilizes image fusion technology to integrate the high-resolution advantage of visible light images into array single-photon imaging, thereby solving the problems of low resolution and lack of confidence in target texture details in the three-dimensional point cloud of array single-photon lidar. The fused image provides a more comprehensive and clear description of the scene, facilitating subsequent target detection, recognition, and tasks. Specifically, for the single-photon and visible light image fusion method based on a multi-scale Markov random field model in this embodiment, please refer to [link to relevant documentation]. Figure 1 It includes the following steps: The first step is to select a laser with a suitable pulse width and a suitable repetition frequency based on the required ranging distance, and then use an array single-photon lidar system to perform ranging and imaging of targets at different distances and collect raw data.
[0030] The second step involves preprocessing the echo signals collected by the lidar based on the time-of-flight method and the time-correlated single-photon counting principle to obtain a single-photon point cloud map after frame accumulation. At this point, the point cloud map has low resolution and sparse information, which is not conducive to subsequent target recognition and other tasks.
[0031] The third step is to use a dichroic mirror to make the single-photon detector and the high-speed camera coaxial.
[0032] The fourth step involves selecting a dichroic mirror of appropriate wavelength for beam splitting. The reflected laser pulses are received by the receiving lens tube. After passing through the dichroic mirror, the received light is split into two paths. One path is the transmitted long-wavelength light, which enters the array single-photon camera for single-photon imaging after background noise is filtered out by a filter. The other path is the reflected short-wavelength light, which is received and imaged by a visible light camera. This allows for the acquisition of a high-resolution intensity map (visible light intensity image) of the same target using a high-speed visible light camera.
[0033] The fifth step involves obtaining multimodal images from two different sensors. The high-resolution visible light image is then unfolded across multiple resolution scales. Independent Markov random field models are constructed at each scale using the relationship between pixels in the single-photon image. The distance, contour texture information, and coarse-scale fusion results contained in the two images are used as constraints and prior information. Multi-scale joint optimization is performed using an optimization algorithm. The intensity image guides the fusion of the two images to obtain the final high-resolution image.
[0034] The following section explains in detail the method for constructing multi-scale Markov random field models: High-resolution visible light images are downsampled using a Gaussian pyramid to generate multi-resolution scale levels from low to high. Each pixel value in the low-resolution image is an observation value in the target high-resolution image. At each scale level, a Markov random field model is used to model the image, defining an energy function at each resolution scale. The energy function essentially measures the interrelationships between pixels and the consistency of image data. A multi-scale overall energy function is then defined, and optimization of the energy function is performed to achieve joint optimization across multiple scales. Pixel values are iteratively updated, and information from the two images is fused to obtain the final high-resolution single-photon lidar point cloud map. Finally, the fused image undergoes post-processing to further enhance detail and quality. The quality of the fused image is evaluated using various metrics, including Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Indicator (SSIM), and Learning Perceptual Patch Similarity Indicator (LPIPS).
[0035] Specifically, the energy function includes a data term and a smoothing term. The data term essentially measures the difference between the image fusion result and the original input image, while the smoothing term ensures spatial consistency in the image, meaning that the values of adjacent pixels change smoothly, avoiding unnatural edges or noise. Furthermore, the regularization term promotes similar values in similar neighboring pixels by penalizing differences between them.
[0036] In the optimization of the energy function, it is necessary to consider both the data fidelity of the low-resolution image to ensure that the fusion result matches the observation value and to adjust the guiding intensity of the high-resolution image to avoid overfitting.
[0037] Furthermore, the final fused image quality can be enhanced by processing methods such as nonlocal mean denoising and edge sharpening.
[0038] A third aspect of this invention discloses an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps in the single-photon and visible light image fusion method based on a multi-scale Markov random field model described in the first aspect of this invention.
[0039] Figure 2 This is a structural diagram of an electronic device according to an embodiment of the present invention, such as... Figure 2 As shown, the electronic device includes a processor, memory, communication interface, display screen, and input device connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, Near Field Communication (NFC), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input device can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse.
[0040] Those skilled in the art will understand that Figure 2 The structure shown is merely a structural diagram of the part related to the technical solution of this disclosure and does not constitute a limitation on the electronic device to which the solution of this application is applied. Specifically, the electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0041] A fourth aspect of this invention discloses a computer-readable storage medium. The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the single-photon and visible light image fusion method based on a multi-scale Markov random field model described in the first aspect of this invention.
[0042] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein, and such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for fusing single-photon and visible light images based on a multi-scale Markov random field model, characterized in that, The method includes: S1 detects targets at different distances using an array single-photon lidar system and acquires echo signals; S2, preprocess the echo signal to obtain a single-photon point cloud image after frame accumulation; S3 utilizes a dichroic mirror to perform coaxial processing on the single-photon detectors in the visible light high-speed camera and array single-photon lidar system in order to achieve image registration. S4, acquires a visible light intensity image of the same target using a high-speed visible light camera; S5 unfolds visible light in the image at multiple resolution scales, constructs an independent Markov random field model at each scale, and uses the distance, contour texture information and coarse-scale fusion results of the single-photon point cloud image and the visible light intensity image as constraints and priors. Multi-scale joint optimization is performed by minimizing the energy function to obtain a high-resolution image.
2. The method according to claim 1, characterized in that, In step S1, the pulse width and pulse repetition frequency of the laser in the array single-photon lidar system are determined according to the required ranging distance.
3. The method according to claim 1, characterized in that, Step S2 specifically includes: Time-correlated single-photon counting is performed on the arrival timestamps of photons detected multiple times in each pixel unit of the single-photon detector to generate a photon count-time delay histogram; background noise suppression and signal peak detection are performed on the photon count-time delay histogram to extract the flight time corresponding to the target echo signal; pixel distance values are calculated based on the flight time, and the final low-resolution single-photon depth image is generated by multi-frame accumulation and fusion.
4. The method according to claim 1, characterized in that, Step S5 specifically includes: S51, Gaussian pyramid downsampling is performed on the visible light image to generate multi-resolution scale levels; S52, at each resolution scale, models each pixel of the single-photon image as a node of a Markov random field and defines neighborhood relations to model spatial dependencies. S53 defines the energy function at each layer resolution scale; the energy function includes a data term and a smoothing term. S54 defines a multi-scale global energy function, optimizes the energy function at each resolution scale and the multi-scale global energy function, iteratively updates the pixel values of the image, and gradually merges the information in the two images. S55 performs post-processing on the fused image and evaluates it using image quality metrics.
5. The method according to claim 4, characterized in that, Before step S51, the following are also included: Image alignment and pixel normalization are performed on single-photon point cloud images and visible light intensity images.
6. The method according to claim 4, characterized in that, In step S55, the post-processing steps include non-local mean denoising and edge sharpening to improve image quality.
7. The method according to claim 4, characterized in that, In step S55, the image quality evaluation metrics include: peak signal-to-noise ratio, structural similarity, and learned perceptual image patch similarity.
8. A single-photon and visible light image fusion system based on a multi-scale Markov random field model, characterized in that, The system includes: The first processing module is configured to detect targets at different distances using an array single-photon lidar system and acquire echo signals. The second processing module is configured to preprocess the echo signal to obtain a single-photon point cloud image after frame accumulation. The third processing module is configured to use a dichroic mirror to perform coaxial processing on the single-photon detector in the visible light high-speed camera and the array single-photon lidar system in order to achieve image registration. The fourth processing module is configured to acquire visible light intensity images of the same target through a high-speed visible light camera. The fifth processing module is configured to unfold the visible light in the image at multiple resolution scales, construct an independent Markov random field model at each scale, and use the distance, contour texture information and coarse-scale fusion results of the single-photon point cloud image and the visible light intensity image as constraints and priors to perform multi-scale joint optimization by minimizing the energy function, and fuse them to obtain a high-resolution image.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, it implements the steps in the single-photon and visible light image fusion method based on a multi-scale Markov random field model according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the single-photon and visible light image fusion method based on a multi-scale Markov random field model according to any one of claims 1 to 7.
Citation Information
Cited By
Active-passive dual-mode single-photon laser radar imaging system and method
CN121741757A