Pixel-level infrared and visible light image pair generation method, device, equipment and medium

By synchronously acquiring point clouds and image sequences, constructing a 3D point cloud map, and performing matching and coloring, the problem of constructing large-scale, cross-time period pixel-level aligned infrared and visible light image pairs is solved. The generated image pairs are pixel-level aligned, and the resulting distribution is close to that of real images, thus achieving efficient dataset construction.

CN118736045BActive Publication Date: 2025-12-16NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410821678.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-24
Publication Date
2025-12-16
Estimated Expiration
2044-06-24

AI Technical Summary

Technical Problem

Constructing large-scale, cross-time-period pixel-aligned infrared and visible light images is difficult due to the limitations of existing methods, which suffer from viewpoint differences and dataset constraints. Existing methods struggle to generate satisfactory nighttime infrared image quality, and the datasets are small, limited in scope, and have low collection efficiency.

Method used

By simultaneously acquiring point cloud sequences and image sequences during the day and night, a 3D point cloud map is constructed. The radar pose is used to match infrared and visible light images to generate a dense depth map and perform coloring, achieving pixel-level alignment.

Benefits of technology

The generated images are pixel-level aligned, and the resulting distribution is close to that of real images. This efficiently constructs large-scale high-resolution datasets and solves the problems of modal differences and cross-time misalignment caused by the thermal imaging principle of infrared cameras.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118736045B_ABST
    Figure CN118736045B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of image pair generation, and relates to a pixel-level infrared and visible light image pair generation method, device, equipment and medium. The method comprises the following steps: synchronously acquiring a daytime point cloud sequence and a corresponding visible light image sequence, synchronously acquiring a nighttime point cloud sequence and a corresponding infrared image sequence; constructing a three-dimensional point cloud map, respectively matching the daytime and nighttime point cloud sequences to obtain radar poses, and matching the infrared image sequence to obtain a matched image pair; respectively projecting the daytime or nighttime point cloud sequence onto the matched image pair to obtain a visible light sparse depth map and an infrared sparse depth map, and respectively optimizing the visible light sparse depth map and the infrared sparse depth map to obtain a visible light dense depth map and an infrared dense depth map; coloring each pixel point on the infrared dense depth map according to the color of each pixel point on the visible light dense depth map to generate a pixel-level aligned image pair. The application can construct an infrared and visible light image pair dataset.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image pair generation, and particularly relates to a pixel-level infrared and visible light image pair generation method and device, equipment and a medium. BACKGROUND

[0002] In the fields of assisted driving, robots and monitoring, visible light cameras are usually used as the "eyes" of machines to provide key input information for accurate perception, decision-making and control. However, under insufficient light conditions, visible light images may lose the texture and structure information of the scene, affecting the performance of related algorithms. In this case, infrared cameras based on thermal imaging principles can stably provide relatively clear visual signals, but the differences in gray scale, appearance and content caused by their unique imaging spectral range usually do not conform to the human eye's perception of the real world.

[0003] In the early stage of research, people enhanced night vision visible light data through infrared image fusion, but due to the continuous emission and scattering of heat by objects and the environment, the "ghost effect" caused the infrared image to be unable to provide detailed semantic information. Later, some researchers turned to explore physical shading methods to achieve more direct perception, but the effect was still unsatisfactory.

[0004] At present, a feasible method is to use a generative model to realize the translation of infrared images to visible light images, which provides a solution for unpaired night infrared to daytime visible light image translation to some extent. For example, the style-level supervised generative model ROMA based on CycleGAN network converts unpaired night infrared video into fine-grained daytime visible light video by designing cross-domain region similarity matching and multi-scale cross-region discriminators. However, pixel-level supervised models can usually learn clearer conversion relationships from paired data sets, but this idea of carefully designing structures for each specific task is difficult to expand; at the same time, in various experiments, style-level supervised generative models are difficult to produce results comparable to pixel-level supervised generative models; in addition, pixel-level supervised models have very high requirements for training data sets, and the performance of deep learning-based models is heavily dependent on a large amount of high-quality training data, making image translation very difficult.

[0005] Therefore, it is very necessary to construct a pixel-level aligned night infrared and visible light image data set in the image translation task.

[0006] In the prior art, commonly used paired infrared and visible image datasets are mostly constructed in the following three ways: one is to use binocular infrared and visible image sensors to simultaneously collect infrared and visible images; two is to collect RGB and NIR images of a scene by fixing a tripod and replacing the camera, similarly, there is also a dual-CCD camera system in which NIR and RGB spectra are separated and directed to dedicated CCD sensors, and NIR and RGB images are recorded simultaneously; three is a physical coloring method.

[0007] However, the first method of generating paired data has some unavoidable defects: on the one hand, due to the thermal imaging principle of infrared cameras, there can be a large difference between infrared images taken during the day and at night, and this domain difference makes it difficult for a generation model trained with daytime infrared images to produce satisfactory results on nighttime infrared images; on the other hand, no matter how close the distance between the cameras is, there is still a perspective difference between the images obtained by the cameras. The dataset obtained by the second method overcomes the perspective difference, but still cannot collect pixel-level aligned images across time periods, and the dataset is small in size, limited in scene, and low in collection efficiency. The physical coloring method is less used, and the dataset obtained is often of poor quality.

[0008] In general, it is very difficult to construct a large-scale cross-time period pixel-level aligned infrared and visible image pair dataset, which greatly limits the application of generation models in the field of nighttime infrared. SUMMARY

[0009] Therefore, it is necessary to provide a pixel-level infrared and visible image pair generation method, device, equipment and medium to construct a large-scale cross-time period pixel-level aligned infrared and visible image pair dataset.

[0010] The pixel-level infrared and visible image pair generation method comprises:

[0011] synchronously acquiring a daytime point cloud sequence and a corresponding visible image sequence, and synchronously acquiring a nighttime point cloud sequence and a corresponding infrared image sequence;

[0012] constructing a three-dimensional point cloud map, matching the daytime point cloud sequence and the nighttime point cloud sequence with the three-dimensional point cloud map respectively to obtain a radar pose, and matching the infrared image sequence according to the radar pose to take any infrared image and a visible image closest to the pose of the any infrared image in the visible image sequence as a matching image pair;

[0013] projecting the daytime point cloud sequence or the nighttime point cloud sequence onto the visible image and the infrared image of the matching image pair respectively to obtain a visible sparse depth map and an infrared sparse depth map;

[0014] Optimizing the visible light sparse depth map and the infrared sparse depth map respectively to obtain a visible light dense depth map and an infrared dense depth map;

[0015] According to the color of each pixel point on the visible light dense depth map, coloring each pixel point on the infrared dense depth map to generate a pixel-level aligned image pair.

[0016] In one embodiment, according to the radar pose, matching the infrared image sequence to take any infrared image and the visible light image in the visible light image sequence closest to the pose of the any infrared image as a matching image pair, including:

[0017]

[0018] Wherein:

[0019] d(P n ,P d )=‖(x n ,y n ,z n ,θ n )-(x d ,y d ,z d ,θ d )‖2

[0020] In the formula, is the Euclidean distance of the jth night pose and the kth day pose, is the jth night pose, is the ith day pose.

[0021] In one embodiment, projecting the day point cloud sequence or the night point cloud sequence onto the visible light image and the infrared image of the matching image pair respectively to obtain a visible light sparse depth map and an infrared sparse depth map, including:

[0022] Projecting the day point cloud sequence onto the visible light image of the matching image pair to obtain a visible light sparse depth map, and projecting the day point cloud sequence onto the infrared image of the matching image pair to obtain an infrared sparse depth map;

[0023] Or, projecting the night point cloud sequence onto the visible light image of the matching image pair to obtain a visible light sparse depth map, and projecting the night point cloud sequence onto the infrared image of the matching image pair to obtain an infrared sparse depth map.

[0024] In one embodiment, projecting the night point cloud sequence onto the infrared image of the matching image pair to obtain an infrared sparse depth map, including:

[0025]

[0026] wherein, is the night-time point cloud sequence, is the infrared sparse depth map, g is a projection function from the night-time point cloud coordinate system to the infrared image coordinate system, CTI n is the night-time intrinsic parameter, LTC n is the night-time extrinsic parameter.

[0027] In one embodiment, the night-time point cloud sequence is projected onto the visible light image of the matching image pair to obtain a visible light sparse depth map, including:

[0028]

[0029] wherein, is the night-time point cloud sequence, is the visible light sparse depth map, f is a projection function from the night-time point cloud coordinate system to the visible light image coordinate system, CTI d is the day-time intrinsic parameter, LTC d is the day-time extrinsic parameter, is the pose conversion parameter.

[0030] In one embodiment, the visible light sparse depth map and the infrared sparse depth map are respectively optimized to obtain a visible light dense depth map and an infrared dense depth map, including:

[0031] The visible light sparse depth map and the infrared sparse depth map are respectively processed using a minimum filter with a mask to assign a blank pixel to the nearest surrounding point, using a morphological closing and dilation operation to fill in the holes, using a median filter and a Gaussian filter to remove noise, to achieve depth completion, and through depth back projection to obtain a point cloud smoothing result, to generate the visible light dense depth map and the infrared dense depth map.

[0032] In one embodiment, the visible light sparse depth map and the infrared sparse depth map are respectively optimized to obtain a visible light dense depth map and an infrared dense depth map, further including:

[0033] When the sparse color map contains a sky area, a set value is assigned to the sky area to distinguish the sky area from other areas.

[0034] A pixel-level infrared and visible light image pair generation device, including:

[0035] An acquisition module is configured to synchronously acquire a day-time point cloud sequence and a corresponding visible light image sequence, and synchronously acquire a night-time point cloud sequence and a corresponding infrared image sequence.

[0036] The matching module is configured to construct a three-dimensional point cloud map, match the daytime point cloud sequence and the nighttime point cloud sequence with the three-dimensional point cloud map respectively, and obtain radar poses; and match the infrared image sequence according to the radar poses, so as to take any infrared image and a visible light image closest to the pose of the any infrared image in the visible light image sequence as a matching image pair.

[0037] The projection module is configured to project the daytime point cloud sequence or the nighttime point cloud sequence onto the visible light image and the infrared image of the matching image pair respectively, so as to obtain a visible light sparse depth map and an infrared sparse depth map.

[0038] The optimization module is configured to optimize the visible light sparse depth map and the infrared sparse depth map respectively, so as to obtain a visible light dense depth map and an infrared dense depth map.

[0039] The output module is configured to color each pixel point on the infrared dense depth map according to the color of each pixel point on the visible light dense depth map, so as to generate a pixel-level aligned image pair.

[0040] A computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the following steps when executing the computer program:

[0041] The daytime point cloud sequence and the corresponding visible light image sequence are synchronously acquired, and the nighttime point cloud sequence and the corresponding infrared image sequence are synchronously acquired;

[0042] A three-dimensional point cloud map is constructed, the daytime point cloud sequence and the nighttime point cloud sequence are matched with the three-dimensional point cloud map respectively, and radar poses are obtained; and the infrared image sequence is matched according to the radar poses, so as to take any infrared image and a visible light image closest to the pose of the any infrared image in the visible light image sequence as a matching image pair.

[0043] The daytime point cloud sequence or the nighttime point cloud sequence is projected onto the visible light image and the infrared image of the matching image pair respectively, so as to obtain a visible light sparse depth map and an infrared sparse depth map.

[0044] The visible light sparse depth map and the infrared sparse depth map are optimized respectively, so as to obtain a visible light dense depth map and an infrared dense depth map.

[0045] Each pixel point on the infrared dense depth map is colored according to the color of each pixel point on the visible light dense depth map, so as to generate a pixel-level aligned image pair.

[0046] A computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the following steps:

[0047] Synchronously acquire a daytime point cloud sequence and a corresponding visible light image sequence, and synchronously acquire a nighttime point cloud sequence and a corresponding infrared image sequence;

[0048] Construct a three-dimensional point cloud map, match the daytime point cloud sequence and the nighttime point cloud sequence with the three-dimensional point cloud map respectively to obtain a radar pose, and match the infrared image sequence according to the radar pose to take any infrared image and a visible light image closest to the pose of the infrared image as a matching image pair;

[0049] Project the daytime point cloud sequence or the nighttime point cloud sequence onto the visible light image and the infrared image of the matching image pair respectively to obtain a visible light sparse depth map and an infrared sparse depth map;

[0050] Optimize the visible light sparse depth map and the infrared sparse depth map respectively to obtain a visible light dense depth map and an infrared dense depth map;

[0051] Color each pixel point on the infrared dense depth map according to the color of each pixel point on the visible light dense depth map to generate a pixel-level aligned image pair.

[0052] The pixel-level infrared visible light image pair generation method described above is an efficient large-scale cross-time period pixel-level aligned infrared visible light image pair generation method, and can be used to construct a large-scale pixel-level aligned infrared visible light dataset by using multiple sensors. The present application collects infrared images and visible light images at night and during the day respectively to solve the modal difference problem caused by the thermal imaging principle of the infrared camera, so as to keep the modal consistency of the training data and the test data. In view of the problem of cross-time period image misalignment, a three-dimensional high-precision point cloud map is constructed in advance, the point cloud data is used as a medium, the pose of the image sensor is obtained through global matching with the map, and then the conversion relationship between the infrared camera coordinate system and the visible light camera coordinate system is calculated through sensor time synchronization and spatial calibration. By using this relationship, the pixel value of the visible light image is projected onto the corresponding pixel point of the infrared image, so that the generated image pair is pixel-level aligned. Compared with other methods, the generated results of the present method are closer to the real images, and the efficient generation framework makes it possible to prepare a large-scale high-resolution dataset. BRIEF DESCRIPTION OF DRAWINGS

[0053] Figure 1 An application scenario diagram of the pixel-level infrared visible light image pair generation method in one embodiment;

[0054] Figure 2 A flowchart of the pixel-level infrared visible light image pair generation method in one embodiment;

[0055] Figure 3Fig. 1 is a schematic diagram of a three-dimensional high-precision point cloud map in one embodiment;

[0056] Figure 4 Fig. 2 is a diagram of specific conversion parameters and coordinate system conversion relationships in one embodiment;

[0057] Figure 5 Fig. 3 is a schematic diagram of point cloud projection and point cloud optimization in one embodiment;

[0058] Figure 6 Fig. 4 is a schematic diagram of a data acquisition platform in one specific embodiment;

[0059] Figure 7 Fig. 5 is a comparison diagram of generated image pairs and other modal data in one specific embodiment;

[0060] Figure 8 Fig. 6 is a schematic diagram of error analysis results in one specific embodiment, where (a) is point error, (b) is labeling method, (c) is distance error, and (d) is angle error;

[0061] Figure 9 Fig. 7 is a structural block diagram of a pixel-level infrared-visible light image pair generation device in one embodiment;

[0062] Figure 10 Fig. 8 is an internal structure diagram of a computer device in one embodiment. DETAILED DESCRIPTION

[0063] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of the present application.

[0064] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present application are only used to explain the relative positional relationship, movement condition, etc. between components in a certain posture (as shown in the drawings), and if the certain posture changes, the directional indications also change accordingly.

[0065] In addition, the descriptions such as "first", "second" and the like in the present application are only for the purpose of description, and cannot be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the features limited by "first" and "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "multiple groups" is at least two groups, such as two groups, three groups, etc., unless otherwise specifically limited.

[0066] In this application, unless otherwise clearly specified and limited, the terms "connection", "fixing", and the like should be understood in a broad sense, for example, "fixing" can be fixed connection, or detachable connection, or integral; can be mechanical connection, or electrical connection, or physical connection, or wireless communication connection; can be direct connection, or indirect connection through intermediate medium, can be internal connection of two elements or interaction relationship between two elements, unless otherwise clearly limited. For those skilled in the art, the specific meaning of the above terms in this application can be understood according to the specific circumstances.

[0067] In addition, the technical solutions of various embodiments of the present application can be combined with each other, but it must be based on the fact that a person skilled in the art can realize it, when the combination of technical solutions appears contradictory or unachievable, it should be considered that the combination of technical solutions does not exist, nor is it within the protection scope required by the present application.

[0068] The pixel-level infrared and visible light image pair generation method provided by the present application can be applied to the application scenario as shown in the figure. Figure 1 The terminal 102 communicates with the server 104 through the network, and the terminal 102 can include but is not limited to various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices, and the server 104 can be a server corresponding to various portal websites and work system backgrounds.

[0069] The present application provides a pixel-level infrared and visible light image pair generation method, as shown in the flowchart. Figure 2 In one embodiment, the method is applied to the terminal in Figure 1 , including:

[0070] Step 202, synchronously acquiring a daytime point cloud sequence and a corresponding visible light image sequence, and synchronously acquiring a nighttime point cloud sequence and a corresponding infrared image sequence.

[0071] Specifically:

[0072] A visible light image sensor, an infrared image sensor and a 128-line laser radar are fixed on a vehicle as a data acquisition platform. During data acquisition, the same route is driven in the daytime and at night, respectively, and the modal image (including visible light image and infrared image) and the corresponding laser radar point are synchronously collected, so as to obtain a daytime visible light image sequence a nighttime infrared image sequence a daytime point cloud sequence and a nighttime point cloud sequence

[0073] In this step, it is necessary to point out that if no processing is done, the data collected by the lidar and the image sensor are not matched, so time synchronization and spatial calibration are set, wherein the time synchronization includes sensor synchronization and motion compensation, and the spatial calibration includes internal parameter optimization and external parameter optimization.

[0074] 1) Sensor synchronization: a synchronization system based on FPGA is designed, the infrared frame synchronization signal is used as a trigger to synchronize the lidar, and the exposure of the RGB camera sensor is triggered at the same time, realizing the synchronization of multiple sensors, that is, the image sensor will be exposed when the laser radar scanning line passes through the front of the vehicle;

[0075] 2) Motion compensation: a fixed time period is needed for the rotary lidar to obtain a complete point cloud, during which the vehicle carrying the sensor will continue to move forward, so the coordinates of each radar point in the data are relative to a moving vehicle coordinate system, causing distortion and misplacement of scene data. In order to solve this problem, the point cloud in the frame is compensated based on the vehicle IMU data, and all data are aligned at a certain moment;

[0076] 3) Internal parameter optimization: based on Zhang's calibration method, the internal parameters are calculated, and the projection method is used to complete the conversion from the camera coordinate system to the image coordinate system;

[0077] 4) External parameter optimization: based on the cross-modal target matching network (ATOP), the external parameters are calculated, and the rigid transformation method is used to complete the conversion from the radar coordinate system to the camera coordinate system.

[0078] Therefore, in this step, on the basis of time synchronization, spatial calibration is carried out, and internal parameters CTI and external parameters LTC are obtained, which describe the conversion relationship between the radar coordinate system, the camera coordinate system and the image coordinate system.

[0079] Step 204, constructing a three-dimensional point cloud map, matching the daytime point cloud sequence and the nighttime point cloud sequence with the three-dimensional point cloud map respectively to obtain the radar pose; according to the radar pose, matching the infrared image sequence, taking any infrared image and the visible light image closest to the pose of any infrared image in the visible light image sequence as a matching image pair.

[0080] Specifically:

[0081] The laser radar scans the entire data acquisition area to obtain raw point cloud data, and the raw point cloud data is preprocessed (including denoising, filtering, point cloud registration, etc.). Then, the preprocessed raw point cloud data is converted into a world coordinate system by using IMU data, three-dimensional modeling is performed by accumulating point cloud frames, and finally, a three-dimensional point cloud map is constructed by using SLAM to establish the correlation between scene points across time periods. The constructed three-dimensional point cloud map is a three-dimensional high-precision point cloud map, as shown in Figure 3 .

[0082] Each frame of daytime point cloud and each frame of nighttime point cloud is matched and positioned with the three-dimensional point cloud map to obtain the radar pose corresponding to the time, which includes daytime radar pose and nighttime radar pose

[0083] The pose can be expressed as where (x, y, z) represents the coordinates of the laser radar in the world coordinate system, represents the deflection angle of the vehicle, and the conversion relationship between the daytime pose coordinates and the nighttime pose coordinates and the origin of the world coordinate system is denoted as RT d and RT n , respectively. The specific conversion parameters and coordinate system conversion relationship are shown in Figure 4 It can be seen that the daytime and nighttime poses (represented by yellow and black trajectories) are difficult to maintain complete consistency in time, that is,

[0084]

[0085] where is the jth frame of nighttime radar pose, is the ith frame of daytime radar pose, and

[0086] That is, there is a perspective difference between the daytime visible light image and the nighttime infrared image .

[0087] To ensure that there are as many coincident scene points as possible between point cloud frames across time periods, the infrared image sequence is matched according to the radar pose, taking any infrared image and the visible light image in the visible light image sequence closest to the pose of the any infrared image as a matching image pair. Specifically, for any given infrared image , the corresponding nighttime radar pose , the daytime radar pose closest to it in the Euclidean distance, and the corresponding visible light image As a preliminary matching image, any given infrared image and the corresponding preliminary matching image form a matching image pair:

[0088]

[0089] wherein:

[0090] d(P n ,P d )=‖(x n ,y n ,z n ,θ n )-(x d ,y d ,z d ,θ d )‖2

[0091] wherein, is the Euclidean distance between the jth night-time pose and the kth day-time pose, is the jth night-time pose, is the ith day-time pose, it should be noted that, is the day-time pose closest to the jth night-time pose.

[0092] In this step, since the vehicle may return along the original route during data collection, a restriction on θ is added in the pose distance calculation to ensure that the modal image corresponding to the closest pose is taken in the same forward direction, so that the obtained image pair can guarantee the highest scene coincidence rate.

[0093] Step 206, projecting the day-time point cloud sequence or the night-time point cloud sequence onto the visible light image and the infrared image of the matching image pair respectively to obtain a visible light sparse depth map and an infrared sparse depth map.

[0094] Specifically:

[0095] projecting the day-time point cloud sequence onto the visible light image of the matching image pair to obtain a visible light sparse depth map, and projecting the day-time point cloud sequence onto the infrared image of the matching image pair to obtain an infrared sparse depth map;

[0096] or, projecting the night-time point cloud sequence onto the visible light image of the matching image pair to obtain a visible light sparse depth map, and projecting the night-time point cloud sequence onto the infrared image of the matching image pair to obtain an infrared sparse depth map.

[0097] More specifically:

[0098] projecting the night-time point cloud sequence onto the infrared image of the matching image pair to obtain an infrared sparse depth map:

[0099]

[0100] wherein, is a night-time point cloud sequence (jth frame night-time point cloud), is an infrared image , g is a projection function from the night-time point cloud coordinate system to the infrared image coordinate system, CTI n is a night-time intrinsic parameter, LTC n is a night-time extrinsic parameter;

[0101] projecting the night-time point cloud sequence onto the visible light image of the matched image pair to obtain a visible light sparse depth map, comprising:

[0102]

[0103] wherein, is a night-time point cloud sequence (jth frame night-time point cloud), is a visible light image , f is a projection function from the night-time point cloud coordinate system to the visible light image coordinate system, CTI d is a day-time intrinsic parameter, LTC d is a day-time extrinsic parameter, is a pose transformation parameter.

[0104] In this step, the night-time point cloud is first transformed to the day-time radar coordinate system through the pose transformation parameter, and then projected onto the plane of the visible light image to obtain a sparse depth map of the infrared image with the same resolution as the infrared image .

[0105] Step 208, respectively optimizing the visible light sparse depth map and the infrared sparse depth map to obtain a visible light dense depth map and an infrared dense depth map.

[0106] Specifically:

[0107] respectively for the visible light sparse depth map and the infrared sparse depth map, using a minimum filter with a mask to assign a blank pixel to the nearest surrounding point, using a morphological closing and dilation operation to fill in the holes, using a median filter and a Gaussian filter to remove noise, to achieve depth completion, and through depth back projection to obtain point cloud smoothing results, to generate a visible light dense depth map and an infrared dense depth map.

[0108] When the sparse color map contains a sky area, a set value is assigned to the sky area to distinguish the sky area from other areas.

[0109] More specifically:

[0110] Estimate the complete object surface under the current view to avoid the false projection of point cloud; adopt a fast depth completion method to complete the sparse depth map with basic image processing operations Completing: Set a depth completion module, assign the blank pixels to the nearest points around using a minimum filter with a mask, fill the holes using morphological closing and dilation operations, remove the noise using a median filter and a Gaussian filter, and get the final point cloud smoothing result through depth back projection, so as to realize image completion.

[0111] Since the laser radar has no echo in the sky area, In the sky area in , there is no corresponding depth value, set a sky projection module to assign a large enough value to the sky area And satisfy:

[0112]

[0113] In the formula, is the dense point cloud corresponding to the jth infrared image, is the depth value of the point cloud p in

[0114] That is, the distance of the radar to the sky is infinite, the surface of the sky can be approximated as a plane, and a large enough value can effectively distinguish the sky from other ground areas.

[0115] The sparse depth map obtained in the previous step is still in a sparse state, especially on the road surface close to the radar, and even if interpolation is used, it is difficult to cover all pixels; in addition, the dense point cloud is formed by accumulating multiple frames of point clouds into the same frame, which will inevitably cause projection errors, and due to the existence of holes in the point cloud of the object surface, points that should be blocked by the object will also be colored.

[0116] In this step, the depth completion module and the sky projection module are set for the visible light sparse depth map and the infrared sparse depth map to avoid the region missing caused by the distance limitation of the radar, the inherent sparsity of the point cloud, and the false projection caused by the accumulation of point clouds.

[0117] Step 210, according to the color of each pixel point on the visible light dense depth map, color each pixel point on the infrared dense depth map to generate a pixel-level aligned image pair.

[0118] In this step, since the cross-period point cloud projection has been completed, the points in are used as the starting point, and ​Correspondence between pixels can be found, and accordingly, the correspondence between some points in and can be found, so as to color the image; that is, the corresponding pixel pair in is found by using point cloud projection to color, so as to realize pixel-level alignment.

[0119] However, if the original point cloud is directly projected, it is found that the projection points in and are very sparse, and a large number of pixel points cannot be indexed, therefore, preferably, the dense point cloud within the range of 70mx70m in front of the vehicle is taken out from the high-precision map by using the pose and as a new medium, and is converted to the camera coordinate system for regional filtering to remove the perspective projection points, so that the night infrared image pixel points can be well colored, thereby completing the point cloud optimization, as shown in Figure 5 .

[0120] In the embodiment, the pose of the image sensor at each time can be obtained by point cloud matching between the point cloud frame and the pre-constructed three-dimensional high-precision map; for the visible light image sensor pose corresponding to each frame of daytime visible light image, the nearest night infrared image sensor pose can be found, and the corresponding night infrared image can be obtained, and naturally, the scene coincidence degree of the two images is the highest; however, due to the fact that the speed, trajectory and fixed position of the image sensor of the vehicle cannot remain completely consistent at different time periods, the modal images are not pixel-level aligned (the same object marked by the rectangular frame in Figure 3 is at different positions), and due to the depth in the three-dimensional scene, the deviation cannot be eliminated by simple affine transformation, therefore, it is proposed to take the point cloud as a medium, realize the pixel point alignment between the modal images by point cloud projection, and fill the pixel value in the visible light image to the corresponding pixel position of the infrared image; further, the point cloud is optimized to obtain the pixel-dense image pair.

[0121] The pixel-level infrared and visible light image pair generation method is an efficient large-scale cross-time period pixel-level alignment infrared and visible light image pair generation method, and can construct a large-scale pixel-level alignment infrared and visible light dataset by using multiple sensors. The present application is aimed at the modal difference problem caused by the thermal imaging principle of the infrared camera, and collects infrared images and visible light images at night and during the day respectively, so as to keep the modal consistency of the training data and the test data; in view of the problem of cross-time period image misalignment, a three-dimensional high-precision point cloud map is constructed in advance, the point cloud data is used as a medium, the pose of the image sensor is obtained through global matching with the map, and then the conversion relationship between the infrared camera coordinate system and the visible light camera coordinate system is calculated through sensor time synchronization and space calibration, and by using the relationship, the pixel value of the visible light image is projected onto the corresponding pixel point of the infrared image, so that the generated image pair is pixel-level aligned. Compared with other methods, the generated results of the present method are closer to the real images, and the efficient generation framework makes it possible to prepare a large-scale high-resolution dataset.

[0122] It should be understood that, although Figure 2 the steps in the flowchart of the method are shown in sequence according to the arrows, these steps are not necessarily executed in sequence according to the arrows. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, Figure 2 at least part of the steps in the method can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these sub-steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or sub-steps or stages of other steps.

[0123] In a specific embodiment, the method of the present application is experimentally verified.

[0124] 1) Experimental environment:

[0125] A data acquisition platform as shown in Figure 6 is used. The vehicle is equipped with a Robosense 128-line laser radar, an Xsens MTI300 IMU, a NovAtel SPAN-CPT GNSS / INS system, a visible light camera and an infrared camera, and a rural ring road is selected as a fixed route for data acquisition.

[0126] As shown in Figure 7The contrast chart of the generated image pair and other modal data is shown, the first column is the night infrared image, and the second column is the visible light image generated result of pixel-level alignment of the first column. When generating the cross-modal image pair, a large amount of other modal data is collected at the same time, and the scene includes: sandy and cement road, turning and uphill, complex texture such as branches and leaves, and the generation effect is good. Due to the difference in viewing angle, part of the infrared image cannot be colored, so a mask of the effective area is provided, and the point cloud data is projected onto its corresponding modal image for visual representation.

[0127] 2) Error analysis:

[0128] In order to illustrate the reliability of the image pair generation method of the present application, 100 pairs of infrared-colored images and corresponding infrared-primarily matched original visible light images are randomly selected, and point error, distance error and angle error are used as three indicators for confidence analysis. Specifically, two clear target points p1=(x1,y1) and p2=(x2,y2) are marked on the image, and respectively represent the same position of the infrared and colored image pair, especially (when calculating the error of the primarily matched original visible light image, will be replaced).

[0129] As shown in Figure 8 (b), the target points are usually selected on a fixed reference object, such as a building or a tree trunk, to avoid errors caused by non-algorithmic reasons. The calculation of each indicator is as follows:

[0130]

[0131]

[0132]

[0133]

[0134] wherein,

[0135] e(p1,p2)=||p1-p2||2

[0136] θ(p1,p2)=arctan((x2-x1) / (y2-y1))

[0137] In the formula, r represents the resolution of the image.

[0138] The above indicators evaluate the accuracy of pixel alignment and the degree of affine transformation of the two images, thereby giving the overall "matching degree" of the image pair.

[0139] As Figure 8 The error analysis results shown in (a), 8(c), 8(d) show that the point error (i.e., colored) of the colored image is basically within 3% of the image resolution, which is significantly lower than 40% (i.e., raw) of the original paired visible light image. Similarly, the distance error of the colored image is within 10%, and the angle error is 0.14 / 2π.

[0140] 3) Conclusion:

[0141] People have poor perception of the night environment, which has given rise to the demand for translation from night infrared images to daytime visible light images, and the thermal imaging characteristics of infrared cameras make it difficult to construct a pixel-level aligned pair of night infrared and visible light images. The present application proposes an efficient pixel-level aligned image pair generation method, which makes it possible to quickly generate a large number of high-resolution night infrared and visible light image pairs. Experimental results show that the data set generated by the method in the present application is reliable, has strong practicality in related image translation, and has applicability in many weather and different scenes, and can further improve the sensor synchronization performance to achieve better coloring effect.

[0142] The present application also provides a pixel-level infrared and visible light image pair generation device, as shown in Figure 9 In one embodiment, it comprises: an acquisition module 902, a matching module 904, a projection module 906, a coloring module 908 and an output module 910, wherein:

[0143] The acquisition module 902 is configured to synchronously acquire a daytime point cloud sequence and a corresponding visible light image sequence, and synchronously acquire a night point cloud sequence and a corresponding infrared image sequence;

[0144] The matching module 904 is configured to construct a three-dimensional point cloud map, match the daytime point cloud sequence and the night point cloud sequence with the three-dimensional point cloud map respectively, and obtain a radar pose; according to the radar pose, match the infrared image sequence, and take any infrared image and the visible light image closest to the pose of the any infrared image in the visible light image sequence as a matching image pair;

[0145] The projection module 906 is configured to project the daytime point cloud sequence or the night point cloud sequence onto the visible light image and the infrared image of the matching image pair respectively, to obtain a visible light sparse depth map and an infrared sparse depth map;

[0146] The coloring module 908 is configured to optimize the visible light sparse depth map and the infrared sparse depth map respectively, to obtain a visible light dense depth map and an infrared dense depth map;

[0147] The output module 910 is configured to color each pixel point on the infrared dense depth map according to the color of each pixel point on the visible light dense depth map, to generate a pixel-level image pair.

[0148] For specific limitations of the pixel-level infrared and visible light image pair generation apparatus, refer to the limitations of the pixel-level infrared and visible light image pair generation method described above, which will not be repeated here. Each module in the apparatus can be realized by software, hardware, and combinations thereof, in whole or in part. Each module described above can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0149] In one embodiment, a computer device is provided, which can be a terminal, and its internal structure diagram can be as shown in Figure 10 The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected by a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the computer device is configured to communicate with external terminals through network connections. The computer program is executed by the processor to implement a pixel-level infrared and visible light image pair generation method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball, or touchpad provided on the shell of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0150] Those skilled in the art can understand that Figure 10 the structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0151] In one embodiment, a computer device is provided, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the method in the above embodiments.

[0152] In one embodiment, a computer readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the steps of the method in the above embodiments.

[0153] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, storage, databases, or other media in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0154] Any combination of the technical features of the above embodiments can be made, and in order to make the description simple, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.

[0155] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A method for generating pixel-level infrared and visible light image pairs, characterized in that, include: Simultaneously acquire daytime point cloud sequences and corresponding visible light image sequences, and simultaneously acquire nighttime point cloud sequences and corresponding infrared image sequences; A three-dimensional point cloud map is constructed, and the daytime point cloud sequence and the nighttime point cloud sequence are matched with the three-dimensional point cloud map to obtain the radar pose. Based on the radar pose, the infrared image sequence is matched, and the visible light image that is closest to the pose of any infrared image in any infrared image and visible light image sequence is taken as a matching image pair. The daytime point cloud sequence or the nighttime point cloud sequence is projected onto the visible light image and infrared image of the matched image pair, respectively, to obtain the visible light sparse depth map and the infrared sparse depth map. The visible light sparse depth map and the infrared sparse depth map are optimized respectively to obtain the visible light dense depth map and the infrared dense depth map. Based on the color of each pixel in the visible light dense depth map, each corresponding pixel in the infrared dense depth map is colored to generate pixel-level aligned image pairs. The daytime point cloud sequence or the nighttime point cloud sequence is projected onto the visible light image and infrared image of the matched image pair, respectively, to obtain a visible light sparse depth map and an infrared sparse depth map, including: The nighttime point cloud sequence is projected onto the visible light image of the matching image pair to obtain a visible light sparse depth map, and the nighttime point cloud sequence is projected onto the infrared image of the matching image pair to obtain an infrared sparse depth map. Projecting the nighttime point cloud sequence onto the visible light image of the matched image pair yields a visible light sparse depth map, including: In the formula, This is a nighttime point cloud sequence. This is a sparse depth map for visible light. This is the projection function from the nighttime point cloud coordinate system to the visible light image coordinate system. This is a daytime internal reference. The daytime extrinsic parameters, intrinsic parameters, and extrinsic parameters describe the transformation relationships between the radar coordinate system, camera coordinate system, and image coordinate system. These are pose transformation parameters. and These represent the transformation relationships between daytime pose coordinates and nighttime pose coordinates and the origin of the world coordinate system, respectively. The nighttime point cloud is transformed to the daytime radar coordinate system through pose transformation parameters. The nighttime point cloud sequence is projected onto the infrared image of the matched image pair to obtain an infrared sparse depth map, including: In the formula, This is a nighttime point cloud sequence. This is an infrared sparse depth map. This is the projection function from the nighttime point cloud coordinate system to the infrared image coordinate system. This is an internal reference report for nighttime use. This is for nighttime external reference.

2. The pixel-level infrared-visible image pair generation method according to claim 1, characterized in that, Based on the radar pose, the infrared image sequence is matched, and a matching image pair is formed by taking any infrared image and the visible light image in the visible light image sequence that is closest to the pose of any infrared image, including: in: In the formula, Let be the Euclidean distance between the nighttime pose of frame j and the daytime pose of frame k. For the nighttime pose of frame j, For the daytime pose of the i-th frame, This indicates the coordinates of the lidar in the world coordinate system.

3. The pixel-level infrared-visible image pair generation method according to claim 1 or 2, characterized in that, Projecting the daytime point cloud sequence or the nighttime point cloud sequence onto the visible light image and infrared image of the matched image pair, respectively, to obtain a visible light sparse depth map and an infrared sparse depth map, further comprising: The daytime point cloud sequence is projected onto the visible light image of the matching image pair to obtain a visible light sparse depth map, and the daytime point cloud sequence is projected onto the infrared image of the matching image pair to obtain an infrared sparse depth map.

4. The pixel-level infrared-visible image pair generation method according to claim 1 or 2, characterized in that, The visible light sparse depth map and the infrared sparse depth map are optimized respectively to obtain the visible light dense depth map and the infrared dense depth map, including: For the visible light sparse depth map and the infrared sparse depth map respectively, a minimum value filter with a mask is used to assign the blank pixel to the nearest point around it, morphological closure and expansion operations are used to fill the hole, and a median filter and a Gaussian filter are used to remove noise to achieve depth completion. The point cloud smoothing result is obtained by depth back projection to generate a visible light dense depth map and an infrared dense depth map.

5. The pixel-level infrared-visible image pair generation method according to claim 4, characterized in that, The visible light sparse depth map and the infrared sparse depth map are optimized respectively to obtain the visible light dense depth map and the infrared dense depth map, which also includes: When a sparsely shaded map contains a sky region, assign a setting value to the sky region to distinguish it from other regions.

6. A pixel-level infrared-visible light image pair generation device, characterized in that, The pixel-level infrared-visible image pair generation method according to any one of claims 1 to 5 includes: The acquisition module is used to simultaneously acquire daytime point cloud sequences and corresponding visible light image sequences, and simultaneously acquire nighttime point cloud sequences and corresponding infrared image sequences. The matching module is used to construct a three-dimensional point cloud map, and match the daytime point cloud sequence and the nighttime point cloud sequence with the three-dimensional point cloud map to obtain the radar pose; based on the radar pose, the infrared image sequence is matched, and the visible light image that is closest to the pose of any infrared image in any infrared image and visible light image sequence is taken as a matching image pair. The projection module is used to project the daytime point cloud sequence or the nighttime point cloud sequence onto the visible light image and infrared image of the matched image pair, respectively, to obtain the visible light sparse depth map and the infrared sparse depth map. The optimization module is used to optimize the visible light sparse depth map and the infrared sparse depth map respectively to obtain the visible light dense depth map and the infrared dense depth map. The output module is used to color each corresponding pixel in the infrared dense depth map according to the color of each pixel in the visible light dense depth map, and generate pixel-aligned image pairs.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.