Cross-spectral feature detection for volumetric image alignment and shading
By combining LIDAR and HDR cameras to perform image feature detection, color photographs are automatically aligned and colored with laser scan intensity images, solving the problem of time-consuming manual matching in existing technologies and achieving efficient image alignment and colorization effects.
Patent Information
- Application Number
- CN202080051155.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-05-26
- Filing Date
- 2020-12-10
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2040-12-10
AI Technical Summary
In existing technologies, manually matching the corresponding points between color photographs and intensity images generated by laser scanning is cumbersome and time-consuming.
Intensity data is captured using a LiDAR scanner and color information is captured using an HDR camera. The intensity image and color image are automatically aligned through image feature detection and matching. The color information is applied to generate a color image, and the GPU is used for fine-tuning and lens distortion correction.
It automates the image alignment and coloring process, improving efficiency, reducing manual operation time, and increasing the accuracy and speed of image matching.
Smart Images

Figure CN114127786B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to image alignment and colorization, and more specifically, to cross-spectral feature detection for volumetric image alignment and colorization. BACKGROUND
[0002] Video systems for video production, studio environments, or virtual production can use a combination of laser scanners and color photography to recreate a scene. This recreation operation can include manually using an image alignment program to match points between color photographs and intensity images generated from laser scans. However, the manual operation of defining and matching corresponding points in two images is very tedious and takes a long time to process. SUMMARY
[0003] The present disclosure provides volumetric image alignment and colorization.
[0004] In one implementation, a method for image alignment and colorization is disclosed. The method includes: capturing intensity data using at least one scanner; generating an intensity image using the intensity data, wherein the intensity image includes at least one feature in a scene, the at least one feature including a sample feature; capturing image data using at least one camera, wherein the image data includes color information; generating a camera image using the image data, wherein the camera image includes the sample feature; matching the sample feature in the intensity image with the sample feature in the camera image to align the intensity image with the camera image; and generating a color image by applying the color information to the aligned intensity image.
[0005] In one implementation, the at least one scanner includes at least one LIDAR scanner. In one implementation, the intensity data captured by the at least one LIDAR scanner is a 3-D intensity image. In one implementation, the method further includes applying color information on the 3-D intensity image. In one implementation, the method further includes generating a 2-D intensity image from the 3-D intensity image. In one implementation, the at least one camera includes at least one HDR camera. In one implementation, the camera data captured by the at least one HDR camera is a 2-D color photograph. In one implementation, generating the camera image using the image data includes exposure stacking and color correction of the image data to generate the camera image. In one implementation, the method further includes: generating one or more control points in the intensity image; and receiving adjustments to the alignment of the intensity image and the camera image. In one implementation, the method further includes: correcting for any lens distortion in the camera image; and receiving adjustments to the alignment of the intensity image and the camera image.
[0006] In another embodiment, a system for aligning and coloring volumetric images is disclosed. The system includes at least one scanner for capturing intensity data; at least one camera for capturing image data, the image data including color information; and a processor for: generating an intensity image using the captured intensity data, wherein the intensity image includes at least one feature in a scene, the at least one feature including a sample feature; generating a camera image using the image data, wherein the camera image includes the sample feature; matching the sample feature in the intensity image to the sample feature in the camera image to align the intensity image to the camera image; and generating a color image by applying the color information to the aligned intensity image.
[0007] In an embodiment, the at least one scanner includes at least one LIDAR scanner. In an embodiment, the intensity data captured by the at least one LIDAR scanner is a 3-D intensity image. In an embodiment, the at least one camera includes at least one HDR camera. In an embodiment, the system further includes a cloud cluster for receiving the intensity image and the camera image from the processor, performing the matching and alignment, and sending the results back to the processor.
[0008] In another embodiment, a non-transitory computer-readable storage medium storing a computer program for aligning and coloring volumetric images is disclosed. The computer program includes executable instructions that cause a computer to: capture video data using a plurality of cameras; capture intensity data; generate an intensity image using the intensity data, wherein the intensity image includes at least one feature in a scene, the at least one feature including a sample feature; capture image data, wherein the image data includes color information; generate a camera image using the image data, wherein the color image includes the sample feature; match the sample feature in the intensity image to the sample feature in the camera image to align the intensity image to the camera image; and generate a color image by applying the color information to the aligned intensity image.
[0009] In an embodiment, the captured intensity data is a 3-D intensity image. In an embodiment, the computer-readable storage medium further includes executable instructions that cause the computer to generate a 2-D intensity image from the 3-D intensity image. In an embodiment, the executable instructions that cause the computer to generate the camera image using the image data include executable instructions that cause the computer to perform exposure stacking and color correction on the image data to generate the camera image. In an embodiment, the computer-readable storage medium further includes executable instructions that cause the computer to: generate one or more control points in the intensity image; and receive adjustments to the alignment of the intensity image and the camera image.
[0010] Other features and advantages will be apparent from this specification, which illustrates various aspects of this disclosure by way of example. Attached Figure Description
[0011] Details of this disclosure (regarding both its structure and operation) can be partially gathered by studying the accompanying drawings, in which the same reference numerals refer to the same parts, and wherein:
[0012] Figure 1A This is a flowchart of a method for volumetric image alignment and coloring according to one embodiment of the present disclosure;
[0013] Figure 1B This is a flowchart of a method for volumetric image alignment and coloring according to another embodiment of the present disclosure;
[0014] Figure 1C It shows Figure 1B The process description of the steps in the document;
[0015] Figure 2 This is a block diagram of a system for volumetric image alignment and coloring according to one embodiment of the present disclosure;
[0016] Figure 3A This refers to a computer system and user representation according to embodiments of the present disclosure; and
[0017] Figure 3B This is a functional block diagram illustrating a managed alignment and coloring application according to embodiments of the present disclosure. Detailed Implementation
[0018] As mentioned above, video systems used in video production, studio environments, or virtual production can reproduce scenes using a combination of laser scanners and color photography. This can involve manually matching points between a color photograph and an intensity image generated from a laser scan using an image alignment procedure. However, manually defining and matching corresponding points in two images is tedious and time-consuming.
[0019] Certain embodiments of this disclosure provide systems and methods for implementing techniques for processing video data. In one embodiment, the video system captures video data of objects and the environment and creates a volumetric dataset with the colors of points. In another embodiment, the system automatically assigns colors to the points.
[0020] After reading the following description, it will become apparent how this disclosure can be implemented in various ways and applications. Although various embodiments of this disclosure will be described herein, it should be understood that these embodiments are presented by way of example only and not by way of limitation. Therefore, the specific description of the various embodiments should not be construed as limiting the scope or breadth of this disclosure.
[0021] In one implementation of the new system, a light detection and ranging (LIDAR) scanner (e.g., a laser scanner) is used to produce a volumetric point cloud (3-D) represented by intensity data without color. However, to replicate a real-world environment scanned using a LIDAR scanner, a high dynamic range (HDR) photo (color image) of the environment is taken and mapped as a 2-D image representation of the LIDAR scan intensity data.
[0022] In one implementation, the new system captures intensity information with one or more LIDAR scanners and captures images and color information with one or more cameras. The new system uses image feature detection in both the color photo and the intensity image and maps, aligns, and / or matches (collectively, "aligns") the two cross-spectral images (e.g., intensity and color images) together. The system can also allow for fine-tuning of detected points to improve alignment.
[0023] In one implementation, the image feature detection method includes automatic detection, matching, and alignment of features within a color photo to features within an intensity image from a real-world LIDAR scan. The image feature detection method also includes a step of correcting for lens distortion during the alignment process. The method also includes manual fine-tuning, adjusting alignment, and utilizing asynchronous compute shaders for a graphics processing unit (GPU) for computation and shading.
[0024] Figure 1A is a flowchart of a method 100 for volumetric image alignment and shading according to one implementation of the present disclosure. In one implementation, the method includes scene reconstruction using a LIDAR scanner and HDR photography. Thus, such scene reconstruction includes scanning an environment from multiple positions using a LIDAR scanner (to produce 3-D intensity data of the scene) and taking a scene from multiple positions in multiple exposures using an HDR camera (to produce 2-D color image data of the scene).
[0025] In one implementation, 2-D intensity images are generated from the 3-D intensity data of the scene. The 2-D color image data is exposed overlaid and color corrected to generate an HDR image corresponding to the intensity image of a particular view of the scene scanned by the LIDAR scanner. Features in the HDR image are then matched to features in the 2-D intensity image. Feature detection techniques include histogram feature detection, edge-based feature detection, spectral feature detection, and co-occurring feature detection, among others.
[0026] In Figure 1AIn the illustrated embodiment, at step 110, an intensity image is generated from intensity data captured using a scanner. In one embodiment, the intensity image includes at least one feature in the scene and the at least one feature includes a sample feature. At step 112, an HDR camera is used to capture image data, which includes color information. At step 114, the image data is used to generate a camera image, which includes the sample feature of the scene. At step 116, the sample feature in the intensity image is matched to the sample feature in the camera image to align the intensity image and the camera image. At step 118, a color intensity image is generated by applying the color information to the aligned intensity image.
[0027] In further embodiments, during and after the alignment process, any lens distortion is corrected and control points available for fine-tuning the alignment are presented using a GPU. In another embodiment, the color is applied directly to the LIDAR 3-D intensity data.
[0028] In an alternative embodiment of the volumetric image alignment and shading method, the matching and alignment operations are done "offline" on a high-performance cloud cluster, to which users can upload 2-D LIDAR intensity images together with HDR photography. The cloud cluster then performs the above-described process and presents the final results to the users.
[0029] Figure 1B is a flowchart of a method 120 for volumetric image alignment and shading according to another embodiment of the present disclosure. Figure 1C A process illustration of the steps described in Figure 1B is shown.
[0030] In Figure 1B and Figure 1C illustrative embodiments, at step 130, two images are received, a 2-D color image 160 and a 2-D intensity image 162. The 2-D color image 160 is then converted to a 2-D grayscale image 164 at step 132 (see process 170 in Figure 1C At step 134, first pass shape detection is performed on both the 2-D grayscale image 164 and the 2-D intensity image 162 (see process 172 in Figure 1C With the detected shapes, at step 136, first pass feature detection is performed to find matching "shape points" in both images. In Figure 1C the example, the color image is not directly "matched" to the intensity image (i.e., the color image is slightly rotated (in P coordinates) relative to the intensity image (in Q coordinates)). Therefore, at step 138, a phase correlation between the images is computed to account for any offset (see process 174 in Figure 1CThe process 174) in FIG. 1. At step 140, an image transform is computed using the phase correlation data, and at step 142, the 2-D color image is transformed using the image transform so that the transformed 2-D color image 166 has the same orientation as the 2-D intensity image.
[0031] In one implementation, the method 120 further includes performing an initial image alignment at step 144, and a more refined feature detection at step 146. For example, the initial image alignment can include pre-processing for edges, and the more refined feature detection can include feeding the pre-processing results to a Hough algorithm or other line / feature detection. At step 148, matching feature points are extracted from the overlapping matched features based on some heuristic properties such as "acute angle matching" or other feature matching. At step 150, a final image transform is computed using the shape points found in step 136, the feature points extracted in step 148, and the phase correlation computed in step 138. Finally, at step 152, the 2-D color image (source image) is aligned with the 2-D intensity image (target image) using the final image transform (see Figure 2 the process 176) in FIG. 1.
[0032] Figure 2 is a block diagram of a system 200 for volumetric image alignment and painting according to one implementation of the present disclosure. In Figure 3A In the illustrated implementation, the system 200 is used for video production, a studio environment, or virtual production to re-stage a scene, and includes a camera 210 for image capture, a scanner and / or sensor 220, and a processor 230 for processing camera and sensor data. In one implementation, the camera 210 includes an HDR camera, and the scanner / sensor 220 includes a LIDAR scanner. Thus, in one particular implementation, such re-creation of a scene includes scanning the scene using a LIDAR scanner 220 to produce 3-D intensity data of the scene and taking a picture of the scene using an HDR camera to produce 2-D color image data of the scene. Although the illustrated implementation shows only one camera and one scanner, one or more additional cameras and / or one or more additional scanners can be used.
[0033] In one implementation, processor 230 receives intensity data captured by LIDAR scanner 220 and generates an intensity image. In one implementation, the intensity image includes features of the scene (including sample features). Processor 230 also receives image data (including color information) captured by HDR camera 210 and generates a camera image (including sample features). Processor 230 matches the sample features in the intensity image with the sample features in the camera image to align the intensity image and the camera image. Processor 230 generates a color intensity image by applying color information to the aligned intensity image.
[0034] In a further embodiment, processor 230 also performs correction for any lens distortion and controls points that can be used for fine-tuning alignment. In one embodiment, processor 230 is configured as a GPU. In an alternative embodiment, processor 230 outsources the matching and alignment operations by uploading 2-D LIDAR intensity images along with HDR photography to a high-performance cloud cluster. The cloud cluster then performs the above process and sends the final result back to processor 230.
[0035] Figure 1A This is a representation of a computer system 300 and a user 302 according to embodiments of the present disclosure. User 302 uses the computer system 300 to implement functions such as those described above. Figure 1B and Figure 2 Methods 100, 120 and... Figure 3B The system 200 describes and illustrates the alignment and coloring of volumetric images, and the alignment and coloring application 390.
[0036] Computer system 300 Store and Execute Figure 3B Alignment and coloring application 390. Furthermore, computer system 300 can communicate with software program 304. Software program 304 may include software code for alignment and coloring application 390. As will be further explained below, software program 304 may be loaded onto external media (such as CD, DVD, or storage drive).
[0037] Furthermore, computer system 300 can be connected to network 380. Network 380 can be connected in various different architectures (e.g., client-server architecture, peer-to-peer network architecture, or other types of architecture). For example, network 380 can communicate with server 385, which coordinates the engine and data used within the alignment and coloring application 390. Moreover, the network can be of different types. For example, network 380 can be the Internet, a local area network (LAN) or any variant of a LAN, a wide area network (WAN), a metropolitan area network (MAN), an intranet or extranet, or a wireless network.
[0038] Figure 3Bis a functional block diagram illustrating a computer system 300 hosting an alignment and shading application 390 in accordance with an embodiment of the present disclosure. The controller 310 is a programmable processor and controls the operation of the computer system 300 and its components. The controller 310 loads instructions (e.g., in the form of a computer program) from the memory 320 or embedded controller memory (not shown) and executes the instructions to control the system, such as to provide data processing to establish depth and rendering data to render visualizations. In its execution, the controller 310 provides a software system to the alignment and shading application 390, such as to implement creation of device groups and transmission of device setup data in parallel using task queues. Alternatively, this service can be implemented as a separate hardware component in the controller 310 or computer system 300.
[0039] The memory 320 temporarily stores data for use by other components of the computer system 300.
[0040] In one embodiment, the memory 320 is implemented as RAM. In one embodiment, the memory 320 also includes long-term or permanent memory, such as flash memory and / or ROM.
[0041] The storage 330 temporarily or long-term stores data for use by other components of the computer system 300. For example, the storage 330 stores data used by the alignment and shading application 390.
[0042] In one embodiment, the storage 330 is a hard disk drive.
[0043] The media device 340 receives removable media and reads and / or writes data to the inserted media. In one embodiment, for example, the media device 340 is an optical disk drive.
[0044] The user interface 350 includes components to accept user input from a user of the computer system 300 and to present information to the user 302. In one embodiment, the user interface 350 includes a keyboard, a mouse, an audio speaker, and a display. The controller 310 uses input from the user 302 to adjust the operation of the computer system 300.
[0045] The I / O interface 360 includes one or more I / O ports to connect to corresponding I / O devices, such as external storage or supplemental devices (e.g., a printer or a PDA). In one embodiment, the ports of the I / O interface 360 include ports such as: USB ports, PCMCIA ports, serial ports, and / or parallel ports. In another embodiment, the I / O interface 360 includes a wireless interface for wireless communication with external devices.
[0046] Network interface 370 includes wired and / or wireless network connections (such as an RJ-45 supporting an Ethernet connection or a "Wi-Fi" interface (including but not limited to 802.11) supporting a wireless connection).
[0047] Computer system 300 includes additional hardware and software typically found in computer systems of this type (e.g., power supplies, cooling, operating systems), although for simplicity, Figure 3B These components are not specifically shown in FIG. 3. In other embodiments, different configurations of computer systems (e.g., different bus or storage configurations or multi-processor configurations) can be used.
[0048] The description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the disclosure. Many modifications to these embodiments will be readily apparent to those skilled in the art, and the principles defined herein can be applied to other embodiments without departing from the spirit or scope of the disclosure. For example, implementations of the system and method can be applied and adapted to other applications besides video production for movies or television (such as virtual production (e.g., virtual reality environments) or other LIDAR or 3D point space shading applications). Thus, the disclosure is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0049] In particular embodiments of the disclosure, not all features of the examples discussed above are necessarily required. Moreover, it is to be understood that the specific embodiments and figures presented herein represent a broad inventive subject matter. It is also to be understood that the scope of the disclosure fully encompasses other embodiments that can become apparent to those skilled in the art and that the scope of the present disclosure is solely limited by the claims that follow.
Claims
1. A method for image alignment and colorization, the method comprising: receiving a 2-D color image and a 2-D intensity image; converting the 2-D color image to a 2-D grayscale image; performing shape detection on the 2-D grayscale image and the 2-D intensity image to obtain detected shapes; performing feature detection on the detected shapes to find matching shape points in the 2-D grayscale image and the 2-D intensity image; computing a phase correlation between the 2-D grayscale image and the 2-D intensity image; transforming the 2-D color image using the phase correlation such that the transformed 2-D color image has the same orientation as the 2-D intensity image; performing initial image alignment including preprocessing of edges to obtain preprocessed results; performing finer feature detection based on the initial image alignment including feeding the preprocessed results to a line / feature detection process to produce matching features; extracting matching feature points from the matching features; computing a final image transformation using the matching shape points, the phase correlation, and the extracted matching feature points; and aligning the 2-D color image with the 2-D intensity image using the final image transformation.
2. The method of claim 1, further comprising: applying color information on the 3-D intensity image.
3. The method of claim 1, further comprising: generating the 2-D intensity image from the 3-D intensity image.
4. The method of claim 1, further comprising: performing exposure overlay and color correction on the 2-D color image.
5. The method of claim 1, further comprising: generating one or more control points in the 2-D intensity image; and receiving adjustments to the alignment of the 2-D intensity image and the 2-D color image.
6. The method of claim 1, further comprising: correcting any lens distortion in the 2-D color image; and receiving adjustments to the alignment of the 2-D intensity image and the 2-D color image.
7. A system for aligning and colorizing volumetric images, the system comprising: at least one scanner for capturing a 3-D intensity image; at least one camera for capturing a 2-D color image including color information; and a processor for: converting the 2-D color image to a 2-D grayscale image; generating the 2-D intensity image from the 3-D intensity image; performing shape detection on the 2-D grayscale image and the 2-D intensity image to obtain detected shapes; performing feature detection on the detected shapes to find matching shape points in the 2-D grayscale image and the 2-D intensity image; computing a phase correlation between the 2-D grayscale image and the 2-D intensity image; transforming the 2-D color image using the phase correlation such that the transformed 2-D color image has the same orientation as the 2-D intensity image; performing initial image alignment including preprocessing of edges to obtain preprocessed results; performing finer feature detection based on the initial image alignment including feeding the preprocessed results to a line / feature detection process to produce matching features; extracting matching feature points from the matching features; computing a final image transformation using the matching shape points, the phase correlation, and the extracted matching feature points; and aligning the 2-D color image with the 2-D intensity image using the final image transformation. aligning the 2-D color image with the 2-D intensity image using the final image transform.
8. The system of claim 7, wherein, The at least one scanner includes at least one LIDAR scanner.
9. The system of claim 7, wherein, The at least one camera includes at least one HDR camera.
10. The system of claim 7, further comprising: a cloud cluster to receive 2-D intensity images and 2-D color images from the processor, perform matching and alignment, and send results back to the processor.
11. A non-transitory computer-readable storage medium storing a computer program for aligning and shading volumetric images, the computer program comprising executable instructions to cause a computer to: receive 2-D color images and 2-D intensity images; convert the 2-D color images to 2-D grayscale images; perform shape detection on the 2-D grayscale images and the 2-D intensity images to obtain detected shapes; perform feature detection on the detected shapes to find matching shape points in the 2-D grayscale images and the 2-D intensity images; compute a phase correlation between the 2-D grayscale images and the 2-D intensity images; transform the 2-D color images using the phase correlation such that the transformed 2-D color images have the same orientation as the 2-D intensity images; perform initial image alignment including preprocessing of edges to obtain preprocessing results; perform more refined feature detection based on the initial image alignment, including feeding the preprocessing results to line / feature detection processing to produce matching features; extract matching feature points from the matching features; compute a final image transform using the matching shape points, the phase correlation, and the extracted matching feature points; and align the 2-D color images with the 2-D intensity images using the final image transform.
12. The non-transitory computer-readable storage medium of claim 11, further comprising executable instructions to cause the computer to generate 2-D intensity images from 3-D intensity images.
13. The non-transitory computer-readable storage medium of claim 11, further comprising: executable instructions to cause a computer to perform exposure stacking and color correction on the 2-D color images.
14. The non-transitory computer-readable storage medium of claim 11, further comprising executable instructions to cause the computer to: generate one or more control points in the 2-D intensity images; and receive adjustments to the alignment of the 2-D intensity images and the 2-D color images.