Image processing device and image processing method

The integration of high-resolution visible light images with 3D scanner data enhances 3D model quality by correcting scanner viewpoints and textures, addressing the sparsity of depth sensor measurements.

WO2026034226A1PCT designated stage Publication Date: 2026-02-12SONY GROUP CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/026311
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-09
Filing Date
2025-07-24
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Existing 3D modeling technologies face challenges in generating high-quality 3D models due to the sparsity of distance measurement points from depth sensors, leading to low-quality textures where segment identification fails, especially when applying visible light images.

Method used

An image processing device and method that utilize a 3D scanner and a visible light camera to generate and correct 3D shape data by integrating high-resolution visible light images as textures onto the 3D shape, enhancing the accuracy of scanner viewpoints and integrating 3D point clouds based on RGB images.

Benefits of technology

Improves the quality of 3D models by accurately restoring 3D shapes and textures, ensuring high-density point clouds and detailed representations even with low-resolution scanners, resulting in high-quality 3D models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025026311_12022026_PF_FP_ABST
    Figure JP2025026311_12022026_PF_FP_ABST
Patent Text Reader

Abstract

The present technology relates to an image processing device and an image processing method that enable generation of a high-quality 3D model. An image processing device according to the present technology comprises: a first shape restoration unit that restores the shape of a subject on the basis of distance information based on a viewpoint of a depth sensor, and thereby generates first shape data representing a 3D shape of the subject; and a shape correction unit that corrects the 3D shape represented by the first shape data, on the basis of a visible light image mapped as a texture on the surface of the 3D shape represented by the first shape data. The visible light image is an image obtained by a visible light sensor capturing the subject at a viewpoint corresponding to the viewpoint of the depth sensor, and is an image with higher resolution than the resolution of the distance information. The present technology can be applied to, for example, a 3D modeling system.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing device and image processing method

[0001] The present technology relates to an image processing device and an image processing method, and more particularly to an image processing device and an image processing method that are capable of generating high-quality 3D models.

[0002] A technology has been proposed that generates a 3D model of a subject by using a 3D scanner to generate shape data that indicates the 3D shape of the subject, and then pasting the RGB image obtained by photographing the subject with an RGB camera onto the surface of the 3D shape indicated by the shape data.

[0003] For example, Patent Document 1 describes a technology that performs segmentation to identify segments in each of the shape data and the RGB image, aligns the shape data and the RGB image based on the segments that are common between the shape data and the RGB image, and then pastes the RGB image onto the surface of the 3D shape represented by the shape data.

[0004] Japanese Patent Application Laid-Open No. 2021-189600

[0005] Generally, because the distance measurement points at which a depth sensor, including a 3D scanner, can measure distances are relatively sparse, it is difficult to improve the quality of the 3D shape represented by shape data generated using a depth sensor. Furthermore, the technology described in Patent Document 1 cannot apply a visible light image (RGB image) to portions of the 3D shape represented by the shape data where segments cannot be identified, which may result in a low quality texture for the 3D model.

[0006] This technology was developed in light of these circumstances, and makes it possible to generate high-quality 3D models.

[0007] According to one aspect of the present technology, an image processing device includes: a first shape restoration unit that generates first shape data indicating a 3D shape of the subject by performing shape restoration on the basis of distance information based on a viewpoint of a depth sensor; and a shape correction unit that corrects the 3D shape indicated by the first shape data on the basis of a visible light image that is mapped as a texture onto a surface of the 3D shape indicated by the first shape data. The visible light image is an image obtained by capturing an image of the subject with a visible light sensor at a viewpoint corresponding to the viewpoint of the depth sensor, and has a higher resolution than the resolution of the distance information.

[0008] An image processing method according to one aspect of the present technology includes: generating first shape data representing a 3D shape of a subject by performing shape reconstruction on the basis of distance information based on a viewpoint of a depth sensor; and correcting the 3D shape represented by the first shape data on the basis of a visible light image that is mapped as a texture onto a surface of the 3D shape represented by the first shape data. The visible light image is an image obtained by capturing an image of the subject with a visible light sensor at a viewpoint corresponding to the viewpoint of the depth sensor, and has a higher resolution than the resolution of the distance information.

[0009] In one aspect of the present technology, first shape data indicating a 3D shape of the subject is generated by performing shape restoration based on distance information relative to a viewpoint of a depth sensor, and the 3D shape indicated by the first shape data is corrected based on a visible light image that is mapped as a texture onto a surface of the 3D shape indicated by the first shape data. The visible light image is an image obtained by capturing an image of the subject with a visible light sensor at a viewpoint corresponding to the viewpoint of the depth sensor, and has a higher resolution than the resolution of the distance information.

[0010] 1 is a block diagram showing an example of the configuration of a 3D modeling system according to an embodiment of the present technology; FIG. 2 is a diagram showing an example of the viewpoint of an RGB camera; FIG. 3 is a diagram explaining the flow of conventional 3D modeling; FIG. 4 is a block diagram showing an example of the configuration of an image processing device; FIG. 5 is a diagram explaining an example of the flow of camera viewpoint estimation by a viewpoint estimation unit and shape restoration by a shape restoration unit; FIG. 6 is a diagram showing an example of shape data generated based on scan data by a shape restoration unit; FIG. 7 is a diagram explaining a method of correcting a 3D shape indicated by shape data; FIG. 8 is a diagram showing examples of shape data based on scan data before correction and shape data based on scan data after correction; FIG. 9 is a diagram showing correction of a UV map; FIG. 10 is a diagram showing an example of texture correction; FIG. 11 is a flowchart showing processing performed by a 3D modeling system; FIG. 12 is a diagram comparing shape data generated by conventional 3D modeling with shape data generated by 3D modeling of the present technology; FIG. 13 is a block diagram showing a modified example of the configuration of an image processing device; FIG. 14 is a diagram explaining that a user can be involved in 3D modeling of the present technology; FIG. 15 is a diagram explaining the flow of user involvement in correcting shape data based on scan data;

[0011] Hereinafter, embodiments of the present technology will be described in the following order: 1. Configuration of 3D modeling system 2. Operation of 3D modeling system 3. Modified examples

[0012] 1. Configuration of 3D Modeling System FIG. 1 is a block diagram showing an example configuration of a 3D modeling system according to an embodiment of the present technology.

[0013] 1, the 3D modeling system of the present technology includes a 3D scanner 1, an RGB camera 2, and an image processing device 3. The 3D scanner 1 and the RGB camera 2 are examples of the depth sensor and the visible light sensor, respectively, in the present disclosure.

[0014] The 3D scanner 1 is configured as a depth sensor that uses, for example, a stereo camera system, a structured light system, or a ToF (Time of Flight) system. The 3D scanner 1 measures the distance to a subject and acquires distance information. The 3D scanner 1 acquires distance information about the same subject as the subject photographed by the RGB camera 2.

[0015] Hereinafter, the points on the surface of the subject at which the distance to the 3D scanner 1 is measured will be referred to as ranging points. Generally, the resolution of the distance information acquired by a 3D scanner (the number of ranging points) is lower than the resolution (number of pixels) of the RGB image captured by an RGB camera. The 3D scanner 1 supplies distance information indicating the distance between the 3D scanner 1 and each ranging point to the image processing device 3 as scan data.

[0016] The RGB camera 2 captures an image of a subject to obtain an RGB image (visible light image), and supplies the image to the image processing device 3 .

[0017] FIG. 2 is a diagram showing an example of the viewpoint of the RGB camera 2.

[0018] 2, the RGB camera 2 captures images of the object Obj1, such as a plastic model, from multiple viewpoints 2-1 to 2-n that are set to surround the object Obj1. The 3D scanner 1 measures the distance to the object Obj1 from viewpoints corresponding to the viewpoints 2-1 to 2-n, respectively (for example, viewpoints near each of the viewpoints 2-1 to 2-n).

[0019] In the following, for simplicity of explanation, it is assumed that the viewpoint of the 3D scanner 1 and the viewpoint of the RGB camera 2 correspond one-to-one, but this does not necessarily have to be the case. For example, there may be a relationship in which one viewpoint of the 3D scanner 1 corresponds to multiple viewpoints of the RGB camera 2, or the RGB camera 2 may capture an image of the subject at a viewpoint that is not corresponded to the viewpoint of the 3D scanner 1.

[0020] A user of the 3D modeling system takes photographs and measures distances while moving, for example, one 3D scanner 1 and one RGB camera 2. If the 3D scanner 1 and the RGB camera 2 are fixed on a turntable, the user can take photographs and measure distances while moving the 3D scanner 1 and the RGB camera so as to rotate around the subject Obj1. Multiple 3D scanners 1 and multiple RGB cameras 2 may be installed to surround the subject Obj1, and photographs and distances may be taken from each viewpoint.

[0021] Returning to Fig. 1, the image processing device 3 is configured with a PC, a cloud server, etc. The image processing device 3 performs shape reconstruction to reconstruct the 3D shape of the subject based on multiple scan data based on multiple viewpoints supplied from the 3D scanner 1, and generates shape data indicating the 3D shape of the subject. The image processing device 3 corrects the shape data generated based on the scan data based on RGB images from multiple viewpoints supplied from the RGB camera 2.

[0022] The image processing device 3 generates a 3D model that represents the 3D shape, color, pattern, texture, etc. of the subject by pasting RGB images from multiple viewpoints onto the surface of the 3D shape represented by the corrected shape data (texture mapping). The 3D model is composed of shape data that represents the 3D shape (geometry information) of the subject using, for example, a polygon mesh or a 3D point cloud, and a UV map that maps the texture of the subject in, for example, two dimensions.

[0023] FIG. 3 is a diagram illustrating the flow of conventional 3D modeling.

[0024] In conventional 3D modeling, segmentation is first performed to identify segments in each of the shape data and RGB image, and then the shape data and RGB image are aligned (viewpoint alignment) using the segments common to both the shape data and RGB image as a reference, as shown in #1 in Figure 3. Next, as shown in #2 in Figure 3, texture mapping is performed to attach the RGB image for each segment to the surface of the 3D shape indicated by the shape data, and a 3D model is generated.

[0025] Generally, because the measurement points of a 3D scanner are relatively sparse, it is difficult to improve the quality of the 3D shape represented by the shape data generated using a 3D scanner. In conventional 3D modeling, it is not possible to attach RGB images to parts of the 3D shape represented by the shape data where segments cannot be identified, which can result in low quality textures for the 3D model.

[0026] This technology was developed with the above points in mind, and enables the generation of high-quality 3D models by correcting shape data generated using a 3D scanner 1 based on RGB images.

[0027] FIG. 4 is a block diagram showing an example of the configuration of the image processing device 3.

[0028] As shown in FIG. 4, the image processing device 3 includes a viewpoint estimation unit 11, shape restoration units 12 and 13, a shape correction unit 14, and a texture correction unit 15.

[0029] The viewpoint estimation unit 11 estimates the camera viewpoint, which is the viewpoint of the RGB camera 2 at the time of shooting, based on the scan data (distance information) acquired by the scanner 1 and the RGB image acquired by the RGB camera 2, and supplies viewpoint information indicating the camera viewpoint to the shape restoration units 12 and 13 and the texture correction unit 15.

[0030] Here, the viewpoint includes the position and orientation of the scanner 1 and the RGB camera 2, and the viewpoint information includes, for example, three-dimensional coordinates indicating the position of the RGB camera 2 and a three-dimensional vector indicating the orientation of the RGB camera 2. Since the correspondence between the camera viewpoint and the scanner viewpoint, which is the viewpoint of the scanner 1 during distance measurement, is known due to prior calibration, the viewpoint information can also be said to be information indicating the scanner viewpoint.

[0031] The shape restoration unit 12 generates shape data by performing shape restoration based on the scan data. Specifically, the shape restoration unit 12 converts distance information as scan data into a 3D point cloud indicating the three-dimensional distribution of each ranging point, and generates shape data by integrating multiple 3D point clouds corresponding to multiple scanner viewpoints based on viewpoint information corresponding to each of the multiple scanner viewpoints.

[0032] FIG. 5 is a diagram illustrating an example of the flow of the camera viewpoint estimation by the viewpoint estimation unit 11 and the shape restoration by the shape restoration unit 12. In FIG.

[0033] 5, the viewpoint estimation unit 11 estimates the camera viewpoint of each RGB image in a predetermined reference space based on RGB images from multiple camera viewpoints. Here, the reference space is a space in which a reference distance (scale) is set arbitrarily, for example, by setting the distance between the camera viewpoints of two specific RGB images to 1.

[0034] 5, the viewpoint estimation unit 11 performs scale conversion to make the scale of the reference space the same as the scale of the real space. Viewpoint information indicating the viewpoints of each camera after the scale conversion is supplied to the shape restoration unit 12.

[0035] In FIG. 5, the shape restoration unit 12 performs processes shown in #13 and #14 in parallel with the processes shown in #11 and #12.

[0036] As shown in #13 of FIG. 5, the shape restoration unit 12 converts the distance information at the scanner viewpoint into distance information at the camera viewpoint based on the calibration data indicating the correspondence between the scanner viewpoint and the camera viewpoint.

[0037] Next, as shown in #14 of FIG. 5, the shape restoration unit 12 converts the distance information from the camera viewpoint into a 3D point cloud.

[0038] After the processing shown in #12 and #14 in Figure 5 is performed, as shown in #15 in Figure 5, the shape restoration unit 12 integrates the 3D point clouds for each camera viewpoint based on the viewpoint information corresponding to each camera viewpoint to generate shape data.

[0039] Accurate reconstruction of the 3D shape of a subject requires not only accurate distance information from the scan data, but also accurate integration of the 3D point clouds corresponding to each scanner viewpoint. The accuracy of integration of the point clouds corresponding to each scanner viewpoint is determined, for example, by the accuracy of the scanner viewpoint estimation. Because the resolution of scan data is generally relatively low, estimating the scanner viewpoint based on multiple scan data sets can result in low accuracy, potentially degrading the quality of the 3D shape represented by the shape data. Therefore, accurate reconstruction of the 3D shape of a subject requires highly accurate estimation of the scanner viewpoint.

[0040] In the viewpoint estimation unit 11 of the present technology, the scanner viewpoint is estimated based on an RGB image having a higher resolution than the scan data, and therefore the scanner viewpoint can be estimated with higher accuracy than when the scanner viewpoint is estimated based on the scan data. Because the scanner viewpoint is estimated with high accuracy, the shape restoration unit 12 can more accurately restore the 3D shape of the subject based on the scan data.

[0041] Fig. 6 is a diagram showing an example of shape data generated based on scan data by the shape restoration unit 12. In the example of Fig. 6, the 3D shape of the motorcycle is represented by a 3D point cloud.

[0042] Even when an inexpensive 3D scanner 1 with low scan data resolution is used for 3D modeling, the shape restoration unit 12 can integrate multiple 3D point clouds based on each scanner viewpoint estimated with high accuracy based on an RGB image with a higher resolution than the scan data, thereby generating shape data of the bike with an accurately restored 3D shape, as shown in Figure 6.

[0043] The viewpoint estimation unit 11 can also be said to be a shape correction unit that improves (corrects) the quality of the 3D shape represented by the shape data generated by the shape restoration unit 12 by having the shape restoration unit 12 perform shape restoration based on highly accurate viewpoint information.

[0044] Returning to FIG. 4, the shape restoration unit 12 supplies the generated shape data to the shape correction unit 14 .

[0045] The shape restoration unit 13 generates shape data by performing shape restoration (photogrammetry) based on the RGB images acquired by the RGB camera 2. Specifically, the shape restoration unit 13 acquires distance information based on each camera viewpoint based on the parallax information of the two RGB images, and converts the distance information into a 3D point cloud indicating the three-dimensional distribution of each ranging point (points on the surface of the subject reflected in each pixel of the RGB image). The shape restoration unit 13 generates shape data by integrating multiple 3D point clouds corresponding to each camera viewpoint based on the viewpoint information corresponding to each camera viewpoint.

[0046] In order to accurately reconstruct the 3D shape of a subject, it is necessary not only to accurately obtain distance information based on the disparity information of the two RGB images, but also to accurately integrate the 3D point clouds corresponding to each camera viewpoint, in other words, to estimate each camera viewpoint with high accuracy.

[0047] The viewpoint estimation unit 11 can improve the accuracy of the camera viewpoint estimation by correcting the estimated camera viewpoint using feature points included in the scan data. Because the camera viewpoint is estimated with high accuracy, the shape restoration unit 13 can more accurately restore the 3D shape of the subject based on the RGB image.

[0048] In addition, the shape restoration unit 13 can also generate shape data by performing image processing called SfM (Structure from Motion) and image processing called MVS (Multiview Stereo) using RGB images from multiple camera viewpoints.

[0049] The shape restoration unit 13 supplies the generated shape data to the shape correction unit 14 .

[0050] The shape correction unit 14 corrects the 3D shape indicated by the shape data (shape data based on scan data) supplied from the shape restoration unit 12 based on the 3D shape indicated by the shape data (shape data based on RGB image) supplied from the shape restoration unit 13.

[0051] FIG. 7 is a diagram illustrating a method for correcting a 3D shape indicated by shape data.

[0052] 7, the shape correction unit 14 adopts the shape data based on the scan data as the shape data for a portion of the 3D shape of the subject for which shape reconstruction based on the scan data was successful but shape reconstruction based on the RGB image was unsuccessful. The shape correction unit 14 adopts the shape data based on the RGB image as the shape data for a portion of the 3D shape of the subject for which shape reconstruction based on the RGB image was successful but shape reconstruction based on the scan data was unsuccessful.

[0053] The shape correction unit 14 uses both the shape data based on the scan data and the shape data based on the RGB image as shape data for a portion of the 3D shape of the subject that has been successfully reconstructed both based on the scan data and the RGB image. Specifically, the shape correction unit 14 increases the 3D point density and the number of polygon meshes for that portion of the shape data based on the scan data, based on the shape data based on the RGB image.

[0054] The shape correction unit 14 does not process any part of the 3D shape of the subject for which shape reconstruction based on both the scan data and the RGB image has failed.

[0055] Whether or not shape restoration is successful is determined based on, for example, the results of a comparison between the 3D shape indicated by the shape data based on the scan data and the 3D shape indicated by the shape data based on the RGB image, and the variation in the distribution of 3D points and polygon meshes that represent the 3D shape in each of the shape data based on the scan data and the shape data based on the RGB image.

[0056] For example, a part where 3D points or polygon meshes that represent a 3D shape are arranged at approximately the same positions in the shape data based on the scan data and the shape data based on the RGB image, and where the 3D points or polygon meshes are not separated, is determined to be a part where shape reconstruction was successful.On the other hand, a part where 3D points or polygon meshes that represent a 3D shape are arranged at approximately the same positions in the shape data based on the scan data and the shape data based on the RGB image, but where the 3D points or polygon meshes are separated, is determined to be a part where shape reconstruction failed.

[0057] A portion in either the shape data based on scan data or the shape data based on RGB images where 3D points or polygon meshes that represent a 3D shape are arranged and where the 3D points or polygon meshes are not separated is determined as a portion where shape reconstruction was successful. On the other hand, a portion in either the shape data based on scan data or the shape data based on RGB images where 3D points or polygon meshes that represent a 3D shape are arranged but where the 3D points or polygon meshes are separated is determined as a portion where shape reconstruction failed.

[0058] In summary, the shape correction unit 14 determines a portion of the 3D shape represented by the shape data based on the scan data to be corrected based on the 3D shape represented by the shape data based on the RGB image, based on a comparison result between the 3D shape represented by the shape data based on the scan data and the 3D shape represented by the shape data based on the RGB image. Also, the shape correction unit 14 determines a portion of the 3D shape represented by the shape data based on the scan data to be corrected based on the 3D shape represented by the shape data based on the RGB image, based on the variation in the distribution of 3D points or polygon meshes representing the 3D shape of the subject, in each of the shape data based on the scan data and the shape data based on the RGB image.

[0059] For example, the shape correction unit 14 corrects portions of shape data based on scan data where 3D points or polygon meshes are missing, using 33D points or polygon meshes of corresponding portions in shape data based on an RGB image.

[0060] FIG. 8 is a diagram showing an example of shape data based on scan data before correction and shape data based on scan data after correction.

[0061] FIG. 8A shows shape data based on scan data before correction, and FIG. 8B shows shape data based on scan data after correction.

[0062] In the shape data before correction, the portion indicated by the dashed ellipse in A of Fig. 8 is missing, but in the shape data after correction, this portion is not missing and the 3D shape of the subject is accurately restored. In this way, by correcting shape data based on scan data based on shape data based on an RGB image, it is possible to improve the quality of the shape data.

[0063] Returning to FIG. 4, the shape correcting unit 14 supplies the corrected shape data to the texture correcting unit 15 .

[0064] Based on the viewpoint information supplied from the viewpoint estimation unit 11, the texture correction unit 15 generates a UV map by mapping an RGB image onto the surface of the 3D shape indicated by the shape data supplied from the shape correction unit 14, and combines the shape data and the UV map into a single 3D model.

[0065] Furthermore, the texture correction unit 15 corrects the UM map (texture information) based on the comparison result between a rendering image of the 3D model rendered from the camera viewpoint indicated by the viewpoint information and an RGB image from the camera viewpoint.

[0066] FIG. 9 is a diagram illustrating the correction of the UV map.

[0067] 9, the texture correction unit 15 compares the RGB image Pi1 with a rendered image Pi2 obtained by rendering a 3D model from the camera viewpoint of the RGB image Pi1. If the quality of the texture information included in the UV map is low, a difference of a certain level or more will occur between the RGB image Pi1 and the rendered image Pi2.

[0068] The texture correction unit 15 compares the RGB image Pi1 with the rendering image Pi2, and if there is a difference of a certain level or more, corrects the texture information in the UV map M1 that corresponds to the portion where there is a difference. For example, the texture correction unit 15 replaces the texture information in the UV map M1 that corresponds to the portion where there is a difference with texture information extracted from the RGB image Pi1, or applies a high-frequency emphasis filter to the texture information in the UV map M1 that corresponds to the portion where there is a difference.

[0069] In this way, the texture correction unit 15 corrects the UM map so that the rendering image becomes closer to an RGB image. For example, as shown in Fig. 10, the texture correction unit 15 can improve the quality of the 3D model, especially the texture, by correcting a blurred texture into a high-definition texture.

[0070] Returning to Figure 4, the texture correction unit 15 notifies the viewpoint estimation unit 11 of a request to correct the viewpoint information, and notifies the shape correction unit 14 of a request to correct the shape data, based on the result of comparing the rendering image of the 3D model from the camera viewpoint with the RGB image from the camera viewpoint.

[0071] Specifically, if a difference of a certain level or more remains between the rendering image and the RGB image even after texture correction, the texture correction unit 15 notifies the viewpoint estimation unit 11 of a correction request indicating the camera viewpoint to be corrected and the amount of correction, and acquires viewpoint information indicating the camera viewpoint corrected with the amount of correction from the viewpoint estimation unit 11. In a similar case, the texture correction unit 15 notifies the shape correction unit 14 of a correction request indicating the location of the 3D shape to be corrected and the amount of correction, and acquires shape data indicating the 3D shape corrected with the amount of correction from the shape correction unit 14.

[0072] The texture correction unit 15 remaps the RGB image onto the surface of the 3D shape represented by the shape data before correction in accordance with the correction request, based on the viewpoint information corrected in accordance with the correction request, and compares a rendered image of the 3D model rendered from the camera viewpoint with the RGB image from the camera viewpoint to check whether the difference is small. Similarly, the texture correction unit 15 remaps the RGB image onto the surface of the 3D shape represented by the shape data corrected in accordance with the correction request, based on the viewpoint information before correction in accordance with the correction request, and compares a rendered image of the 3D model rendered from the camera viewpoint with the RGB image from the camera viewpoint to check whether the difference is small.

[0073] The texture correction unit 15 determines the shape data and UV map that results in the smaller difference between the rendering image and the RGB image when the viewpoint information is corrected or when the shape data is corrected as the data that constitutes the final 3D model.

[0074] The texture correction unit 15 notifies a correction request and corrects the viewpoint information and shape data so as to reduce the difference between the rendering image and the RGB image, thereby making it possible to improve the quality of the 3D model.

[0075] In addition, when a request for correction of viewpoint information is notified by the texture correction unit 15, shape restoration by the shape restoration units 12 and 13 and correction of shape data by the shape correction unit 14 may be performed again based on the viewpoint information corrected in accordance with the correction request.

[0076] A part of the configuration of the image processing device 3 may be provided in a device other than the image processing device 3. For example, the processes performed by the shape correction unit 14 and the texture correction unit 15 may be performed by a cloud server.

[0077] 2. Operation of the 3D Modeling System Next, the processing performed by the 3D modeling system having the above-described configuration will be described with reference to the flowchart of FIG.

[0078] In step S1, the viewpoint estimation unit 11 performs calibration. For example, the viewpoint estimation unit 11 calculates the positions of corner points of a calibration board based on scan data acquired by the scanner 1 measuring the distance to the calibration board on which a black and white checkered pattern is formed. The viewpoint estimation unit 11 also calculates the positions of corner points of the calibration board based on an RGB image acquired by the RGB camera 2 by photographing the calibration board. The viewpoint estimation unit 11 calculates the correspondence between the scanner viewpoint and the camera viewpoint based on the calculation results of the positions of corner points based on the scan data and the calculation results of the positions of corner points based on the RGB image.

[0079] In step S2, the RGB camera 2 captures images of the subject from multiple viewpoints. The 3D scanner 1 measures the distance to the subject from multiple scanner viewpoints corresponding to the multiple camera viewpoints.

[0080] In step S3, the viewpoint estimation unit 11 estimates the camera viewpoint of each RGB image based on the multiple scan data acquired by the 3D scanner 1 and the multiple RGB images acquired by the RGB camera 2. The viewpoint estimation unit 11 generates viewpoint information based on the estimation results of the camera viewpoint and the calibration data.

[0081] In step S4, the shape restoration unit 12 generates shape data by performing shape restoration based on the multiple scan data and viewpoint information corresponding to each scanner viewpoint, while the shape restoration unit 13 generates shape data by performing shape restoration based on the multiple RGB images and viewpoint information corresponding to each camera viewpoint.

[0082] In step S5, the image processing device 3 accepts corrections to the shape data in response to, for example, a user's operation on the input unit. Here, the user can delete portions of the shape data based on the scan data and the shape data based on the RGB image where the 3D shapes of objects other than the subject are erroneously restored.

[0083] In step S6, the shape correcting unit 14 corrects the 3D shape indicated by the shape data based on the scan data, based on the 3D shape indicated by the shape data based on the RGB image.

[0084] In step S7, the shape correction unit 14 determines whether the quality of the shape data is sufficient. Here, for example, the user checks the corrected shape data to determine whether the quality of the shape data is sufficient, and inputs the determination result using the input unit.

[0085] If it is determined in step S7 that the quality of the shape data is not sufficient, the process returns to step S5, and the subsequent steps are carried out.

[0086] On the other hand, if it is determined in step S7 that the quality of the shape data is sufficient, in step S8, the texture correction unit 15 maps an RGB image onto the surface of the 3D shape indicated by the corrected shape data based on the viewpoint information, and generates a UV map.

[0087] In step S9, the texture correction unit 15 compares the rendering images of the 3D model rendered from each camera viewpoint indicated by the viewpoint information with the RGB images from each camera viewpoint, and obtains the difference (error) between the rendering images and the RGB images.

[0088] In step S10, the texture correction unit 15 analyzes the error between the rendering image and the RGB image.

[0089] In step S11, the texture correction unit 15 determines whether the quality of the 3D model is sufficient based on the error analysis result.

[0090] If it is determined in step S11 that the quality of the 3D model is insufficient, then in step S12, the texture correction unit 15 corrects at least one of the UV map, viewpoint information, and shape data based on the error analysis results. If the viewpoint information or shape data has been corrected, an RGB image is mapped onto the 3D shape indicated by the shape data. Then, the process returns to step S9, and subsequent processes are performed.

[0091] On the other hand, if it is determined in step S11 that the quality of the 3D model is sufficient, the process ends.

[0092] As described above, in the image processing device 3 of the present technology, shape data (first shape data) indicating the 3D shape of the subject is generated by restoring the shape of the subject based on distance information (scan data) based on the viewpoint of the 3D scanner 1, and the 3D shape indicated by the first shape data is corrected based on an RGB image that is mapped as a texture onto the surface of the 3D shape indicated by the first shape data.

[0093] The accuracy of estimating the scanner viewpoint is improved based on an RGB image having a higher resolution than the scan data, and the quality of the 3D shape represented by the first shape data is improved based on the 3D shape represented by the second shape data based on the RGB image. Furthermore, the quality of the texture is improved based on a comparison between the RGB image and a rendering image of the improved quality 3D shape. In this way, the quality of the data constituting the 3D model is improved based on the RGB image in various 3D modeling processes, making it possible to generate a high-quality 3D model.

[0094] FIG. 12 is a diagram comparing shape data generated by conventional 3D modeling with shape data generated by 3D modeling according to the present technology.

[0095] FIG. 12A shows shape data generated by conventional 3D modeling, and FIG. 12B shows shape data generated by 3D modeling according to the present technology.

[0096] In the shape data generated by conventional 3D modeling, thin parts, curved surfaces, flat surfaces, etc., as shown by the dashed ovals in A of Figure 12, are missing. However, in the shape data generated by the 3D modeling of this technology, these parts are not missing, and the 3D shape of the subject is accurately restored.

[0097] 3. Modifications Example of shape restoration based on cross-polarized images Fig. 13 is a block diagram showing a modification of the configuration of the image processing device 3. In Fig. 13, the same components as those in Fig. 4 are assigned the same reference numerals. Duplicate descriptions will be omitted as appropriate.

[0098] The image processing device 3 in FIG. 13 differs from the image processing device 3 in FIG. 4 in that a cross-polarized image and a parallel-polarized image are input instead of an RGB image.

[0099] A cross-polarized image is an RGB image generated by cross-polarized photography. An RGB image is generated when light from the light-emitting unit of the RGB camera 2 is irradiated onto a subject and the light reflected from the subject is received by an image sensor. In the case of cross-polarized photography, the light from the light-emitting unit and the light reflected from the subject are polarized in directions perpendicular to each other. That is, the light from the light-emitting unit is polarized in a first polarization direction, and the light reflected from the subject is polarized in a second polarization direction perpendicular to the first polarization direction.

[0100] For example, in the case of cross-polarized photography, a light emitting unit emits light, and the light emitted from the light emitting unit is polarized in a first direction by a first polarizing filter or the like disposed between the light emitting unit and the subject, and the polarized light is irradiated onto the subject. The light reflected from the subject is polarized in a second direction perpendicular to the first direction by a second polarizing filter or the like disposed between the subject and the image sensor, and is received by the image sensor.

[0101] Light reflected from an object can be classified into specular reflection (regular reflection) and diffuse reflection. That is, light reflected from an object can contain both specular and diffuse reflection components. In the case of cross-polarized light photography, the polarization direction of the first polarizing filter is perpendicular to the polarization direction of the second polarizing filter, which prevents the image sensor from receiving the specular reflection component of the light irradiated from the light emitter.

[0102] A parallel polarization image is an RGB image generated by parallel polarization photography. An RGB image is generated when light from a light-emitting unit is irradiated onto a subject and the light from the subject is received by an image sensor. In parallel polarization photography, the light from the light-emitting unit and the light reflected from the subject are polarized in the same direction. That is, the light from the light-emitting unit is polarized in a first polarization direction, and the light reflected from the subject is polarized in the first polarization direction.

[0103] In parallel polarization photography, a light emitting unit emits light, and the light emitted from the light emitting unit is polarized in a first direction by a first polarizing filter or the like arranged between the light emitting unit and the subject, and the polarized light is irradiated onto the subject. The light reflected from the subject is polarized in the first direction by a second polarizing filter or the like arranged between the subject and the image sensor, and is received by the image sensor.

[0104] A user can switch between cross-polarized and parallel-polarized image capture by, for example, rotating one of the first and second polarizing filters by 90°. Therefore, a user can obtain cross-polarized and parallel-polarized images from the same camera viewpoint by, for example, fixing the RGB camera 2 on a tripod or the like and taking images before and after rotating the polarizing filter. If the RGB camera 2 is fixed on a turntable, a user can obtain cross-polarized and parallel-polarized images from the same camera viewpoint by taking cross-polarized and parallel-polarized images, respectively, each time the RGB camera 2 is moved.

[0105] In the case of parallel polarization photography, the polarization direction of the second polarizing filter is perpendicular to the polarization direction of the first polarizing filter, so the image sensor receives both specular and diffuse reflection components. In other words, the cross-polarized image has less gloss (specular reflection) than the parallel polarization image.

[0106] The viewpoint estimation unit 11 estimates the camera viewpoint based on the scan data acquired by the scanner 1 and the cross-polarized image captured by the RGB camera 2 .

[0107] The shape restoration unit 13 generates shape data by performing shape restoration based on the cross-polarized image captured by the RGB camera 2 .

[0108] Based on the viewpoint information, the texture correction unit 15 generates a UV map by mapping the parallel polarization image captured by the RGB camera 2 onto the surface of the 3D shape indicated by the shape data supplied from the shape correction unit 14. Based on the result of comparing a rendering image of the 3D model rendered from the camera viewpoint with the parallel polarization image from the camera viewpoint, the texture correction unit 15 corrects at least one of the UV map, viewpoint information, and shape data.

[0109] Because cross-polarized images have a small gloss component, estimating the camera viewpoint and restoring the shape based on cross-polarized images allows for more accurate estimation of the camera viewpoint and shape restoration than when estimating the camera viewpoint and restoring the shape based on RGB images that contain a gloss component. On the other hand, because parallel-polarized images contain a gloss component, mapping parallel-polarized images as texture onto shape data allows for more accurate representation of the texture of the subject than mapping cross-polarized images.

[0110] Example in which the user is involved in correcting shape data The user may select a portion of the 3D shape indicated by the shape data based on the scan data to be corrected based on the 3D shape indicated by the shape data based on the RGB image.

[0111] FIG. 14 is a diagram illustrating the user's involvement in the 3D modeling of the present technology.

[0112] The image processing device 3 presents to the user the 3D shape indicated by the shape data based on the scan data generated by the shape restoration unit 12. As shown in #31 in Fig. 14 , the user can confirm and correct the 3D shape of the shape data based on the scan data presented by the image processing device 3.

[0113] The image processing device 3 presents to the user the 3D shape indicated by the shape data based on the RGB image generated by the shape restoration unit 13. As shown in #32 in Fig. 14 , the user can confirm and modify the 3D shape of the shape data based on RGB presented by the image processing device 3.

[0114] As shown in #33 of Figure 14, when correcting shape data by the shape correction unit 14, the user can select a portion of the 3D shape represented by the shape data based on the scan data to be corrected based on the 3D shape represented by the shape data based on the RGB image.

[0115] FIG. 15 is a diagram illustrating the flow of user involvement in correcting shape data based on scan data.

[0116] As shown in #51 in FIG. 15, the shape correcting unit 14 obtains the difference between the 3D shape indicated by the shape data D1 based on the RGB image and the 3D shape indicated by the shape data D2 based on the scan data.

[0117] As shown in #52 of FIG. 15, the shape corrector 14 presents the user with the difference (comparison result) between the 3D shape indicated by the shape data D1 and the 3D shape indicated by the shape data D2.

[0118] As shown in #53 of Figure 15, the user can check the difference between the 3D shape represented by the shape data D1 presented by the image processing device 3 and the 3D shape represented by the shape data D2, and specify the part of the difference that he or she wants to add to the 3D shape represented by the shape data D2 or delete from the 3D shape represented by the shape data D2.

[0119] As shown in #54 of FIG. 15, the shape correction unit 14 corrects the shape data D2, for example, by combining the 3D shape of the part specified by the user with the 3D shape of the shape data D2 based on the scan data.

[0120] - Example of computer configuration The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the program that constitutes the software is installed from a program recording medium into a computer built into dedicated hardware or a general-purpose personal computer.

[0121] FIG. 16 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.

[0122] A CPU (Central Processing Unit) 501 , a ROM (Read Only Memory) 502 , and a RAM (Random Access Memory) 503 are interconnected by a bus 504 .

[0123] An input / output interface 505 is also connected to the bus 504. An input unit 506 including a keyboard, a mouse, etc., and an output unit 507 including a display, a speaker, etc. are connected to the input / output interface 505. Also connected to the input / output interface 505 are a storage unit 508 including a hard disk, a nonvolatile memory, etc., a communication unit 509 including a network interface, etc., and a drive 510 that drives removable media 511.

[0124] In a computer configured as described above, the CPU 501 performs the above-described series of processes by, for example, loading a program stored in the storage unit 508 into the RAM 503 via the input / output interface 505 and the bus 504 and executing it.

[0125] The program executed by the CPU 501 is installed in the storage unit 508 by being recorded on, for example, a removable medium 511 or provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital broadcasting.

[0126] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.

[0127] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are housed in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.

[0128] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0129] The embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible without departing from the spirit of the present technology.

[0130] For example, the present technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by a plurality of devices via a network.

[0131] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.

[0132] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.

[0133] Example of Combination of Configurations The present technology can also be configured as follows.

[0134] (1) An image processing device comprising: a first shape restoration unit that generates first shape data indicating a 3D shape of the subject by performing shape restoration on the basis of distance information based on a viewpoint of a depth sensor; and a shape correction unit that corrects the 3D shape indicated by the first shape data on the basis of a visible light image that is mapped as a texture onto a surface of the 3D shape indicated by the first shape data, wherein the visible light image is an image acquired by a visible light sensor capturing an image of the subject at a viewpoint corresponding to the viewpoint of the depth sensor and has a higher resolution than a resolution of the distance information. (2) An image processing device according to (1), wherein the shape correction unit estimates the viewpoint of the visible light sensor based on the visible light image, generates viewpoint information indicating at least one of the viewpoint of the depth sensor and the viewpoint of the visible light sensor based on calibration data indicating a correspondence between the viewpoint of the depth sensor and the viewpoint of the visible light sensor and on the estimated result of the viewpoint of the visible light sensor, and corrects the 3D shape indicated by the first shape data on the basis of the viewpoint information. (3) The image processing device according to (2), wherein the shape correction unit causes the first shape restoration unit to restore the shape of the subject based on the viewpoint information. (4) The image processing device according to (3), wherein the first shape restoration unit restores the shape of the subject by integrating multiple pieces of distance information acquired by the depth sensor measuring distances to the subject from multiple viewpoints based on the viewpoint information corresponding to each of the multiple viewpoints of the depth sensor. (5) The image processing device according to (3) or (4), further comprising a second shape restoration unit that generates second shape data indicating the 3D shape of the subject by restoring the shape of the subject based on the visible light image and the viewpoint information, wherein the shape correction unit corrects the 3D shape indicated by the first shape data based on the 3D shape indicated by the second shape data. (6) The image processing device according to (5), wherein the shape correction unit determines a portion of the 3D shape represented by the first shape data to be corrected based on the 3D shape represented by the second shape data based on a comparison result between the 3D shape represented by the first shape data and the 3D shape represented by the second shape data.(7) The image processing device according to (5) or (6), wherein the shape correction unit determines a portion of the 3D shape represented by the first shape data to be corrected based on the 3D shape represented by the second shape data, based on a variation in distribution of 3D points or polygon meshes representing the 3D shape of the subject, in each of the first shape data and the second shape data. (8) The image processing device according to any of (5) to (7), wherein the shape correction unit presents a comparison result between the 3D shape represented by the first shape data and the 3D shape represented by the second shape data to a user, and corrects a portion of the 3D shape represented by the first shape data, specified by the user, based on the 3D shape represented by the second shape data. (9) The image processing device according to (5), wherein the second shape restoration unit restores the shape of the object by acquiring the plurality of pieces of distance information based on each of the plurality of viewpoints of the visible light sensor, based on a plurality of visible light images acquired by the visible light sensor capturing images of the object from a plurality of viewpoints, and integrating the plurality of pieces of distance information based on each of the plurality of viewpoints of the visible light sensor, based on the viewpoint information corresponding to each of the plurality of viewpoints of the visible light sensor. (10) The image processing device according to any of (2) to (9), further comprising a texture correction unit that corrects the texture mapped on the surface of the 3D shape represented by the corrected first shape data, based on the viewpoint information. (11) The image processing device according to (10), wherein the texture correction unit maps the visible light image as the texture on the surface of the 3D shape represented by the corrected first shape data, based on the viewpoint information, and corrects the texture based on a result of comparing a rendering image obtained by rendering the first shape data with the texture-mapped first shape data from the viewpoint of the visible light sensor indicated by the viewpoint information with the visible light image captured from the viewpoint of the visible light sensor. (12) The image processing device according to (11), wherein the texture correction unit corrects at least one of the viewpoint information and the first shape data based on a comparison result between the rendering image and the visible light image.(13) The image processing device according to (11) or (12), wherein the texture correction unit corrects the texture based on the visible light image compared with the rendering image, or applies a high-frequency emphasis filter to the texture. (14) The image processing device according to any of (2) to (13), wherein the shape correction unit estimates a viewpoint of the visible light sensor based on the visible light image and the distance information. (15) The image processing device according to any of (1) to (14), further comprising a second shape restoration unit that generates second shape data indicating the 3D shape of the object by restoring the shape of the object based on the visible light image and the viewpoint information, wherein the shape correction unit corrects the 3D shape indicated by the first shape data based on the 3D shape indicated by the second shape data. (16) The image processing device according to any one of (1) to (15), wherein the shape correction unit corrects the 3D shape indicated by the first shape data based on a cross-polarized image acquired by the visible light sensor by photographing the subject, and a parallel polarization image corresponding to the cross-polarized image is mapped as the texture onto a surface of the 3D shape indicated by the first shape data. (17) An image processing method including: generating first shape data indicating the 3D shape of the subject by performing shape reconstruction on the basis of distance information based on a viewpoint of a depth sensor; and correcting the 3D shape indicated by the first shape data based on a visible light image mapped as a texture onto the surface of the 3D shape indicated by the first shape data, wherein the visible light image is an image acquired by photographing the subject with a visible light sensor at a viewpoint corresponding to the viewpoint of the depth sensor and has a higher resolution than a resolution of the distance information.

[0135] 1 3D scanner, 2 RGB camera, 3 image processing device, 11 viewpoint estimation unit, 12, 13 shape restoration unit, 14 shape correction unit, 15 texture correction unit

Claims

1. An image processing device comprising: a first shape restoration unit that generates first shape data indicating the 3D shape of a subject by restoring the shape of the subject based on distance information based on the viewpoint of a depth sensor; and a shape correction unit that corrects the 3D shape indicated by the first shape data based on a visible light image that is mapped as a texture onto the surface of the 3D shape indicated by the first shape data, wherein the visible light image is an image obtained by a visible light sensor photographing the subject at a viewpoint corresponding to the viewpoint of the depth sensor, and is an image having a higher resolution than the resolution of the distance information.

2. The image processing device described in claim 1, wherein the shape correction unit estimates the viewpoint of the visible light sensor based on the visible light image, generates viewpoint information indicating at least one of the viewpoint of the depth sensor and the viewpoint of the visible light sensor based on calibration data indicating the correspondence between the viewpoint of the depth sensor and the viewpoint of the visible light sensor and the estimated result of the viewpoint of the visible light sensor, and corrects the 3D shape indicated by the first shape data based on the viewpoint information.

3. The image processing device according to claim 2, wherein the shape correction section causes the first shape restoration section to restore the shape of the subject based on the viewpoint information.

4. The image processing device described in claim 3, wherein the first shape restoration unit restores the shape of the subject by integrating multiple pieces of distance information obtained by the depth sensor measuring the distance to the subject from multiple viewpoints based on the viewpoint information corresponding to each of the multiple viewpoints of the depth sensor.

5. The image processing device of claim 3, further comprising a second shape restoration unit that generates second shape data indicating the 3D shape of the subject by restoring the shape of the subject based on the visible light image and the viewpoint information, and wherein the shape correction unit corrects the 3D shape indicated by the first shape data based on the 3D shape indicated by the second shape data.

6. The image processing device described in claim 5, wherein the shape correction unit determines the portion of the 3D shape represented by the first shape data to be corrected based on the 3D shape represented by the second shape data based on the comparison result between the 3D shape represented by the first shape data and the 3D shape represented by the second shape data.

7. The image processing device described in claim 5, wherein the shape correction unit determines a portion of the 3D shape represented by the first shape data to be corrected based on the 3D shape represented by the second shape data, based on the variation in the distribution of 3D points or polygon meshes representing the 3D shape of the subject in each of the first shape data and the second shape data.

8. The image processing device described in claim 5, wherein the shape correction unit presents to the user a comparison result between the 3D shape indicated by the first shape data and the 3D shape indicated by the second shape data, and corrects a portion of the 3D shape indicated by the first shape data specified by the user based on the 3D shape indicated by the second shape data.

9. The image processing device described in claim 5, wherein the second shape restoration unit restores the shape of the subject by acquiring the plurality of pieces of distance information based on each of the plurality of viewpoints of the visible light sensor based on the plurality of visible light images acquired by the visible light sensor photographing the subject from multiple viewpoints, and integrating the plurality of pieces of distance information based on each of the plurality of viewpoints of the visible light sensor based on the viewpoint information corresponding to each of the plurality of viewpoints of the visible light sensor.

10. An image processing device according to claim 2, further comprising a texture correction unit that corrects the texture mapped onto the surface of the 3D shape represented by the corrected first shape data based on the viewpoint information.

11. The image processing device according to claim 10, wherein the texture correction unit maps the visible light image as the texture onto the surface of the 3D shape indicated by the corrected first shape data based on the viewpoint information, and corrects the texture based on a comparison result between a rendering image of the first shape data onto which the texture has been mapped, rendered from the viewpoint of the visible light sensor indicated by the viewpoint information, and the visible light image captured from the viewpoint of the visible light sensor.

12. The image processing device according to claim 11, wherein the texture correction unit corrects at least one of the viewpoint information and the first shape data based on a result of comparing the rendering image with the visible light image.

13. The image processing device according to claim 11, wherein the texture correction unit corrects the texture based on the visible light image compared with the rendering image, or applies a high-frequency emphasis filter to the texture.

14. The image processing device according to claim 2, wherein the shape correction unit estimates the viewpoint of the visible light sensor based on the visible light image and the distance information.

15. The image processing device of claim 1, further comprising a second shape restoration unit that generates second shape data indicating the 3D shape of the subject by restoring the shape of the subject based on the visible light image and the viewpoint information, and wherein the shape correction unit corrects the 3D shape indicated by the first shape data based on the 3D shape indicated by the second shape data.

16. The image processing device described in claim 1, wherein the shape correction unit corrects the 3D shape represented by the first shape data based on a cross-polarized image acquired by the visible light sensor photographing the subject, and a parallel polarization image corresponding to the cross-polarized image is mapped as the texture onto the surface of the 3D shape represented by the first shape data.

17. An image processing method comprising: generating first shape data indicating the 3D shape of the subject by restoring the shape of the subject based on distance information based on the viewpoint of a depth sensor; and correcting the 3D shape indicated by the first shape data based on a visible light image that is mapped as a texture onto the surface of the 3D shape indicated by the first shape data, wherein the visible light image is an image obtained by photographing the subject with a visible light sensor at a viewpoint corresponding to the viewpoint of the depth sensor and has a higher resolution than the resolution of the distance information.

Citation Information

Patent Citations

  • Semantic segmentation method based on visible light image and low-resolution depth image

    CN113920317A

  • Information processing device and control method

    JP2023131258A

  • Information processing device and method, and information processing system

    WO2024080120A1