Image processing device and image processing method
The image processing apparatus optimizes 3D model alignment and texture mapping using a reference model to address shape and material challenges, resulting in high-quality 3D models with improved texture accuracy.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SONY GROUP CORP
- Filing Date
- 2025-10-28
- Publication Date
- 2026-05-21
AI Technical Summary
Existing 3D scanning technologies face challenges in generating high-quality 3D models due to variations in subject shape during scanning, material limitations, and alignment issues between 3D scanner and RGB camera data, leading to poor texture quality.
An image processing apparatus and method that optimizes the alignment of 3D models using a third model as a reference, correcting shape and texture data based on material information and reliability, and interpolating occluded areas to generate high-quality 3D models.
Enables the creation of high-quality 3D models with accurate textures by aligning and optimizing 3D scanner and RGB camera data, even with subject shape changes and material limitations, improving texture mapping and occlusion interpolation.
Smart Images

Figure JP2025037729_21052026_PF_FP_ABST
Abstract
Description
Image Processing Apparatus and Image Processing Method
[0001] The present technology relates to an image processing apparatus and an image processing method, and particularly relates to an image processing apparatus and an image processing method that enable a high-quality 3D model of a subject to be obtained using a 3D scanner and an RGB camera.
[0002] A technique has been proposed for generating a 3D model of a subject by generating shape data indicating the 3D shape of the subject using a 3D scanner and attaching the RGB image acquired by the RGB camera capturing the subject as a texture to the surface of the 3D shape indicated by the shape data.
[0003] For example, Patent Document 1 describes a technique in which segmentation for identifying segments is performed on each of the shape data and the RGB image, alignment of the shape data and the RGB image is performed based on the segments common between the shape data and the RGB image, and then the RGB image is attached to the surface of the 3D shape indicated by the shape data.
[0004] Japanese Unexamined Patent Application Publication No. 2021-189600
[0005] In a 3D scanner, the 3D shape of the entire subject is measured by irradiating the subject with light from various viewpoints. If the shape of the subject changes even slightly while the viewpoint of the 3D scanner is being moved, the quality of the shape data generated using the 3D scanner may deteriorate.
[0006] In a 3D scanner, due to the principle of measuring the 3D shape of a subject by irradiating it with light, depending on the material of the surface of the subject, the 3D shape may not be accurately measured, and the quality of the shape data may deteriorate.
[0007] If there is even a slight difference in the shape of the subject between the measurement by the 3D scanner and the shooting by the RGB camera, it is difficult to accurately align the shape data and the RGB image, and the quality of the texture of the 3D model may deteriorate.
[0008] The present technology has been made in view of such a situation, and enables a high-quality 3D model of a subject to be obtained using a 3D scanner and an RGB camera.
[0009] One aspect of this technology is an image processing apparatus which includes: an alignment unit that aligns a first 3D model which includes first shape data which shows the 3D shape of an object, and a second 3D model which includes second shape data which shows the 3D shape of the object and second texture data which shows the color information of the object; an optimization unit that optimizes the alignment of the first 3D model and the second 3D model based on a third 3D model which is generated by a method different from the method for generating the first 3D model and includes third shape data which shows the 3D shape of the object; and a texture mapping unit which generates first texture data of the first 3D model by attaching the color information shown by the second texture data to the surface of the 3D shape shown by the first shape data based on the alignment result of the first 3D model and the second 3D model.
[0010] One aspect of this technology is an image processing method which includes: aligning a first 3D model which includes first shape data which shows the 3D shape of an object, and a second 3D model which includes second shape data which shows the 3D shape of the object and second texture data which shows the color information of the object; optimizing the alignment of the first 3D model and the second 3D model based on a third 3D model which is generated by a method different from the method for generating the first 3D model and includes third shape data which shows the 3D shape of the object; and generating first texture data for the first 3D model by applying the color information shown by the second texture data to the surface of the 3D shape shown by the first shape data based on the alignment result of the first 3D model and the second 3D model.
[0011] In one aspect of this technology, alignment is performed between a first 3D model including first shape data representing the 3D shape of an object, and a second 3D model including second shape data representing the 3D shape of the object and texture data representing the color information of the object. The alignment of the first 3D model and the second 3D model is generated by a method different from the method for generating the first 3D model, and is optimized based on a third 3D model including third shape data representing the 3D shape of the object. Based on the alignment result of the first 3D model and the second 3D model, the color information represented by the second texture data is applied to the surface of the 3D shape represented by the first shape data, thereby generating first texture data for the first 3D model.
[0012] This is a block diagram showing an example configuration of a 3D modeling system according to one embodiment of this technology. This is a diagram showing an example of viewpoints for a 3D scanner and an RGB camera. This is a diagram illustrating an example of rotating an object. This is a block diagram showing an example configuration of an image processing device. This is a diagram illustrating an example of a completed 3D model. This is a flowchart explaining the processes performed by the 3D modeling system. This is a flowchart explaining the processes performed by the 3D modeling system when the pose of the object during 3D scanning differs from the pose of the existing 3D model. This is a block diagram showing an example configuration of computer hardware.
[0013] The following describes the configurations for implementing this technology. The explanation will proceed in the following order: 1. Overview of the 3D modeling system 2. Configuration and operation of each device 3. Application examples
[0014] <1. Overview of the 3D Modeling System> Figure 1 is a block diagram showing an example configuration of a 3D modeling system according to one embodiment of this technology.
[0015] As shown in Figure 1, the 3D modeling system of this technology consists of a 3D scanner 1, an RGB camera 2, and an image processing device 3. The 3D scanner 1 and the RGB camera 2 are examples of a depth sensor and a visible light sensor, respectively, in this disclosure.
[0016] 3D scanner 1 is configured as a depth sensor that utilizes methods such as a stereo camera system, a structured light system, or a ToF (Time of Flight) system. 3D scanner 1 performs a 3D scan by illuminating the subject with light, thereby measuring the 3D shape of the subject as the distance to the subject. 3D scanner 1 measures the distance to the same subject captured by RGB camera 2.
[0017] Hereinafter, points on the surface of an object whose distance from the 3D scanner 1 is measured will be referred to as distance measurement points. The 3D scanner 1 measures the distance between the 3D scanner 1 and each distance measurement point from multiple viewpoints and acquires a 3D point cloud showing the three-dimensional distribution of each distance measurement point for each viewpoint. The 3D scanner 1 supplies the multiple 3D point clouds based on each viewpoint as scan data to the image processing device 3.
[0018] The RGB camera 2 captures the subject from multiple viewpoints, acquires multiple RGB images (visible light images), and supplies them to the image processing device 3.
[0019] Figure 2 shows an example of the viewpoints of the 3D scanner 1 and the RGB camera 2.
[0020] The 3D scanner 1 measures the distance to the subject Obj1, such as a plastic model, from multiple viewpoints 1-1 to 1-N, which are set up on a circle surrounding the subject Obj1, as shown in Figure 2.
[0021] The RGB camera 2 takes pictures from viewpoints 2-1 to 2-N, which photograph the subject Obj1 from below looking up; viewpoints 2'-1 to 2'-N, which photograph the subject Obj1 from a nearly horizontal direction; and viewpoints 2''-1 to 2''-N, which photograph the subject Obj1 from above looking down. In the example in Figure 2, the lower viewpoints 2-1 to 2-N, the middle viewpoints 2'-1 to 2'-N, and the upper viewpoints 2''-1 to 2''-N are set to positions on the circumference that surround the subject in each row.
[0022] Users of the 3D modeling system can, for example, move one 3D scanner 1 and one RGB camera 2 around to perform photography and 3D scanning from various viewpoints. Alternatively, multiple 3D scanners 1 and multiple RGB cameras 2 may be set up at various viewpoints, allowing photography and 3D scanning to be performed from each viewpoint.
[0023] Here, the viewpoint includes, for example, the position and orientation (tilt) of the 3D scanner 1 and RGB camera 2, based on the position and orientation of the subject.
[0024] Instead of changing the viewpoints of the 3D scanner 1 and RGB camera 2 by moving the 3D scanner 1 and RGB camera 2, the viewpoints of the 3D scanner 1 and RGB camera 2 may be changed by rotating the subject Obj1, as shown in Figure 3.
[0025] In the example shown in Figure 3, the subject Obj1 is placed on the turntable 21, the 3D scanner 1 is fixed at one position 1A, and the RGB cameras 2 are fixed at positions 2A to 2C on the lower, middle, and upper levels. When the user rotates the turntable 21, the subject Obj1 rotates horizontally, allowing the 3D scanner 1 and RGB cameras 2 to capture images and perform 3D scans from various viewpoints.
[0026] Returning to Figure 1, the image processing device 3 is composed of a PC, a cloud server, and the like.
[0027] The image processing device 3 performs shape reconstruction to restore the 3D shape of the subject based on multiple scan data (3D point clouds) supplied from the 3D scanner 1, each based on a different viewpoint, and generates a measured 3D model. The measured 3D model is a 3D model generated based on scan data, and includes measured shape data and measured texture data. The measured shape data is shape data that represents the 3D shape (geometry information) of the subject, for example, using polygon meshes or 3D point clouds, and the measured texture data is texture data that represents the color information of the subject, for example, using UV maps that map textures applied to each polygon mesh or each 3D point in two dimensions. The measured texture data included in the measured 3D model generated based on scan data is also called the original measured texture data.
[0028] Furthermore, the image processing device 3 performs 3D reconstruction of RGB images acquired from multiple viewpoints supplied by the RGB camera 2 to generate a real-world 3D model. The real-world 3D model is a 3D model generated based on the RGB image and includes real-world shape data and real-world texture data. The real-world shape data is shape data that represents the 3D shape (geometry information) of the subject, for example, using polygon meshes or 3D point clouds, and the real-world texture data is texture data that represents the color information of the subject, for example, using UV maps that map textures applied to each polygon mesh or each 3D point in two dimensions.
[0029] Generally, the color information displayed in original, measured texture data is of lower quality than the color information displayed in photorealistic texture data.
[0030] The image processing device 3 aligns the measured 3D model with the actual photographic 3D model based on an existing 3D model (existing 3D model) that shows the 3D shape, color, pattern, texture, etc. of the subject. The existing 3D model is CT data, which is data showing the 3D shape of the subject obtained by irradiating the subject with X-rays using a CT (Computed Tomography) device, or CAD (Computer-Aided Design) data created when designing the subject, such as a plastic model. CT data does not contain color information. CAD data does not contain color information, but it does contain material information that shows the material of the subject. In addition, CAD data contains part information that shows the part of the CAD data corresponding to each of the multiple parts that make up the subject.
[0031] The image processing device 3 generates measured texture data (updates the original measured texture data) by applying color information indicated by the actual texture data to the surface of the 3D shape indicated by the measured shape data (texture mapping) based on the alignment result of the measured 3D model and the actual 3D model.
[0032] Aligning measured 3D models with real-world 3D models is performed using methods such as ICP (Iterative Closest Point). ICP is a method that iteratively calculates the optimal rotation and translation parameters for each 3D model to minimize the error between the 3D models.
[0033] Generally, even when alignment is performed using methods such as ICP, the measured 3D model and the actual photographic 3D model will not perfectly overlap. For example, if there is even a slight difference in the shape of the subject between the 3D scan by 3D scanner 1 and the capture by RGB camera 2 (the shape changes due to the passage of time or the shaking of the turntable), the 3D models will not perfectly overlap even if alignment is performed to minimize the error.
[0034] Furthermore, if the quality of at least one of the measured shape data or the actual photographic shape data is low, and outliers or missing values are present in only one of the shape data sets, the alignment discrepancy will be large.
[0035] In particular, 3D scanner 1 and RGB camera 2 perform 3D scanning and photography from various viewpoints to acquire shape data that shows the 3D shape of the entire object. Therefore, if the shape of the object changes even slightly while the viewpoint of 3D scanner 1 or RGB camera 2 is moved, the quality of the shape data may decrease. Also, 3D scanner 1 and RGB camera 2 have materials that they are not good at acquiring shape data from. For example, 3D scanner 1 is not good at acquiring shape data from objects with low reflectivity surfaces or transparent surfaces, and RGB camera 2 is not good at acquiring shape data from metallic objects or objects with transparent surfaces.
[0036] Thus, if the shape of the subject changes while the viewpoint of 3D scanner 1 or RGB camera 2 is being moved, or if the subject is made of a material that is not well-suited for acquiring shape data for either 3D scanner 1 or RGB camera 2, the alignment error is likely to become large.
[0037] If the alignment is significantly off, texture mapping will not be performed properly, resulting in a decrease in the quality of the textures of the final measured 3D model.
[0038] This technology was developed in light of these circumstances, and by optimizing the alignment of measured 3D models and real-world 3D models based on existing 3D models, it enables the acquisition of high-quality 3D models of subjects using a 3D scanner 1 and an RGB camera 2.
[0039] <2. Configuration and Operation of Each Device> Figure 4 is a block diagram showing an example of the configuration of the image processing device 3.
[0040] As shown in Figure 4, the image processing device 3 is composed of a point cloud integration unit 41, a shape restoration unit 42, conversion units 43, 44, a subject pose estimation unit 45, a positioning unit 46, optimization units 47, 48, a texture mapping unit 49, a texture correction unit 50, an occlusion interpolation unit 51, and a texture mapping unit 52.
[0041] The point cloud integration unit 41 generates a measured 3D model that represents the 3D shape of the entire subject in the form of a 3D point cloud by integrating a plurality of 3D point clouds (scan data) based on each of the plurality of viewpoints acquired by the 3D scanner 1. The point cloud integration unit 41 functions as a first shape restoration unit that generates a measured 3D model by performing shape restoration based on a plurality of scan data. The point cloud integration unit 41 performs meshing to convert the measured 3D model into a polygon mesh, and supplies the meshed measured 3D model to the subject pose estimation unit 45 and the alignment unit 46.
[0042] The shape restoration unit 42 (second shape restoration unit) generates a photographed 3D model that represents the 3D shape of the entire subject in the form of a 3D point cloud by performing 3D reconstruction (photogrammetry) on a plurality of RGB images acquired by the RGB camera 2. The shape restoration unit 42 supplies the generated photographed 3D model to the conversion unit 43.
[0043] The conversion unit 43 performs meshing to convert the photographed 3D model generated by the shape restoration unit 42 into a polygon mesh, and supplies the meshed photographed 3D model to the alignment unit 46.
[0044] The conversion unit 44 performs meshing to convert the existing 3D model into a polygon mesh or a volume mesh, and supplies the meshed existing 3D model to the subject pose estimation unit 45.
[0045] The subject pose estimation unit 45 estimates the pose of the subject at the time of 3D scanning based on the measured shape data of the measured 3D model supplied from the point cloud integration unit 41. Here, the pose of the subject is, for example, the position and angle of each joint of the subject. The subject pose estimation unit 45 deformes the pose of the existing 3D model supplied from the conversion unit 44 based on the estimation result of the pose of the subject at the time of 3D scanning.
[0046] Specifically, the subject pose estimation unit 45 recursively repeats evaluating the difference between the measured shape data and the shape data of the existing 3D model while moving the movable parts of the shape data of the existing 3D model, so as to match the pose of the subject at the time of 3D scanning, and deformes the pose of the existing 3D model.
[0047] The subject pose estimation unit 45 supplies the existing 3D model with its pose deformed to the alignment unit 46.
[0048] The alignment unit 46 performs alignment among the measured 3D model supplied from the point cloud integration unit 41, the captured 3D model supplied from the conversion unit 43, and the existing 3D model supplied from the subject pose estimation unit 45 using a method such as ICP. The alignment unit 46 supplies the optimized unit 47 with the aligned existing 3D model and the measured 3D model. Also, the alignment unit 46 supplies the optimized unit 48 with the aligned existing 3D model and the captured 3D model.
[0049] The optimization units 47 and 48 optimize the alignment between the measured 3D model and the captured 3D model performed by the alignment unit 46 based on the existing 3D model supplied from the alignment unit 46. For example, the optimization units 47 and 48 optimize the alignment between the measured 3D model and the captured 3D model by correcting the measured shape data and the captured shape data based on the shape data of the existing 3D model.
[0050] Specifically, first, the optimization unit 47 calculates the reliability of the measured model for each coordinate in the virtual space including the aligned measured 3D model and the existing 3D model. Next, the optimization unit 47 corrects the measured shape data by integrating the measured shape data and the shape data of the existing 3D model based on the reliability for each coordinate. For example, at coordinates with high reliability, the geometry information of the measured shape data is adopted, and at coordinates with low reliability, the geometry information of the shape data of the existing 3D model is adopted.
[0051] The reliability of the measured 3D model is calculated based on the volume difference between the measured shape data and the shape data of the existing 3D model, the color information indicated by the original measured texture data, and the material information included in the existing 3D model. For example, the larger the volume difference between the measured shape data and the shape data of the existing 3D model, the lower the reliability for each coordinate will be set. Also, for example, the material of the subject corresponding to each coordinate is estimated based on the color information indicated by the original measured texture data, and the reliability is set higher for coordinates where the estimated material matches the material indicated by the material information. Furthermore, for example, if a part of the subject composed of a material that is unsuitable for 3D scanner 1 is identified based on the material information, the reliability for the coordinates corresponding to that part of the subject will be set lower.
[0052] Similarly, first, the optimization unit 48 calculates the reliability of the real-life 3D model for each coordinate in the virtual space, which includes the aligned real-life 3D model and the existing 3D model. Next, the optimization unit 48 corrects the measured shape data by integrating the real-life shape data and the shape data of the existing 3D model based on the reliability for each coordinate.
[0053] The optimization unit 47 corrects the measured shape data so that the measured 3D model and the existing 3D model overlap in the virtual space. The optimization unit 48 corrects the real-world shape data so that the real-world 3D model and the existing 3D model overlap in the virtual space. In other words, the optimization units 47 and 48 correct the difference between the 3D shape of the subject at the time of 3D scanning, which is reflected in the measured shape data, and the 3D shape of the subject at the time of shooting, which is reflected in the real-world shape data, using the 3D shape shown by the existing 3D model as a reference. By correcting the difference between the 3D shape of the subject at the time of 3D scanning and the 3D shape of the subject at the time of shooting, the optimization units 47 and 48 can optimize the alignment of the measured 3D model and the real-world 3D model.
[0054] Furthermore, the optimization unit 47 corrects outliers and missing parts in the measured shape data based on the shape data of the existing 3D model. The optimization unit 48 corrects outliers and missing parts in the actual shape data based on the shape data of the existing 3D model. The optimization unit 48 supplies the actual 3D model, with outliers and missing parts corrected, to the alignment unit 46.
[0055] The alignment unit 46 performs alignment again between the measured 3D model, the existing 3D model, and the real-life 3D model corrected by the optimization unit 48. Generally, real-life shape data tends to be of lower quality than measured shape data, which can easily lead to larger alignment discrepancies. The optimization unit 48 improves the quality of the real-life shape data by correcting outliers and missing parts contained in the real-life shape data, and then has the alignment unit 46 perform alignment again, thereby optimizing the alignment between the measured 3D model and the real-life 3D model.
[0056] The optimization unit 47 discards the original measured texture data (original texture data) and supplies only the corrected measured shape data to the texture mapping unit 49. The optimization unit 48 supplies the corrected real-world 3D model to the texture mapping units 49 and 52.
[0057] The texture mapping unit 49 generates (updates) new texture data for the live-action 3D model by pasting the color information indicated by the live-action texture data onto the surface of the 3D shape indicated by the measured shape data, based on the alignment results of the measured 3D model and the live-action 3D model. For example, the texture mapping unit 49 copies the color information to be pasted onto the live-action shape data for each coordinate in the virtual space including the measured shape data and the live-action 3D model, and pastes it onto the measured shape data. The texture mapping unit 49 combines the measured shape data and the updated texture data into an updated measured 3D model and supplies the updated measured 3D model to the texture correction unit 50.
[0058] Furthermore, instead of performing texture mapping on shape data generated by integrating measured shape data and shape data from existing 3D models, texture mapping may be performed on shape data that integrates measured shape data, photographic shape data, and shape data from existing 3D models based on their reliability.
[0059] The texture correction unit 50 corrects the measured texture data based on a comparison between the measured 3D model supplied by the texture mapping unit 49 and a rendered image obtained from the viewpoint of the RGB camera 2 when a certain RGB image was acquired.
[0060] If the quality of the measured texture data is low, a difference of a certain magnitude will occur between the RGB image and the rendered image. The texture correction unit 50 compares the RGB image and the rendered image, and if there is a difference of a certain magnitude or more, it corrects the color information in the measured texture data corresponding to the part with the difference. For example, the texture correction unit 50 replaces the color information in the measured texture data corresponding to the part with the difference with color information extracted from the RGB image, or applies a high-frequency enhancement filter to the color information in the texture data corresponding to the part with the difference.
[0061] In this way, the texture correction unit 50 recursively repeats the process of correcting the UV map while changing the reference RGB image, so as to minimize the error between the rendered image and the RGB image. The texture correction unit 50 then supplies the corrected measured 3D model to the occlusion interpolation unit 51.
[0062] The occlusion interpolation unit 51 interpolates the occlusion portions of the surface of the subject that could not be obtained during 3D scanning, in the measured shape data of the measured 3D model supplied from the texture correction unit 50.
[0063] Actual measured shape data does not reflect the 3D shape of occlusion areas such as joints and the backs of parts that are hidden by parts of the object itself when viewed from the viewpoint used during 3D scanning.
[0064] Therefore, the occlusion interpolation unit 51 acquires the difference between the measured shape data and the shape data of the existing 3D model supplied from the alignment unit 46 as occlusion data, which is data that interpolates the occluded portion of the measured shape data. The occlusion interpolation unit 51 integrates the measured shape data and the occlusion data to reflect the 3D shape of the occluded portion in the measured shape data.
[0065] The occlusion interpolation unit 51 supplies the texture mapping unit 52 with an existing 3D model and a measured 3D model that has measured shape data integrated with occlusion data.
[0066] The texture mapping unit 52 obtains part information about the occlusion area and its surrounding area from an existing 3D model, and based on this part information, generates color information to be pasted onto the occlusion area in the measured shape data of the measured 3D model supplied by the occlusion interpolation unit 51. Specifically, the texture mapping unit 52 identifies the part of the measured 3D model that constitutes the same part as the occlusion area based on the part information, and copies the color information to be pasted onto the part that constitutes the same part as the occlusion area and pastes it onto the occlusion area.
[0067] The texture mapping unit 52 applies color information to the occlusion areas, completing a 3D model that shows the actual 3D shape, color, pattern, and texture of the subject.
[0068] Figure 5 illustrates an example of a completed 3D model.
[0069] For example, if a wooden box with a closed lid, as shown on the left side of Figure 5A, is used as the subject for 3D scanning and photography, a 3D model will ultimately be generated that reproduces not only the 3D shape of the outside of the box, but also the 3D shape of the inside of the box. Therefore, it is possible to move the lid of the 3D model of the wooden box to expose the inside of the box, as shown by the dotted ellipse on the right side of Figure 5A.
[0070] For example, if a 3D scan and photograph are performed on the arm portion of a plastic model with an unbent elbow, as shown on the left side of Figure 5B, the final 3D model generated will also reproduce the 3D shape of the joint parts stored inside the upper arm and forearm parts. Therefore, it is possible to move the forearm part of the 3D model of the arm to expose the joint parts, as shown by the dotted ellipse on the right side of Figure 5B.
[0071] For example, if a 3D scan and photograph are performed on the leg portion of a plastic model with the left knee bent at a 90-degree angle, as shown on the left side of Figure 5C, the final 3D model generated will also reproduce the 3D shape of the joint parts housed inside the thigh and lower leg parts. Therefore, it is possible to move the lower leg portion of the 3D model of the leg portion to expose the joint parts, as shown by the dotted ellipse on the right side of Figure 5C.
[0072] Next, referring to the flowchart in Figure 6, we will explain the processes performed by the 3D modeling system having the above configuration.
[0073] In step S1, the 3D scanner 1 performs a 3D scan from multiple viewpoints to measure the distance to the subject and acquires multiple scan data.
[0074] In step S2, the point cloud integration unit 41 of the image processing device 3 generates a measured 3D model by integrating multiple scan data.
[0075] In step S3, the point cloud integration unit 41 meshes the measured 3D model.
[0076] In step S4, the RGB camera 2 takes pictures of the subject from multiple viewpoints and acquires multiple images.
[0077] In step S5, the shape restoration unit 42 of the image processing device 3 generates a real-life 3D model by performing shape restoration based on multiple captured images.
[0078] In step S6, the conversion unit 43 of the image processing device 3 converts the real-life 3D model into a mesh.
[0079] In step S7, the conversion unit 44 of the image processing device 3 converts the existing 3D model into a mesh.
[0080] In step S8, the alignment unit 46 of the image processing device 3 performs alignment between the measured 3D model, the actual 3D model, and the existing 3D model.
[0081] In step S9, the optimization unit 47 of the image processing device 3 corrects the measured shape data of the measured 3D model based on the shape data of an existing 3D model.
[0082] In step S10, the optimization unit 47 determines whether the quality of the measured shape data is sufficient.
[0083] If it is determined in step S10 that the quality of the measured shape data is insufficient, the process returns to step S3, and the meshing and alignment of the measured 3D model and the existing 3D model are performed again. In this case, for example, steps S4 to S6 are skipped.
[0084] On the other hand, if it is determined in step S10 that the quality of the measured shape data is sufficient, in step S11 the optimization unit 48 corrects the actual shape data of the actual 3D model based on the shape data of the existing 3D model.
[0085] In step S12, the optimization unit 48 of the image processing device 3 determines whether the quality of the actual shape data is sufficient.
[0086] If it is determined in step S12 that the quality of the live-action shape data is insufficient, the process returns to step S6, and the meshing and alignment of the live-action 3D model and the existing 3D model are performed again.
[0087] On the other hand, if it is determined in step S12 that the quality of the actual shape data is sufficient, in step S13, the texture mapping unit 49 of the image processing device 3 performs texture mapping on the measured shape data to generate measured texture data. The image processing device 3 combines the measured shape data and the measured texture data into a new measured 3D model.
[0088] In step S14, the texture correction unit 50 of the image processing device 3 corrects the measured texture data of the measured 3D model based on the RGB image.
[0089] In step S15, the occlusion interpolation unit 51 of the image processing device 3 interpolates the occlusion portions in the measured shape data of the measured 3D model based on the shape data of the existing 3D model. The texture mapping unit 52 of the image processing device 3 performs texture mapping on the occlusion portions of the measured shape data.
[0090] Next, referring to the flowchart in Figure 7, we will explain the process that the 3D modeling system performs when the pose of the subject during 3D scanning differs from the pose of the existing 3D model.
[0091] The processes in steps S31 to S37 are the same as those in steps S1 to S7 in Figure 6, so their explanation will be omitted.
[0092] In step S38, the subject pose estimation unit 45 of the image processing device 3 estimates the pose of the subject during 3D scanning based on the measured shape data of the measured 3D model, and deforms the pose of the existing 3D model based on the estimation result of the subject's pose during 3D scanning.
[0093] In step S39, the subject pose estimation unit 45 determines whether the pose of the subject during 3D scanning matches the pose of the existing 3D model, and repeatedly deforms the pose of the existing 3D model until the pose of the subject during 3D scanning matches the pose of the existing 3D model.
[0094] After the pose of the existing 3D model is deformed to match the pose of the subject during 3D scanning, the process proceeds to step S40. The processes from steps S40 to S48 are basically the same as the processes from steps S8 to S16 in Figure 6, so their explanation is omitted.
[0095] As described above, in the image processing apparatus 3 of this technology, alignment is performed between a first 3D model (e.g., measured 3D model) which includes first shape data (e.g., measured shape data) indicating the 3D shape of an object, and a second 3D model (e.g., live-action 3D model) which includes second shape data (e.g., live-action shape data) indicating the 3D shape of an object and second texture data (e.g., live-action texture data) indicating the color information of an object. The alignment of the first 3D model and the second 3D model is generated by a method different from the generation method of the first 3D model, and is optimized based on a third 3D model (e.g., existing 3D model) which includes third shape data indicating the 3D shape of an object. Based on the alignment result of the first 3D model and the second 3D model, the color information indicated by the second texture data is applied to the surface of the 3D shape indicated by the first shape data, thereby generating first texture data (e.g., updated live-action texture data).
[0096] The image processing device 3 optimizes the alignment between the measured 3D model and the real-world 3D model, enabling it to appropriately apply texture mapping to the measured shape data of the measured 3D model and generate a 3D model with high-quality textures applied.
[0097] Furthermore, in the image processing device 3 of this technology, outliers and missing values included in the shape data of the second 3D model are corrected based on the shape data of the third 3D model. As a result, the image processing device 3 can apply textures to high-quality measured shape data from which outliers and missing values have been corrected, and consequently, it becomes possible to generate a high-quality 3D model.
[0098] Generally, the quality of original measured texture data is lower than that of photorealistic texture data. Therefore, the color information shown in the photorealistic texture data is applied to the surface of the 3D shape shown by the measured shape data. However, it is also possible to apply the color information shown in the original measured texture data to the surface of the 3D shape shown by the photorealistic shape data.
[0099] The above describes an example of 3D modeling based on measured 3D models and real-world 3D models. However, it is also possible to perform 3D modeling based on either the measured 3D model or the real-world 3D model and CT data. In this case, a 3D model that does not become the basis for the final 3D model may be used as an existing 3D model. Since CT scanners are not good at acquiring shape data of metal objects, when performing 3D modeling of metal objects, for example, using a measured 3D model as an existing 3D model and performing 3D modeling based on CT data and real-world 3D models may result in a higher quality 3D model than using CT data as an existing 3D model.
[0100] <3. Application Examples> This technology's 3D modeling system can be applied to 3D modeling of industrial products such as vehicles and plastic models for which CAD data can be obtained. Even if CAD data cannot be obtained, if CT data can be obtained for the subject, this technology's 3D modeling system can perform 3D modeling of that subject.
[0101] This technology's 3D modeling system can be applied to 3D modeling of buildings and entire cities using drones equipped with 3D scanners and RGB cameras. In this case, existing 3D models would be design data of the building or city.
[0102] The series of processes described above can be executed by hardware or by software. When the series of processes are executed by software, the programs that make up the software are installed from a program storage medium onto a computer built into dedicated hardware, or a general-purpose personal computer.
[0103] Figure 8 is a block diagram showing an example of the hardware configuration of a computer that executes the series of processes described above using a program.
[0104] The CPU (Central Processing Unit) 501, ROM (Read Only Memory) 502, and RAM (Random Access Memory) 503 are interconnected by a bus 504.
[0105] An input / output interface 505 is further connected to the bus 504. An input unit 506 consisting of a keyboard, mouse, etc., and an output unit 507 consisting of a display, speakers, etc. are connected to the input / output interface 505. In addition, a storage unit 508 consisting of a hard disk, non-volatile memory, etc., a communication unit 509 consisting of a network interface, etc., and a drive 510 that drives removable media 511 are connected to the input / output interface 505.
[0106] In a computer configured as described above, the CPU 501 loads, for example, a program stored in the memory unit 508 into the RAM 503 via the input / output interface 505 and the bus 504, and executes it, thereby performing the series of processes described above.
[0107] The program executed by the CPU 501 is recorded on removable media 511, for example, or provided via a wired or wireless transmission medium such as a local area network, the internet, or digital broadcasting, and installed in the storage unit 508.
[0108] The programs executed by the computer may be programs that are processed chronologically in the order described herein, or they may be programs that are processed in parallel or at necessary times, such as when a call is made.
[0109] In this specification, a system means a collection of multiple components (devices, modules (parts), etc.), regardless of whether all components are located in the same enclosure. Therefore, multiple devices housed in separate enclosures and connected via a network, and a single device containing multiple modules in one enclosure, are both considered systems.
[0110] The effects described herein are illustrative and not limited to those described herein, and other effects may also occur.
[0111] The embodiments of this technology are not limited to those described above, and various modifications are possible without departing from the spirit of this technology.
[0112] For example, this technology can be configured as cloud computing, where a single function is shared and processed collaboratively by multiple devices via a network.
[0113] Furthermore, each step described in the flowchart above can be performed by a single device, or it can be divided and performed by multiple devices.
[0114] Furthermore, if a single step includes multiple processes, those processes can be executed by a single device or shared among multiple devices.
[0115] • Example configurations: This technology can also take the following configurations.
[0116] (1) An image processing apparatus comprising: an alignment unit that aligns a first 3D model including first shape data indicating the 3D shape of a subject, and a second 3D model including second shape data indicating the 3D shape of the subject and second texture data indicating color information of the subject; an optimization unit that optimizes the alignment of the first 3D model and the second 3D model based on a third 3D model which is generated by a method different from the method for generating the first 3D model and includes third shape data indicating the 3D shape of the subject; and a texture mapping unit that generates first texture data of the first 3D model by applying the color information indicated by the second texture data to the surface of the 3D shape indicated by the first shape data based on the alignment result of the first 3D model and the second 3D model. (2) The image processing apparatus according to (1), wherein the optimization unit optimizes the alignment of the first 3D model and the second 3D model by correcting the first shape data and the second shape data based on the third shape data. (3) The image processing apparatus according to (2), wherein the optimization unit corrects outliers and missing values in the second shape data based on the third shape data, and then causes the alignment unit to perform alignment between the first 3D model and the second 3D model, thereby optimizing the alignment of the first 3D model and the second 3D model. (4) The image processing apparatus according to (2) or (3), further comprising a first shape restoration unit that generates the first 3D model by restoring the shape of the subject based on a plurality of scan data obtained by a depth sensor measuring the distance to the subject from a plurality of viewpoints. (5) The image processing apparatus according to (4), further comprising a second shape restoration unit that generates the second 3D model by performing 3D reconstruction of a plurality of visible light images obtained by a visible light sensor photographing the subject from a plurality of viewpoints. (6) The image processing apparatus according to (5), wherein the optimization unit optimizes the alignment of the first 3D model and the second 3D model by correcting the difference between the shape of the subject when measuring distance with the depth sensor and the shape of the subject when capturing with the visible light sensor, based on the third shape data.(7) The image processing apparatus according to (5) or (6), further comprising a interpolation unit that interpolates occlusion portions of the surface of the subject that could not be obtained when measuring distance by the depth sensor in the first shape data, based on the third shape data. (8) The image processing apparatus according to (7), wherein the third 3D model includes part information indicating portions of the third shape data corresponding to each of a plurality of parts constituting the subject, and the texture mapping unit generates the color information to be pasted onto the occlusion portions based on the first texture data and the part information. (9) The image processing apparatus according to any one of (5) to (8), further comprising a subject posture estimation unit that estimates the posture of the subject when measuring distance by the depth sensor based on the first 3D model, and deforms the posture of the third 3D model based on the estimation result of the posture of the subject. (10) The image processing apparatus according to any one of (5) to (9), further comprising a texture correction unit that corrects the first texture data generated by the texture mapping unit based on the visible light image. (11) The image processing apparatus according to any one of (2) to (10), wherein the alignment unit performs alignment between the first 3D model, the second 3D model, and the third 3D model. (12) The image processing apparatus according to any one of (2) to (11), wherein the optimization unit corrects the first shape data by integrating the first shape data and the third shape data based on the reliability of the first 3D model calculated for each coordinate in the virtual space including the aligned first 3D model and the third 3D model, and corrects the second shape data by integrating the second shape data and the third shape data based on the reliability of the second 3D model calculated for each coordinate in the virtual space including the aligned second 3D model and the third 3D model.(13) The image processing apparatus according to (12), wherein the first 3D model includes first shape data and original texture data of lower quality than the second texture data, and the texture mapping unit generates the first texture data by updating the original texture data based on the second texture data. (14) The image processing apparatus according to (13), wherein the third 3D model includes material information indicating the material of the subject and third shape data, and the optimization unit calculates the reliability of the first 3D model based on any of the volume difference between the first shape data and the third shape data, the original texture data, and the material information, and calculates the reliability of the second 3D model based on any of the volume difference between the second shape data and the third shape data, the second texture data, and the material information. (15) An image processing method comprising: aligning a first 3D model including first shape data indicating the 3D shape of a subject, and a second 3D model including second shape data indicating the 3D shape of the subject and second texture data indicating color information of the subject; optimizing the alignment of the first 3D model and the second 3D model based on a third 3D model which is generated by a method different from the method for generating the first 3D model and includes third shape data indicating the 3D shape of the subject; and generating first texture data of the first 3D model by applying the color information indicated by the second texture data to the surface of the 3D shape indicated by the first shape data based on the alignment result of the first 3D model and the second 3D model.
[0117] 1 3D scanner, 2 RGB camera, 3 Image processing device, 41 Point cloud integration unit, 42 Shape reconstruction unit, 43, 44 Conversion unit, 45 Subject pose estimation unit, 46 Alignment unit, 47, 48 Optimization unit, 49 Texture mapping unit, 50 Texture correction unit, 51 Occlusion interpolation unit, 52 Texture mapping unit
Claims
1. An image processing apparatus comprising: a first 3D model including first shape data indicating the 3D shape of a subject; an alignment unit that aligns a first 3D model including second shape data indicating the 3D shape of the subject and second texture data indicating color information of the subject; an optimization unit that optimizes the alignment of the first 3D model and the second 3D model based on a third 3D model which is generated by a method different from the method for generating the first 3D model and includes third shape data indicating the 3D shape of the subject; and a texture mapping unit that generates first texture data of the first 3D model by applying the color information indicated by the second texture data to the surface of the 3D shape indicated by the first shape data based on the alignment result of the first 3D model and the second 3D model.
2. The image processing apparatus according to claim 1, wherein the optimization unit optimizes the alignment of the first 3D model and the second 3D model by correcting the first shape data and the second shape data based on the third shape data.
3. The image processing apparatus according to claim 2, wherein the optimization unit corrects outliers and missing values included in the second shape data based on the third shape data, and then causes the alignment unit to perform alignment between the first 3D model and the second 3D model, thereby optimizing the alignment of the first 3D model and the second 3D model.
4. The image processing apparatus according to claim 2, further comprising a first shape restoration unit that generates the first 3D model by restoring the shape of the subject based on a plurality of scan data obtained by the depth sensor measuring the distance to the subject from multiple viewpoints.
5. The image processing apparatus according to claim 4, further comprising a second shape reconstruction unit that generates the second 3D model by performing 3D reconstruction of a plurality of visible light images acquired by a visible light sensor capturing the subject from a plurality of viewpoints.
6. The image processing apparatus according to claim 5, wherein the optimization unit optimizes the alignment of the first 3D model and the second 3D model by correcting the difference between the shape of the subject when measuring distance with the depth sensor and the shape of the subject when capturing with the visible light sensor, based on the third shape data.
7. The image processing apparatus according to claim 5, further comprising a interpolation unit that interpolates the occlusion portion of the surface of the subject that could not be obtained when measuring distance by the depth sensor in the first shape data, based on the third shape data.
8. The image processing apparatus according to claim 7, wherein the third 3D model includes part information indicating a portion of the third shape data corresponding to each of the plurality of parts constituting the subject, and the texture mapping unit generates the color information to be applied to the occlusion portion based on the first texture data and the part information.
9. The image processing apparatus according to claim 5, further comprising a subject posture estimation unit that estimates the posture of the subject when measuring distance by the depth sensor based on the first 3D model, and deforms the posture of the third 3D model based on the estimation result of the subject's posture.
10. The image processing apparatus according to claim 5, further comprising a texture correction unit that corrects the first texture data generated by the texture mapping unit based on the visible light image.
11. The image processing apparatus according to claim 2, wherein the alignment unit performs alignment between the first 3D model, the second 3D model, and the third 3D model.
12. The image processing apparatus according to claim 2, wherein the optimization unit corrects the first shape data by integrating the first shape data and the third shape data based on the reliability of the first 3D model calculated for each coordinate in the virtual space including the aligned first 3D model and the third 3D model, and corrects the second shape data by integrating the second shape data and the third shape data based on the reliability of the second 3D model calculated for each coordinate in the virtual space including the aligned second 3D model and the third 3D model.
13. The image processing apparatus according to claim 12, wherein the first 3D model includes first shape data and original texture data of lower quality than the second texture data, and the texture mapping unit generates the first texture data by updating the original texture data based on the second texture data.
14. The image processing apparatus according to claim 13, wherein the third 3D model includes material information indicating the material of the subject and the third shape data, the optimization unit calculates the reliability of the first 3D model based on any of the volume difference between the first shape data and the third shape data, the original texture data, and the material information, and calculates the reliability of the second 3D model based on any of the volume difference between the second shape data and the third shape data, the second texture data, and the material information.
15. An image processing method comprising: aligning a first 3D model including first shape data indicating the 3D shape of a subject, and a second 3D model including second shape data indicating the 3D shape of the subject and second texture data indicating color information of the subject; optimizing the alignment of the first 3D model and the second 3D model based on a third 3D model which is generated by a method different from the method for generating the first 3D model and includes third shape data indicating the 3D shape of the subject; and generating first texture data of the first 3D model by applying the color information indicated by the second texture data to the surface of the 3D shape indicated by the first shape data based on the alignment result of the first 3D model and the second 3D model.