Information processing apparatus, information processing method, and program

By acquiring and combining multiple 3D shape models from different viewpoints, the method addresses the loss of detail in existing 3D reconstruction techniques, resulting in a more precise 3D shape model through likelihood-based alignment and adjustment.

JP2026011336AActive Publication Date: 2026-01-23CANON KK
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024111842
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-11
Publication Date
2026-01-23
Estimated Expiration
2044-07-11

Smart Images

  • Figure 2026011336000001_ABST
    Figure 2026011336000001_ABST
Patent Text Reader

Abstract

To generate a more detailed three dimensional shape model on the basis of a plurality of different three dimensional shape models.SOLUTION: A plurality of three dimensional shape models of the same object are acquired, and a likelihood (first likelihood) in each polygon constituting the acquired three dimensional shape models is calculated. A plurality of three dimensional shape models having different postures are converted into the same posture by using the estimated skeleton data, and scale conversion is performed on the posture-converted three dimensional shape model including the posture-converted skeleton data. Then, the likelihood (second likelihood) of the polygon for the arbitrary region is calculated from the calculated likelihood (first likelihood) of the polygon. Partial regions having the highest polygon likelihood (second likelihood) are extracted from the plurality of three dimensional shape models and combined to generate one three dimensional shape model.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing technique for processing three-dimensional shape information. [Background technology]

[0002] In recent years, many 3D reconstruction techniques using deep learning have been proposed as a technique for generating 3D shape models from 2D images. With this method, regions that cannot be observed from 2D images can be reconstructed using the estimation results of a pre-trained model.

[0003] For example, a trained model that estimates a 3D shape model using a 2D image of a person captured from one side and a corresponding correct 3D shape model as training data can estimate a statistically plausible 3D shape model from a 2D image showing one side of a person. Non-Patent Document 1 discloses a method for estimating a 3D shape model from a single input image based on statistical information obtained through pre-training. For example, a single frontal image of a person can be used to make a plausible estimation, including the back side, which is not included in the input information. However, compared to observed surfaces that can be estimated using observation information, non-observed surfaces that are estimated only from statistical information obtained through pre-training have a lower level of shape detail.

[0004] Non-Patent Document 2 discloses a method for obtaining a 3D shape model in a standard posture from a video sequence by optimization. In Non-Patent Document 2, the posture of the object to be reconstructed in the observation space is normalized, the radiance field of the person in the normalized space is optimized across the entire sequence, and a 3D shape model of one body is obtained from information on the entire sequence. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] Shunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima, Angjoo Kanazawa, and Hao Li. PIFu: Pixel-aligned implicit function for high-resolution clothed human digitization. In Proc. of the IEEE International Conf. on Computer Vision (ICCV), 2019. [Non-patent document 2] Chung-Yi Weng, Brian Curless, Pratul P. Srinivasan, Jonathan T. Barron, and Ira Kemelmacher-Shlizerman. “HumanNeRF: Free-viewpoint rendering of moving people from monocular video” In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022 Summary of the Invention [Problem to be solved by the invention]

[0006] However, in the technique of Non-Patent Document 2, a single 3D shape model is estimated from observation information averaged over the entire input sequence, which may result in the loss of detailed observation information in individual frames observed instantaneously.

[0007] The technology of the present disclosure aims to generate a more detailed three-dimensional shape model based on a plurality of different three-dimensional shape models. [Means for solving the problem]

[0008] The technology disclosed herein is characterized by comprising an acquisition means for acquiring multiple different 3D shape models for one object, a calculation means for calculating likelihood for each subregion of the multiple 3D shape models, and a generation means for combining the subregions of the multiple 3D shape models based on the likelihood to generate one 3D shape model for the one object. [Effects of the Invention]

[0009] According to the present disclosure, it is possible to generate a more detailed three-dimensional shape model based on a plurality of different three-dimensional shape models. [Brief explanation of the drawings]

[0010] [Figure 1] 1 is a diagram illustrating an example of a hardware configuration of an information processing device according to a first embodiment. [Figure 2] 1A and 1B are diagrams illustrating example input and output images of a 3D reconstruction DNN. [Figure 3] 4 is a flowchart illustrating a process for generating a three-dimensional shape model according to the first embodiment. [Figure 4] FIG. 10 is a diagram illustrating an example of a skeleton representation. [Figure 5] 1A to 1C are diagrams illustrating examples of input and output of a 3D reconstruction DNN according to the first embodiment. [Figure 6] FIG. 2 is a diagram showing the positional relationship between an object placed in a virtual three-dimensional space and a virtual camera. [Figure 7] FIG. 2 is a diagram showing the positional relationship between an object and an imaging camera. [Figure 8] FIG. 10 is a diagram for explaining a method for estimating a surface position of a three-dimensional shape model. [Figure 9] 5A to 5C are diagrams illustrating a process for generating a three-dimensional shape model in the same pose from a two-dimensional image according to the first embodiment. [Figure 10] 5A to 5C are diagrams illustrating a scale adjustment process according to the first embodiment. [Figure 11]1 is a conceptual diagram illustrating a method for generating a three-dimensional shape model based on a plurality of three-dimensional shape models according to the first embodiment. [Figure 12] 1A to 1C are diagrams illustrating examples of input and output of a posture conversion DNN according to the first embodiment. [Figure 13] FIG. 4 is a diagram for explaining a method for combining scale-adjusted three-dimensional shape models according to the first embodiment. [Figure 14] 10 is a flowchart illustrating a process for generating a three-dimensional shape model according to the second embodiment. [Figure 15] FIG. 10 is a diagram showing an example of a UI screen according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments according to the present disclosure will be described with reference to the drawings. The following embodiments do not limit the technology of the present disclosure, and not all of the combinations of features described in the present embodiments are necessarily essential to the solutions of the present disclosure. The configurations of the embodiments may be modified or changed as appropriate depending on the specifications of the device to which the technology of the present disclosure is applied and various conditions (such as usage conditions and usage environment). In the following embodiments, the same or similar components are designated by the same reference symbols, and redundant explanations will be omitted.

[0012] <Embodiment 1> 1(a) to 1(c) show an example of the hardware configuration of an information processing device in this embodiment. The information processing device 10 shown in Fig. 1(a) is intended for use as a PC, smartphone, tablet terminal, or the like, and is composed of an imaging unit 101, a CPU 102, a RAM 103, a ROM 104, a storage device 105, an operation unit 106, and a display unit 107.

[0013] The imaging unit 101 includes an imaging element and an image generation processing unit, and outputs the captured image to the storage unit 105 .

[0014] The CPU 102 executes various processes using computer programs and data stored in the RAM 103 and the ROM 104. In this way, the CPU 102 executes or controls various processes that will be described as controlling the operation of the information processing device 10 as a whole.

[0015] The RAM 103 has an area for storing computer programs and data loaded from the ROM 104 or storage device 105, and an area for storing data received from the multiple capture groups. The RAM 103 also has a work area used by the CPU 102 when executing various processes. In this way, the RAM 103 can provide various areas as needed.

[0016] The ROM 104 stores setting data for the information processing device 10, computer programs and data relating to startup, computer programs and data relating to basic operations, and the like.

[0017] The storage device 105 is a hard disk drive device or the like. The large-capacity storage device 104 stores an OS (operating system) and computer programs and data for causing the CPU 102 to execute or control various processes described as being performed by the information processing device 10. The data stored in the storage device 105 includes captured images generated by the imaging unit 101 and data related to a DNN model for performing 3D reconstruction. The computer programs and data stored in the large-capacity storage device 105 are loaded into the RAM 103 as appropriate under the control of the CPU 102 and become targets for processing by the CPU 102.

[0018] The operation unit 106 is a user interface such as a keyboard, a mouse, and a touch panel, and allows the user to input various instructions to the CPU 102 by operating it.

[0019] Display unit 107 has a screen such as a liquid crystal screen or a touch panel screen, and can display the processing results by CPU 102 as images, text, etc. Note that display unit 107 may also be a projection device such as a projector that projects images and text.

[0020] The imaging unit 101, CPU 102, RAM 103, ROM 104, mass storage device 105, operation unit 106, and display unit 107 are all connected to a system bus 108.

[0021] 1(b), the imaging unit 101 may be independent of the information processing device. In that case, imaging information including the captured image is transmitted to the information processing device 11 by a transmission unit 109. Alternatively, the information processing device may have only the information processing device 11, or may have a data acquisition unit 10A and a database 10B storing imaging information independent of the information processing device 12 as shown in FIG. 1(c). The configuration of the information processing device is not limited to the configurations shown in FIGS. 1(a) to 1(c), and the operation unit 106 and the display unit 107 may be included in an information processing device other than the information processing device 10.

[0022] The above-described information processing devices 10, 11, and 12 perform the 3D shape model acquisition, first likelihood calculation, second likelihood calculation, and 3D shape model generation of the present disclosure. The processing of each component of the present disclosure will be described below.

[0023] Using a DNN model, it is possible to estimate a 3D shape model from a single input image, but the estimated 3D shape model will have different regions with high estimation accuracy due to differences in the observation information contained in the input image. In this embodiment, partial regions with high estimation accuracy are extracted from multiple 3D shape models of the same object estimated from multiple input images with different observation information, and then these are combined to generate a highly accurate 3D shape model.

[0024] Figure 2 shows example input and output images of a 3D reconstruction DNN. A 3D shape model 1303 estimated from an input 2D image 1301, which contains information about the bumps and grooves of a jacket, shirt, and tie captured from a frontal view of an object, reproduces the bumps and grooves of the jacket, shirt, and tie. In contrast, a 3D shape model 1304 estimated from an input 2D image 1302, which captured an object from a rearward view, shows that the area around the chest, where the bumps and grooves of the jacket and tie should be, is estimated as a flat area. Because a 3D shape model estimated using a DNN model is based only on observational information contained in the input image and parameter information acquired through prior training, regions with less observational information tend to have lower reproducibility and estimation accuracy than regions with more observational information.

[0025] 3(a) to 3(e) show flowcharts for explaining the generation process of a 3D shape model according to this embodiment. Although the technology of the present disclosure does not limit the type of object for which a 3D shape model can be generated, this embodiment will be described taking a 3D shape model of a person as an example.

[0026] FIG. 3(a) shows a flowchart for explaining an outline of the process for generating a three-dimensional shape model according to this embodiment.

[0027] In S310, CPU 102 acquires multiple 3D shape models of the same object. Here, the same object may be an object in the same category, such as a "person," and is not limited to an entirely identical object (for example, the same person in the case of a "person").

[0028] In S320, the CPU 102 calculates the likelihood (first likelihood) of each polygon constituting the three-dimensional shape model acquired using the 3D reconstruction DNN in S310. The first likelihood here represents the estimated reliability of each polygon constituting the three-dimensional shape model estimated by the 3D reconstruction DNN.

[0029] In S330, the CPU 102 calculates the likelihood (second likelihood) of a polygon for an arbitrary region from the likelihood (first likelihood) of the polygon calculated in S320.

[0030] In S340, the CPU 102 generates a three-dimensional shape model by combining the partial regions with the highest polygon likelihood (second likelihood) calculated in S330.

[0031] <Process for acquiring multiple 3D shape models of the same object (S310)> The process of acquiring multiple 3D shape models of the same object in S310 will be described below. Fig. 3(b) shows a flowchart for explaining the process of acquiring 3D shape models of the same object in S310.

[0032] In S311, the CPU 102 acquires two-dimensional moving images or a plurality of still images of a person for whom a three-dimensional shape model is to be generated.

[0033] Next, the following steps S312 and S313 are repeated for the number of acquired images.

[0034] In S312, the CPU 102 inputs multiple frames of the two-dimensional moving image acquired in S311 or multiple still images to the 3D reconstruction DNN.

[0035] In S313, the CPU 102 acquires the data of the three-dimensional shape model and skeleton (skeleton model) output from the 3D reconstruction DNN and stores them in the storage unit 105.

[0036] FIG. 5(a) shows an example of input and output of the 3D reconstruction DNN used in this embodiment. In this embodiment, as shown in FIG. 5(a), a 3D shape model 505, 506, and a skeleton 507, 508 are estimated by a 3D reconstruction DNN 503, 504 from multiple input images 501, 502 of the same person. FIG. 5(b) is used for explanation. An input image 511 is input to a 3D reconstruction DNN 512, which outputs a 3D shape model 513 and a skeleton 514. Although the 3D reconstruction DNN architecture in FIGS. 5(a) and 5(b) shows an Hourglass Network, the DNN architecture is not limited to this. Alternatively, separate DNNs, such as a 3D shape model estimation DNN 515 and a skeleton estimation DNN 516, may be used, as shown in FIG. 5(c).

[0037] Here, the training method of the 3D reconstruction DNN will be explained. The 3D reconstruction DNN in this embodiment estimates a 3D voxel map of a 3D shape model and skeleton data indicating joint positions, but first, the training method for estimating the 3D voxel map of a 3D shape model will be explained. The training data used is a pair of a 3DCG object that serves as Ground Truth (hereinafter referred to as Ground Truth object) and a 2D virtual viewpoint image obtained by rendering the Ground Truth object with a virtual camera.

[0038] FIG. 6 shows a Ground Truth object 601 and a virtual camera placed in a virtual three-dimensional space. The input images are two-dimensional virtual viewpoint images obtained by rendering the Ground Truth object 601 from the viewpoint of each virtual camera. The loss used to optimize the 3D reconstruction DNN is defined as the difference between the Ground Truth object 601 and the three-dimensional shape model estimated from the 3D reconstruction DNN 512. The data format of the three-dimensional shape model used in learning is a 3D Occupancy Field, where the inside of the target object area is 1 and the outside of the target object area is 0. Therefore, the Ground Truth object 601 holds a two-valued correct answer value, and the voxel value f of the Ground Truth object at the voxel position coordinate X is v * (X) is expressed as equation (1).

[0039]

number

[0040]

number

[0041] The 3D shape model can be obtained by converting the 3D voxel map with continuous values ​​output by the 3D reconstruction DNN into binary values ​​of 0 and 1 using threshold processing, or by converting it into a mesh format using the Marching Cube method, etc.

[0042] Next, we will explain the training method for skeleton estimation. The training data used for skeleton estimation training is a pair of input images, one of which is a Ground Truth 3D skeleton and the other is a rendered CG object with this Ground Truth 3D skeleton. There are multiple skeleton data formats, and the main difference is the number of joints, but this proposal does not impose any restrictions on the skeleton data format.

[0043] An example of a skeleton representation is shown in FIG. 4(a). The position coordinates of each joint in the skeleton representation shown in FIG. 4(a) are expressed as J h (h=1,...,H). As shown in Figure 4(b), h (h=1,...,H). In the skeleton representation shown in Figure 4(a), H=21.

[0044] The 3D reconstruction DNN 512 outputs H 3D confidence maps. The coordinate system of these 3D confidence maps is the same as the coordinate system of the 3D voxel map used when estimating the 3D shape model. Each 3D confidence map maps the probability of the existence of a corresponding joint to voxel position coordinates, and the voxel position coordinate with the highest probability of the joint being present becomes the joint position coordinate. Position coordinate X on the 3D confidence map W The probability that a joint exists in P(X W ), the estimated joint position coordinate J h can be expressed as equation (3).

[0045]

number

[0046]

number

[0047]

number

[0048]

number

[0049]

number

[0050] The three-dimensional shape model and skeleton data acquired in S310 may be acquired from a database 10B in which a plurality of three-dimensional shape models of the object to be reconstructed that have been estimated in the past and the corresponding skeleton data are stored in advance.

[0051] <First Likelihood Calculation Process (S320)> The calculation process of the first likelihood in S320 will be described below.

[0052] 3(c) shows a flowchart for explaining the first likelihood calculation process of S320. Here, the object 701 in FIG. 7 is expressed as x WConsider estimating a three-dimensional shape model from an image captured by an imaging camera 702 on the axis. In this case, the x coordinate in the camera coordinate system of the imaging camera 702 is c′ ,y c′ direction (y in global coordinate system) W ,z W It is easy to estimate with high accuracy the position of the surface of a 3D shape model having a normal consisting only of a z component (direction) from the information of the captured image. c′ direction (x in global coordinates) W It is difficult to estimate with high accuracy the position of the surface of a three-dimensional shape model that also has a (directional) component. Conversely, when estimating a three-dimensional shape model from an image captured by the imaging camera 703, the direction in which the surface position can be estimated with high accuracy is the x direction in the global coordinate system. W z W The direction that is difficult to estimate with high accuracy is the y W In this way, the area that can be easily estimated with high accuracy by the 3D reconstruction DNN changes depending on the positional relationship between the object and the imaging camera.

[0053] The 3D reconstruction DNN outputs the probability that each voxel exists inside the object region as a continuous value between 0 and 1. If the 3D reconstruction DNN can estimate with high accuracy that the voxel being estimated exists inside the object region, it outputs a value close to 1, and if it can estimate with high accuracy that it exists outside the object region, it outputs a value close to 0. On the other hand, if the 3D reconstruction DNN cannot estimate with high accuracy whether the voxel being estimated exists inside or outside the object region, it outputs a value close to the intermediate value of 0.5.

[0054] Figure 8(b) shows a cross-sectional view of an example of applying the Marching Cube method to an estimated 3D voxel map that assumes the surface position of a 3D shape model is estimated with high accuracy (high likelihood). Here, the threshold value in the Marching Cube method is set to 0.5, and the surface exists in the region with a voxel value of 0.5. Figure 8(c) shows a cross-sectional view of an example of applying the Marching Cube method to an estimated 3D voxel map that assumes the surface position is estimated with high accuracy (low likelihood). The greater the difference in estimated voxel values ​​between adjacent voxels across the surface, the higher the likelihood, and the smaller the difference, the lower the likelihood.

[0055] In S321, the CPU 102 calculates the normal vector of each polygon. The normal vector of a polygon formed by three points A, B, and C as shown in Fig. 8(a) can be calculated from the cross product of vector AB and vector AC.

[0056] In S322, the CPU 102 extracts adjacent voxels using the normal direction calculated in S321.

[0057] In S323, the CPU 102 calculates the likelihood of a polygon based on the difference in voxel values ​​between a voxel that contacts the polygon and a voxel that is adjacent to that voxel. As shown in FIG. 8(b), when it is easy to estimate the surface position with high accuracy, the difference in the estimated values ​​between the two voxels becomes large, and when it is difficult to estimate the surface position with high accuracy, as shown in FIG. 8(c), the difference in the estimated values ​​between the two voxels becomes small. The likelihood of a polygon is calculated from the difference in the two voxel estimated values. The voxels adjacent to the inside and outside directions of the polygon M are respectively designated as V M ,V M′ Let each estimate be

[0058]

number

[0059]

number

[0060] In S331, the CPU 102 converts a plurality of 3D shape models with different poses into the same pose using the estimated skeleton data. Here, the explanation is given assuming that the conversion is to a canonical T-pose. The coordinates of points on the 3D shape model before the pose conversion are expressed as s o (∈S o ), and the point coordinates after the posture transformation are s T (∈S T ), then Canonical T-pose S T Pose S o The process of converting this into can be expressed as in equation (7).

[0061]

number

[0062]

number

[0063]

number

[0064]

number

[0065]

number

[0066]

number

[0067]

number

[0068] In S332, the CPU 102 performs scale conversion on the pose-converted 3D shape models including the multiple skeleton data obtained in S331. As shown in FIG. 9B, the pose-converted 3D shape model 915 has a narrower face and torso width than the pose-converted 3D shape model 914, while the pose-converted 3D shape model 916 has a wider torso width and longer hands than the pose-converted 3D shape model 914. Thus, even if the same object representing the same person is reconstructed and converted to the same pose, the scales of the pose-converted 3D shape models, such as height and thickness, may not match, resulting in inconsistent 3D shape models in the global coordinate system. This is because, as described with reference to FIG. 7, it is difficult for the 3D reconstruction DNN to accurately estimate the depth size in the camera coordinate system of the imaging camera, resulting in variations in the accuracy of the 3D shape models. Therefore, in S332, the scale of the 3D shape models is adjusted using skeleton data corresponding to each 3D shape model, making it easier to achieve correspondence between each region of the multiple pose-converted 3D shape models.

[0069] Therefore, in S332, first, the most reliable skeleton vector among the skeleton data after posture transformation is determined as the reference scale that will be the scale after adjustment, as follows: Object IDi is set for each 3D shape model, and the skeleton vector after posture transformation is determined as

[0070]

number

[0071]

number

[0072]

number

[0073]

number

[0074] In S332, the most reliable skeleton vector determined in this way is

[0075]

number

[0076]

number

[0077]

number

[0078]

number

[0079]

number

[0080]

number

[0081]

number

[0082]

number

[0083] In S334, the CPU 102 calculates the likelihood (second likelihood) for each partial region. The likelihood for each partial region is calculated using the first likelihood of the polygon calculated in S323. If the number of division grids is G, then the likelihood C for each grid g is g is the number of polygons in grid g, N g As a result, we get equation (12).

[0084]

number

[0085]

number

[0086] <Processing for generating a 3D shape model by combining multiple 3D shape models (S340)> The following describes the process of generating a 3D shape model by combining the most likely subregions in S340. A flowchart is shown in Figure 3(e). For ease of explanation, consider combining the most likely subregions 1005, 1006, 1007, and 1008 of the scale-adjusted 3D shape models 1001, 1002, 1003, and 1004 shown in Figure 13(a).

[0087] In S341, the CPU 102 extracts the most polygon-likely model from the plurality of three-dimensional shape models acquired in S310 for each grid.

[0088] In S342, the CPU 102 converts the grid polygons extracted in S341 into voxels and combines them to generate a 3D voxel map of a single 3D shape model. When grid-based 3D shape models extracted from different 3D shape models are combined, polygons are misaligned at the grid boundaries, as shown in FIG. 13(b). When a 3D shape model is generated by combining polygons with such misalignment at the grid boundaries, holes are created in the 3D shape model. Therefore, in the voxels near each polygon, estimated voxel values ​​of voxels adjacent to the polygon and voxel values ​​of adjacent voxels in the direction of the polygon normal used to calculate the polygon likelihood are stored. Of the voxels near the polygon, voxels within the object region are stored with a voxel value of 1, and voxels outside the object region are stored with a voxel value of 0.

[0089] In S343, the CPU 102 converts the 3D voxel map generated in S342 into a mesh format using the Marching Cube method or the like.

[0090] As described above, in this embodiment, a highly accurate 3D shape model can be generated by combining the highly accurate portions of multiple 3D shape models that are independently estimated from multiple frames of a moving image or multiple still images obtained by capturing the same object.

[0091] <Embodiment 2> This embodiment is a variation of the first embodiment. The system configuration diagram in this embodiment is the same as that in the first embodiment. In this embodiment, a user can generate any combination of 3D shape models by applying any 3D shape model to any partial region from multiple 3D shape models based on the user's preference through GUI operations. In the second embodiment, the user operates the GUI via the operation unit 106 and the display device 107. The hardware for operating and displaying the envisioned GUI may be a tablet terminal, a smartphone, a PC, or the like, as long as it is connected to the information processing device 10 and functions as the operation unit 106 and the display device 107.

[0092] 14(a) shows a flowchart for explaining the generation process of a 3D shape model in this embodiment. S1410, S1420, S1430, and S1440 are the same as S310, S320, S331, and S332 in the flowchart of embodiment 1, and therefore their explanation will be omitted here.

[0093] In S1450, the CPU 102 generates a GUI displaying a reference representative 3D shape model and displays it on the display unit 107. The representative 3D shape model may be one generated by dividing a 3D shape model estimated independently from different pieces of observation information into regions based on grids of a predetermined size, as described in the first embodiment, and combining partial regions in each grid where the likelihood of a polygon is highest. The representative 3D shape model is not limited to this, and may be a single appropriate 3D shape model, or may be one generated from multiple frames of a video as disclosed in Non-Patent Document 1.

[0094] FIG. 15 shows an example of a GUI screen displayed on a tablet terminal according to this embodiment. FIG. 15(a) shows a representative 3D shape model 1501 displayed on the screen of the tablet terminal. This tablet terminal is capable of accepting user operations on the displayed representative 3D shape model on a screen that serves as both the operation unit 106 and the display unit 107. For example, as shown in FIG. 15(b), when the representative 3D shape model displayed on the tablet terminal is touched and slid in the direction the user wants to see, the representative 3D shape model rotates. This allows a representative 3D shape model that represents the view from the direction the user wants to see to be displayed.

[0095] In S1460, the CPU 102 generates a highly accurate three-dimensional shape model based on user input of a GUI operation received by the operation unit 106. Fig. 14(b) shows a flowchart for explaining the three-dimensional shape model generation processing in S1460.

[0096] In S1461, the CPU 102 accepts a selection of an area of ​​the three-dimensional shape model based on a user input of a GUI operation accepted by the operation unit 106. As shown in Fig. 15(c), the user can tap on the area they want to modify to enlarge the display. Then, as shown in Fig. 15(d), when the user drags and selects the area they want to modify on the enlarged three-dimensional shape model 1502, a grid 1503 is generated, and the polygons within the generated grid become the selected area.

[0097] In S1462, the CPU 102 calculates and scores the likelihood of the partial region, which means the selected region, for the three-dimensional shape model acquired in S1410 according to equation (12). Note that the score may be the likelihood itself or a processed likelihood.

[0098] In S1463, the CPU 102 generates a GUI that displays the calculated likelihood scoring results and displays the GUI on the display unit 107. Here, as shown in FIG. 15(e), the top three 3D shape models are displayed as candidates. If the user does not like these candidate 3D shape models, the user can press the "Change" button to switch to another three 3D shape models with the next highest scores.

[0099] In S1464, the CPU 102 accepts the selection of one three-dimensional shape model from the candidate three-dimensional shape models by the user based on the user input of the GUI operation accepted by the operation unit 106.

[0100] In S1465, the CPU 102 replaces the selected area of ​​the representative three-dimensional shape model with the selected area of ​​the three-dimensional shape model selected in S1464, generates a GUI displaying a three-dimensional shape model reflecting the user's selection, and displays it on the display unit 107. The method of combining the three-dimensional shape models at this time is the same as S341 and S342 in the first embodiment. As shown in FIG. 15(f), when the user selects one candidate three-dimensional shape model 1504, a three-dimensional shape model 1505 is generated in which the selected area is the three-dimensional shape model 1504. When "End" is pressed, the three-dimensional shape model generation process ends.

[0101] As described above, the GUI allows users to check, select, and combine highly accurate partial regions of a 3D shape model, making it easier to generate the 3D shape model they desire. It also makes it easier for users to set the granularity of the partial regions to be corrected as desired.

[0102] (Other Examples) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0103] The present disclosure includes the following configurations and methods. [Configuration 1] an acquisition means for acquiring a plurality of different 3D shape models for one object; a calculation means for calculating likelihood for each partial region of the plurality of three-dimensional shape models; a generating means for generating a single 3D shape model for the single object by combining partial regions of the plurality of 3D shape models based on the likelihood; An information processing device comprising: [Configuration 2] The acquisition means is a plurality of 3D shape models generated by a trained model that generates a 3D shape model from a captured image obtained by capturing an image of the one object. 2. The information processing device according to configuration 1, [Configuration 3] the calculation means calculates likelihoods for polygons of the three-dimensional shape model, and calculates likelihoods for each of the partial regions based on the likelihoods for the polygons. 3. The information processing device according to configuration 1 or 2. [Configuration 4] the acquisition means acquires the plurality of three-dimensional shape models as voxel maps; the calculation means calculates the likelihood for the polygon based on voxel values ​​in the voxel map of voxels that contact the polygon; 4. The information processing device according to configuration 3. [Configuration 5] the calculation means calculates the likelihood of the polygon from a difference between voxel values ​​adjacent to the polygon in a normal direction; 5. The information processing device according to configuration 4. [Configuration 6] the generating means generates the single 3D shape model by combining voxel maps corresponding to the partial regions. 6. The information processing device according to configuration 4 or 5. [Configuration 7] the acquisition means acquires a skeleton corresponding to the three-dimensional shape model, the generating means performs posture transformation on the plurality of 3D shape models based on the skeleton so that the plurality of 3D shape models have the same posture before combining the partial regions. 7. The information processing device according to any one of configurations 1 to 6. [Configuration 8] the generating means performs scale conversion on the plurality of 3D shape models based on the skeleton so that the models have the same size before combining the partial regions. 8. The information processing device according to configuration 7. [Configuration 9] a display means for displaying the three-dimensional shape model; Accepting means for accepting user input; Furthermore, the generating means uses, for the subregion selected based on the user input, the selected subregion in one of the plurality of three-dimensional shape models selected based on the user input; 9. The information processing device according to any one of configurations 1 to 8. [Configuration 10] the receiving means receives a selection of a partial region of the three-dimensional shape model; The generating means changes the partial area selected by the user to the selected partial area of ​​another three-dimensional shape model. 10. The information processing device according to configuration 9. [Configuration 11] The generating means Referring to the likelihood for the selected subregion, selecting a 3D shape model of the selected partial region of the plurality of 3D shape models based on the likelihood for the partial region and displaying it on the display means; the accepting means accepts a user's selection of a 3D shape model of the partial region of the plurality of displayed 3D shape models; 11. The information processing device according to configuration 10. [Configuration 12] obtaining a plurality of different three-dimensional shape models for one object; calculating a likelihood for each partial region of the plurality of three-dimensional shape models; generating a single 3D shape model for the single object by combining subregions of the plurality of 3D shape models based on the likelihood; 1. A method for controlling an information processing device, comprising: [Configuration 13] 12. A program for causing a computer to function as the information processing device according to any one of configurations 1 to 11.

Claims

1. an acquisition means for acquiring a plurality of different three-dimensional shape models for one object; a calculation means for calculating likelihood for each partial region of the plurality of three-dimensional shape models; a generating means for generating a single three-dimensional shape model for the single object by combining partial regions of the plurality of three-dimensional shape models based on the likelihood; An information processing device comprising:

2. The acquisition means is a plurality of 3D shape models generated by a trained model that generates a 3D shape model from a captured image obtained by capturing an image of the one object.

2. The information processing apparatus according to claim 1, wherein:

3. the calculation means calculates likelihoods for polygons of the three-dimensional shape model, and calculates likelihoods for each of the partial regions based on the likelihoods for the polygons.

2. The information processing apparatus according to claim 1, wherein:

4. the acquisition means acquires the plurality of three-dimensional shape models as voxel maps; the calculation means calculates the likelihood for the polygon based on voxel values ​​in the voxel map of voxels that contact the polygon; 4. The information processing apparatus according to claim 3,

5. the calculation means calculates the likelihood of the polygon from a difference between voxel values ​​adjacent to the polygon in a normal direction; 5. The information processing apparatus according to claim 4,

6. the generating means generates the single three-dimensional shape model by combining voxel maps corresponding to the partial regions.

5. The information processing apparatus according to claim 4,

7. the acquiring means acquires a skeleton corresponding to the three-dimensional shape model; the generating means performs posture transformation on the plurality of three-dimensional shape models based on the skeleton so that the plurality of three-dimensional shape models have the same posture before combining the partial regions.

2. The information processing apparatus according to claim 1, wherein:

8. the generating means performs scale conversion on the plurality of three-dimensional shape models based on the skeleton so that the models have the same size before combining the partial regions.

8. The information processing apparatus according to claim 7,

9. a display means for displaying the three-dimensional shape model; Accepting means for accepting user input; Furthermore, the generating means uses, for the subregion selected based on the user input, the selected subregion in one of the plurality of three-dimensional shape models selected based on the user input; 2. The information processing apparatus according to claim 1, wherein:

10. the receiving means receives a selection of a partial region of the three-dimensional shape model; The generating means changes the partial area selected by the user to the selected partial area of ​​another three-dimensional shape model. The information processing device according to claim 9 .

11. The generating means Referring to the likelihood for the selected subregion, selecting a three-dimensional shape model of the selected partial region of the plurality of three-dimensional shape models based on the likelihood for the partial region, and displaying the selected three-dimensional shape model on the display means; the accepting means accepts a user's selection of a 3D shape model of the partial region of the plurality of displayed 3D shape models; 11. The information processing apparatus according to claim 10,

12. obtaining a plurality of different three-dimensional shape models for one object; calculating a likelihood for each partial region of the plurality of three-dimensional shape models; generating a single 3D shape model for the single object by combining partial regions of the plurality of 3D shape models based on the likelihood; 1. A method for controlling an information processing device, comprising:

13. A program for causing a computer to function as the information processing device according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Modeling method based on oblique photography of unmanned aerial vehicle

    CN114359503A

  • Voxel merging method and device, equipment and storage medium

    CN116310149A

  • Image processing apparatus, method, and program

    JP2021140594A

  • Merging three-dimensional models based on confidence scores

    WO2013166023A1

  • Mixed three dimensional scene reconstruction from plural surface models

    WO2017003673A1