Information processing system, information processing method, and information processing program

A single-camera system reconstructs 3D shape and texture data to generate a 3D model of a person, addressing the complexity of existing methods by integrating pre-created data and real-time updates.

JP7759654B2Active Publication Date: 2025-10-24NAT INST OF INFORMATION & COMM TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022001280
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-01-06
Publication Date
2025-10-24
Estimated Expiration
2042-01-06

AI Technical Summary

Technical Problem

Existing methods for reconstructing a 3D model of a person require depth sensors or multiple cameras, complicating the device configuration.

Method used

A system that uses a single camera to capture 2D images, reconstructs 3D shape and texture data, and integrates them to generate a 3D model, utilizing pre-created 3D shape and texture data, and includes units for facial and body reconstruction, pose estimation, and texture blending.

Benefits of technology

Enables the reproduction of a 3D model with a simpler configuration, allowing for real-time reflection of movements and facial expressions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007759654000001
    Figure 0007759654000001
  • Figure 0007759654000002
    Figure 0007759654000002
  • Figure 0007759654000003
    Figure 0007759654000003
Patent Text Reader

Abstract

To provide an arrangement for reproducing a 3D model of a person with a simpler arrangement.SOLUTION: An information processing system according to the present invention has a face texture reconstruction unit for reconstructing a face texture from a 2D image of a person captured by a camera, a face shape reconstruction unit for reconstructing a 3D shape of a face from the 2D image of the person, a posing estimation unit for estimating a pose of the person from the 2D image of the person, a shape integration unit for reconstructing the 3D shape of a body corresponding to the pose estimated based on the 3D shape data and for reconstructing the 3D shape of the person by integrating the reconstructed 3D shape of the body and the reconstructed 3D shape of the face, a texture reconstruction unit for reconstructing a texture image of the person by blending the texture image of the reconstructed face into the texture image included in the texture data, and a model generation unit for generating the 3D model of the person based on the 3D shape of the person captured by the camera and the texture image of the person captured by the camera.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing system, an information processing method, and an information processing program for reproducing a 3D model. [Background technology]

[0002] A technology has been proposed that reconstructs a 3D model that more realistically represents a person, transmits the reconstructed 3D model to a remote location, and uses XR (VR / AR / MR) technology to share a 3D space, providing remote communication.

[0003] For example, Non-Patent Document 1 discloses a system that uses a depth sensor to acquire the 3D shape and texture of a person, transmits the acquired 3D shape and texture to a remote location, and uses an MR (mixed reality) headset to communicate with the person in a state where the 3D model is superimposed on real space. The 3D shape of a person may also be acquired using multiple cameras. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] M. Joachimczak, J. Liu, H. Ando. 2017. Real Time Mixed Reality Telepresence via 3D Reconstruction with HoloLens and Commodity Depth Sensors. In Proceedings of 19th ACM International Conference on Multimodal Interaction (ICMI' 17). ACM, New York, NY, USA, 2 pages. https: / / doi.org / 10. 1145 / 3136755.3143031 Summary of the Invention [Problem to be solved by the invention]

[0005] The above-mentioned prior art requires a depth sensor or multiple cameras, which can complicate the device configuration. Therefore, there is a need for a method that can acquire the 3D shape and texture of a person from a 2D image of the person captured by a single camera, and transmit and reproduce the 3D model of the person at a remote location.

[0006] One object of the present invention is to provide a configuration that can reproduce a 3D model of a person with a simpler configuration. [Means for solving the problem]

[0007] an information processing system according to an embodiment includes a camera; a memory unit that stores pre-created 3D shape data indicating the 3D shape of a body and texture data indicating the texture of the body; a facial texture reconstruction unit that reconstructs the texture of a face from a 2D image of the person captured by the camera; a facial shape reconstruction unit that reconstructs the 3D shape of the face from the 2D image of the person captured by the camera; a pose estimation unit that estimates the pose of the person from the 2D image of the person captured by the camera; a shape integration unit that reconstructs the 3D shape of the body corresponding to the estimated pose based on the 3D shape data and integrates the reconstructed 3D shape of the body and the reconstructed 3D shape of the face to reconstruct the 3D shape of the person captured by the camera; a texture reconstruction unit that reconstructs the texture image of the face by blending the texture image included in the texture data with the reconstructed texture image of the face; and a model generation unit that generates a 3D model of the person captured by the camera based on the 3D shape of the person captured by the camera and the texture image of the person captured by the camera.

[0008] The texture data may include a texture image corresponding to the 3D shape of the reconstructed body and a texture image corresponding to the 3D shape of the reconstructed face, a texture map corresponding to the 3D shape of the reconstructed body and a texture map corresponding to the 3D shape of the reconstructed face.

[0009] The information processing system may further include a body shape reconstruction unit that reconstructs a 3D shape of the body from multiple 2D images of the person captured by the camera, a head shape reconstruction unit that reconstructs a 3D shape of the head from multiple 2D images of the person captured by the camera, and a texture integration unit that determines a correspondence between the reconstructed 3D shape of the body and the reconstructed 3D shape of the head, determines a correspondence between a texture map corresponding to the 3D shape of the body and a texture map corresponding to the 3D shape of the head based on the determined correspondence of the 3D shapes, and generates a texture image corresponding to the 3D shape of the head from a texture image corresponding to the 3D shape of the body based on the determined correspondence of the texture maps.

[0010] The shape integration unit may integrate the reconstructed 3D shape of the body and the reconstructed 3D shape of the face based on a texture map included in the texture data.

[0011] The model generation unit may integrate the 3D shape of the person with the texture image of the person based on a texture map included in the texture data.

[0012] The texture reconstruction unit may superimpose the result of passing the mask through a texture image of the person captured by the camera onto the texture image included in the texture data.

[0013] The mask may be configured to have a continuously varying transmittance.

[0014] A partial image corresponding to a window set in a 2D image of a person captured by a camera may be input to the facial texture reconstruction unit and the facial shape reconstruction unit. The information processing system may further include a stabilization unit that sets the window by temporally smoothing the position of the person in the 2D image.

[0015] The texture integration unit may generate texture data by integrating a texture image corresponding to the 3D shape of the body with a texture image corresponding to the 3D shape of the head, and by integrating a texture map corresponding to the 3D shape of the body with a texture map corresponding to the 3D shape of the head.

[0016] An information processing method according to another embodiment includes the steps of: reconstructing facial texture from a 2D image of a person captured by a camera; reconstructing a 3D shape of the face from the 2D image of the person captured by the camera; estimating a pose of the person from the 2D image of the person captured by the camera; reconstructing a 3D shape of the body corresponding to the estimated pose based on pre-created 3D shape data indicating the 3D shape of the body; integrating the reconstructed 3D shape of the body and the reconstructed 3D shape of the face to reconstruct a 3D shape of the person captured by the camera; reconstructing a texture image of the person captured by the camera by blending a texture image included in pre-created texture data indicating the texture of the body with the texture image of the reconstructed face; and generating a 3D model of the person captured by the camera based on the 3D shape of the person captured by the camera and the texture image of the person captured by the camera.

[0017] According to yet another embodiment, there is provided an information processing program for causing a computer to execute the above method. [Effects of the Invention]

[0018] According to the present invention, a 3D model of a person can be reproduced with a simpler configuration. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is a schematic diagram showing an example of a system configuration of an information processing system according to an embodiment of the present invention. [Figure 2] FIG. 1 is a schematic diagram showing an example of a hardware configuration of an information processing device that constitutes an information processing system according to the present embodiment. [Figure 3] 10 is a flowchart showing a processing procedure in an initial model construction stage of the information processing system according to the present embodiment. [Figure 4] 10 is a flowchart showing a processing procedure at a 3D model reproduction stage in the information processing system according to the present embodiment. [Figure 5] FIG. 2 is a schematic diagram showing an example of a functional configuration for realizing an initial model construction stage of the information processing system according to the present embodiment. [Figure 6] FIG. 4 is a diagram showing an example of data generated in an initial model construction stage of the information processing system according to the present embodiment. [Figure 7] FIG. 10 is a schematic diagram for illustrating texture integration processing in the initial model construction stage of the information processing system according to the present embodiment. [Figure 8] FIG. 2 is a schematic diagram showing an example of a functional configuration for realizing a 3D model reproduction stage of the information processing system according to the present embodiment. [Figure 9] FIG. 2 is a diagram showing an example of data generated in a 3D model reproduction stage of the information processing system according to the present embodiment. [Figure 10] FIG. 10 is a schematic diagram for illustrating a blending process in the information processing system according to the present embodiment. [Figure 11] FIG. 10 is a schematic diagram showing another example of the system configuration of the information processing system according to the present embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0020] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention will be described in detail with reference to the accompanying drawings, in which the same or corresponding parts are designated by the same reference numerals and will not be described repeatedly.

[0021] In this specification, three-dimensional or solid is abbreviated as "3D", and two-dimensional or planar is abbreviated as "2D".

[0022] [A. System Configuration] Fig. 1 is a schematic diagram showing an example of a system configuration of an information processing system 1 according to the present embodiment. Fig. 1 shows an example of a configuration in which, for example, information processing devices 100-1 and 100-2 (hereinafter sometimes collectively referred to as "information processing devices 100") and an information processing device 200 are connected via a network 2. A camera 140-1 is connected to the information processing device 100-1, and a camera 140-2 is connected to the information processing device 100-2.

[0023] The information processing device 100 acquires an initial model of the person 10 in advance. The information processing device 100 recreates a 3D model of the person 10 by continuously capturing images of the person 10 with a camera 140. The recreated 3D model changes in real time to reflect the movements and facial expressions of the captured person 10. The recreated 3D model of the person 10 is also called a 3D avatar, or simply an avatar.

[0024] 1, a person 10-1 is present within the field of view of camera 140-1, and a person 10-2 is present within the field of view of camera 140-2. The information processing device 100-1 captures an image of the person 10-1 to reproduce a 3D model 20-1 of the person 10-1 on the screen of the information processing device 200, for example. Similarly, the information processing device 100-2 captures an image of the person 10-2 to reproduce a 3D model 20-2 of the person 10-2 on the screen of the information processing device 200, for example. The 3D models 20-1 and 20-2 reproduced on the screen of the information processing device 200 can exist in any 3D space.

[0025] [B. Hardware configuration example] 2 is a schematic diagram showing an example of a hardware configuration of information processing device 100 constituting information processing system 1 according to the present embodiment. Typically, information processing device 100 can be realized using a general-purpose computer.

[0026] Referring to FIG. 2, the information processing device 100 includes, as its main hardware components, a CPU 102, a GPU 104, a main memory 106, a display 108, a network interface (I / F) 110, an input device 112, an optical drive 114, a camera interface (I / F) 118, and storage 120.

[0027] The CPU 102 and / or the GPU 104 are processors that execute the information processing method according to the present embodiment. A plurality of CPUs 102 and GPUs 104 may be provided, and each may have a plurality of cores.

[0028] The main memory 106 is a storage area that temporarily stores (or caches) program code, work data, etc. when the processor (CPU 102 and / or GPU 104) executes processing, and is composed of a volatile storage device such as a DRAM (Dynamic Random Access Memory) or an SRAM (Static Random Access Memory).

[0029] The display 108 is a display unit that outputs a user interface related to processing, processing results, etc., and is configured, for example, by an LCD (liquid crystal display) or an organic EL (electroluminescence) display.

[0030] The network interface 110 exchanges data with any information processing device connected to the network 2 .

[0031] The input device 112 is a device that accepts instructions and operations from the user, and is configured by, for example, a keyboard, a mouse, a touch panel, a pen, and the like.

[0032] The optical drive 114 reads information stored on an optical disk 116, such as a CD-ROM (compact disc read only memory) or a DVD (digital versatile disc), and outputs the information to other components. The optical disk 116 is an example of a non-transitory recording medium, and is distributed with any program stored therein in a non-volatile manner. The optical drive 114 reads the program from the optical disk 116 and installs it in the storage 120 or the like, causing the computer to function as the information processing device 100. Therefore, the subject matter of the present invention may also be the program itself installed in the storage 120 or the like, or a recording medium such as the optical disk 116 that stores a program for realizing the functions and processing according to the present embodiment.

[0033] FIG. 2 shows an optical recording medium such as an optical disk 116 as an example of a non-transitory recording medium, but is not limited to this. Alternatively, a semiconductor recording medium such as a flash memory, a magnetic recording medium such as a hard disk or storage tape, or a magneto-optical recording medium such as an MO (magneto-optical disk) may be used.

[0034] The camera interface 118 acquires images captured by the camera 140 and issues commands to the camera 140 regarding image capture.

[0035] The storage 120 stores programs and data necessary for the computer to function as the information processing device 100. For example, the storage 120 is configured by a nonvolatile storage device such as a hard disk or a solid state drive (SSD).

[0036] More specifically, the storage 120 stores an operating system (OS) (not shown), an initial model construction program 122 that implements processing for constructing an initial model (initial model construction stage), and a 3D model reproduction program 124 that implements processing for generating a 3D model (3D model reproduction stage). These information processing programs cause an information processing device 100, which is an example of a computer, to execute various processes according to this embodiment.

[0037] Additionally, the initial 3D shape data 162 and initial texture data 168 generated in the initial model construction stage may be stored in the storage 120. That is, the storage 120 corresponds to a memory unit that stores the 3D shape data 126 that indicates the 3D shape of the body and the initial texture data 168 (texture data) that indicates the texture of the body, which have been created in advance.

[0038] FIG. 2 shows an example in which information processing device 100 is configured using a single computer, but this is not limited to this. Multiple computers connected via a computer network may work together explicitly or implicitly to realize the information processing method according to this embodiment.

[0039] All or part of the functions realized by the processor (CPU 102 and / or GPU 104) executing the program may be realized using a hard-wired circuit such as an integrated circuit, for example, an application specific integrated circuit (ASIC) or a field-programmable gate array (FPGA).

[0040] Those skilled in the art will be able to implement information processing device 100 according to the present embodiment by appropriately using technology suited to the era in which the present invention is implemented.

[0041] Moreover, the hardware configuration of information processing device 200 constituting information processing system 1 is also the same as that in FIG. 2, and therefore detailed description thereof will not be repeated.

[0042] [C. Processing procedures for 3D model reproduction] To reproduce a 3D model, typically, a process of constructing an initial model (initial model construction stage) and a process of generating a 3D model (3D model reproduction stage) are performed.

[0043] In this specification, "texture data" is a general term for texture images and texture maps.

[0044] (c1: Initial model construction stage) Fig. 3 is a flowchart showing a processing procedure in the initial model construction stage of the information processing system 1 according to the present embodiment. Each process shown in Fig. 3 is typically realized by the processor of the information processing device 100 executing a program (initial model construction program 122 shown in Fig. 2).

[0045] 3, information processing device 100 acquires 2D video (one frame) captured by camera 140 (step S100). Information processing device 100 determines whether a predetermined number of frames of 2D video have been acquired (step S102). If the predetermined number of frames of 2D video have not been acquired (NO in step S102), the processes from step S100 onwards are repeated.

[0046] The information processing device 100 may start capturing images using the camera 140 upon receiving an explicit instruction from the user, or may repeat capturing images at a predetermined cycle.

[0047] Next, the information processing device 100 reconstructs body 3D shape data 160 indicating the 3D shape of the captured body based on the acquired multiple 2D images (multiple viewpoint images 144) (step S104). Then, the information processing device 100 flattens an area corresponding to the face area in a displacement map included in the body 3D shape data 160 (step S106). Finally, the information processing device 100 outputs the shape parameters and the flattened displacement map as initial 3D shape data 162 (step S108).

[0048] Furthermore, the information processing device 100 reconstructs body texture data (body texture image 1642 and body texture map 1644) indicating the texture of the body based on the acquired plurality of 2D images (multiple viewpoint images 144) (step S110).

[0049] Furthermore, the information processing device 100 reconstructs head 3D shape data 167 indicating the 3D shape of the captured head based on the acquired plurality of 2D images (multiple viewpoint images 144) (step S112).

[0050] In the information processing device 100, the texture integration unit 158 ​​integrates the body texture data 164 and the face texture data 166 to reconstruct the initial texture data 168 (initial texture image 1682 and initial texture map 1684) (step S114).

[0051] The order of execution of steps S104 to S108 and steps S110 to S114 does not matter, or these processes may be executed in parallel.

[0052] Finally, the information processing device 100 stores the initial 3D shape data 162 and the initial texture data 168 of the person as an initial model (step S116).

[0053] (c2: 3D model reproduction stage) Fig. 4 is a flowchart showing a processing procedure at a 3D model reproduction stage of information processing system 1 according to the present embodiment. Each process shown in Fig. 4 is typically realized by the processor of information processing device 100 executing a program (3D model reproduction program 124 shown in Fig. 2).

[0054] Referring to FIG. 4, information processing device 100 acquires 2D video (one frame) captured by camera 140 (step S200).

[0055] Information processing device 100 detects a face area included in the acquired 2D video (one frame) (step S202), and determines the position and size of the current window based on the past face area detection result (step S204).

[0056] The information processing device 100 reconstructs a facial texture image 1666 indicating the captured image of the face based on the 2D video of the portion corresponding to the determined window (step S206). That is, the information processing device 100 reconstructs the facial texture from the 2D video of the person captured by the camera 140.

[0057] Next, the information processing device 100 blends the facial texture image 1666 with the initial texture image 1682 (initial facial texture image 1686) reconstructed in the initial model construction stage to reconstruct a blended facial texture image 1824 (step S208). That is, the information processing device 100 blends the reconstructed facial texture image (facial texture image 1666) with the texture image (initial facial texture image 1686) included in texture data indicating the texture of the body created in advance to reconstruct a texture image (blended facial texture image 1824) of the person captured by the camera 140.

[0058] Furthermore, information processing device 100 reconstructs parameters (facial expression parameters 184) indicating facial expression, movement, and 3D shape based on the 2D video of the portion corresponding to the determined window (step S210). That is, information processing device 100 reconstructs the 3D shape of the face from the 2D video of the person captured by camera 140.

[0059] Furthermore, information processing device 100 estimates a body pose (posture) for each frame from the 2D video (one frame) (step S212). That is, information processing device 100 estimates the pose of a person from the 2D video of the person captured by camera 140. The estimated pose is output for each frame as body pose data 186.

[0060] The process of step S210 and the process of step S212 may be executed in parallel or serially, and the order of execution of the processes is not important.

[0061] The information processing device 100 inputs the body pose data 186 and the facial expression parameters 184 into the initial 3D shape data 162 reconstructed in the initial model construction stage, thereby reconstructing integrated 3D shape data 188 indicating a 3D shape obtained by integrating the 3D shape of the body and the 3D shape of the face (step S214). More specifically, the information processing device 100 reconstructs a 3D shape of the body (integrated 3D shape data 188) corresponding to the pose estimated based on 3D shape data indicating the 3D shape of the body (initial 3D shape data 162) created in advance. Furthermore, the information processing device 100 integrates the reconstructed 3D shape of the body and the reconstructed 3D shape of the face to reconstruct a 3D shape of the person captured by the camera 140 (integrated 3D shape data 188).

[0062] The processing of steps S202 to S208 and the processing of steps S210 to S214 may be executed in parallel or serially, and the processing may be executed in any order.

[0063] The information processing device 100 integrates the integrated 3D shape data 188 and the blended face texture image 1824 (step S216), and outputs a 3D model viewed from a specified viewpoint (step S218). That is, the information processing device 100 generates a 3D model 190 of the person captured by the camera 140, based on the 3D shape of the person captured by the camera 140 and the texture image of the person captured by the camera 140.

[0064] The processing of steps S200 to S218 is repeated for each frame.

[0065] [D. Details of the process in the initial model building stage] In the initial model construction stage of the information processing system 1 according to the present embodiment, an image of a person is captured to construct an initial model for reproducing a 3D model. The constructed initial model reflects information about the person's body and face.

[0066] Fig. 5 is a schematic diagram showing an example of a functional configuration for realizing the initial model construction stage of information processing system 1 according to the present embodiment. Fig. 6 is a diagram showing an example of data generated in the initial model construction stage of information processing system 1 according to the present embodiment.

[0067] 5 is typically realized by the processor of the information processing device 100 executing a program (initial model construction program 122 shown in FIG. 2). Referring to FIG. 5, the information processing device 100 includes an image acquisition unit 142, a body 3D shape reconstruction unit 150, a 3D shape correction unit 152, a body texture reconstruction unit 154, a face texture reconstruction unit 156, a head 3D shape reconstruction unit 157, and a texture integration unit 158.

[0068] (d1: video acquisition unit 142) The video acquisition unit 142 acquires 2D video captured by the camera 140. At this time, the video acquisition unit 142 acquires multiple 2D videos (multiple viewpoint videos 144) captured from multiple viewpoints of a person for whom a 3D model is to be reproduced. The camera 140 may be positioned at different positions relative to the person to capture images from multiple viewpoints, or the person may rotate their body while the camera 140 is fixed to capture images from multiple viewpoints. Alternatively, multiple cameras 140 may be provided, and multiple 2D videos may be acquired by capturing images of the person with each camera 140. FIG. 6(A) shows an example of multiple viewpoint videos 144 captured from eight viewpoints of a person.

[0069] It is preferable that the multi-viewpoint video 144 used to reconstruct the initial model is a 2D video of 5 to 10 frames.

[0070] (d2:Body 3D shape reconstruction part 150) The body 3D shape reconstruction unit 150 reconstructs the 3D shape of the body based on the multiple viewpoint images 144. That is, the body 3D shape reconstruction unit 150 reconstructs the 3D shape of the body from multiple 2D images of the person captured by the camera 140, and outputs body 3D shape data 160 that indicates the 3D shape of the captured body. Figure 6(B) shows an example of a visual representation of the reconstructed body 3D shape data 160.

[0071] More specifically, the body 3D shape reconstruction unit 150 reconstructs a model representing the 3D shape of a person's body from the 2D image. To reconstruct data representing such a 3D shape, a known algorithm such as "Tex2Shape" (Alldieck, T.; Pons-Moll, G.; Theobalt, C.; Magnor, M. Tex2Shape: Detailed Full Human Body Geometry From a Single Image. In 2019 IEEE / CVF International Conference on Computer Vision (ICCV); 2019; pp 2293-2303. https: / / doi.org / 10.1109 / ICCV.2019.00238.) can be used.

[0072] "Tex2Shape" outputs shape parameters (principal component features β that indicate shape) and a displacement map. When "Tex2Shape" outputs a model in SMPL format, it may be further converted to SMPL-X format, which has four times the resolution of the SMPL format.

[0073] The body 3D shape reconstruction unit 150 outputs information indicating the 3D shape of the person's body as body 3D shape data 160. The body 3D shape data 160 is typically made up of data in a mesh format.

[0074] (d3:3D shape modification section 152) The 3D shape correction unit 152 flattens the face region of the body 3D shape data 160 reconstructed by the body 3D shape reconstruction unit 150. In the 3D model reproduction stage, a separate model is used to reproduce the person's face, so it is preferable not to mutate the face region of the reconstructed 3D shape.

[0075] Therefore, the 3D shape correction unit 152 corrects the area in the displacement map corresponding to the estimated face area to a flat area. That is, the 3D shape correction unit 152 corrects the face area to a flat area without undulations. Such flattening makes it possible to more efficiently reproduce the person's head in the 3D model reproduction stage.

[0076] More specifically, the 3D shape correction unit 152 extracts a person included in the 2D image used to reconstruct the body 3D shape data 160 and estimates the extracted person's human body region (body parts). For example, regions corresponding to the person's face, hands, feet, etc. may be estimated. For such human body region estimation, a known algorithm such as "DensePose" (Gueler, R.A.; Neverova, N.; Kokkinos, I. DensePose: Dense Human Pose Estimation in the Wild. In 2018 IEEE / CVF Conference on Computer Vision and Pattern Recognition; 2018; pp. 7297-7306. https: / / doi.org / 10.1109 / CVPR.2018.00762.) may be used.

[0077] Then, the 3D shape correcting unit 152 updates the value of the area in the displacement map that corresponds to the estimated face area to a value that indicates a flat area.

[0078] Furthermore, since a person's fingers and the like are also likely to be modeled as a variation region, it is preferable to correct them to a flat region.

[0079] Finally, the 3D shape correction unit 152 outputs initial 3D shape data 162 that indicates the 3D shape in which the facial region has been flattened. Fig. 6(C) shows an example in which the initial 3D shape data 162 is visually expressed.

[0080] (d4: Body texture reconstruction unit 154) The body texture reconstruction unit 154 reconstructs the body texture from multiple 2D images (multiple viewpoint images 144) of a person captured by the camera 140. More specifically, the body texture reconstruction unit 154 reconstructs a body texture image 1642 and a body texture map 1644. The body texture image 1642 and the body texture map 1644 may be collectively referred to as "body texture data 164."

[0081] FIG. 6(D) shows an example of a body texture image 1642 and a body texture map 1644 (body texture data 164).

[0082] The body texture reconstructor 154 reconstructs the body texture data 164 according to the following process.

[0083] First, the body texture reconstruction unit 154 detects key points of a person from the 2D images included in the multi-viewpoint images 144. To detect such key points, a known algorithm such as "OpenPose" (Cao, Z.; Hidalgo, G.; Simon, T.; Wei, S.-E.; Sheikh, Y. OpenPose: Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields. IEEE Transactions on Pattern Analysis and Machine Intelligence 2021, 43 (1), 172-186. https: / / doi.org / 10.1109 / TPAMI.2019.2929257.) can be used.

[0084] Next, the body texture reconstruction unit 154 uses the detected keypoints to perform semantic segmentation on the 2D video to estimate the person's body regions (body parts). For such semantic segmentation, a known algorithm such as "PGN" (Gong, K.; Liang, X.; Li, Y.; Chen, Y.; Yang, M.; Lin, L. Instance-Level Human Parsing via Part Grouping Network. In Computer Vision - ECCV 2018; Ferrari, V., Hebert, M., Sminchisescu, C., Weiss, Y., Eds.; Lecture Notes in Computer Science; Springer International Publishing: Cham, 2018; pp. 805-822. https: / / doi.org / 10.1007 / 978-3-030-01225-0_47.) can be used.

[0085] Finally, the body texture reconstruction unit 154 uses the estimated human body region to reconstruct texture data (body texture image 1642 and body texture map 1644) from multiple 2D images (multi-viewpoint images 144). Reconstructing such texture data can use a known algorithm such as "Semantic Human Texture Stitching" (Alldieck, T.; Magnor, M.; Xu, W.; Theobalt, C.; Pons-Moll, G. Detailed Human Avatars from Monocular Video. In 2018 International Conference on 3D Vision (3DV); 2018; pp. 98-109. https: / / doi.org / 10.1109 / 3DV.2018.00022.).

[0086] "Semantic Human Texture Stitching" can output texture data in either the SMPL format or the SMPL-X format. As described above, when the SMPL-X formatted body 3D shape data 160 is used, the texture data also in the SMPL-X format is used.

[0087] Here, the SMPL format / SMPL-X format uses the same format as the texture map (UV mapping) included in the texture data.

[0088] (d5: Face texture reconstruction unit 156) The facial texture reconstruction unit 156 reconstructs facial texture from 2D images of a person captured by the camera 140. In the initial model construction stage, the facial texture reconstruction unit 156 reconstructs facial texture based on the 2D images included in the multi-viewpoint images 144. More specifically, the facial texture reconstruction unit 156 reconstructs a facial texture image 1662 and a facial texture map 1664. The facial texture image 1662 and the facial texture map 1664 may be collectively referred to as "facial texture data 166." As will be described later, the facial texture image is reconstructed by the texture integration unit 158, and therefore the facial texture image 1662 reconstructed by the facial texture reconstruction unit 156 may be discarded.

[0089] The facial texture reconstruction unit 156 reconstructs the facial texture data 166 according to the following process. That is, a known algorithm such as "DECA" (Feng, Y.; Feng, H.; Black, MJ; Bolkart, T. Learning an Animatable Detailed 3D Face Model from In-the-Wild Images. ACM Trans. Graph. 2021, 40 (4), 88:1-88:13. https: / / doi.org / 10.1145 / 3450626.3459936.) can be used.

[0090] "DECA" outputs FLAME model parameters (indicating the shape and expression of the face) for reproducing a person's face and texture data conforming to the FLAME format. In this way, facial texture data 166 conforming to the FLAME format is output from a 2D image of a person captured by a camera. As will be described later, the texture integration unit 158 ​​integrates the two by applying the facial texture data 166 conforming to the FLAME format to the body texture data 164.

[0091] The facial texture reconstruction unit 156 also reconstructs the facial texture data 166 for each frame during the 3D model reproduction stage.

[0092] (d6: Head 3D shape reconstruction part 157) The head 3D shape reconstruction unit 157 reconstructs the 3D shape of the head from multiple 2D images (multiple viewpoint images 144) of the person captured by the camera 140. That is, the head 3D shape reconstruction unit 157 reconstructs head 3D shape data 167 that indicates the 3D shape of the captured head.

[0093] The head 3D shape reconstruction unit 157 reconstructs a model representing the 3D shape of a person's head from the 2D video using the same algorithm as the body 3D shape reconstruction unit 150. The head 3D shape reconstruction unit 157 outputs head 3D shape data 167 as information representing the 3D shape of the head. The head 3D shape data 167 is typically made up of mesh-format data.

[0094] (d7: Texture integration part 158) The texture integration unit 158 ​​integrates the body texture data 164 and the face texture data 166 to reconstruct initial texture data 168 (initial texture image 1682 and initial texture map 1684). The texture integration unit 158 ​​integrates the body texture data 164 and the face texture data 166 based on the correspondence between the body 3D shape data 160 and the head 3D shape data 167.

[0095] As shown in FIG. 6(E), each of the initial texture image 1682 and the initial texture map 1684 is made up of a portion for the head including the face and a portion for the body other than the head.

[0096] More specifically, the initial texture image 1682 consists of an initial face texture image 1686 reconstructed by processing as described below, and a modified body texture image 1642A obtained by invalidating the head portion image 1642H corresponding to the head from the body texture image 1642.

[0097] The initial texture map 1684 is made up of the face texture map 1664 and a modified body texture map 1644A obtained by invalidating the head portion map 1644H corresponding to the head from the body texture map 1644.

[0098] As shown in FIG. 6(E), the initial texture data 168 (texture data) includes a texture image corresponding to the 3D shape of the body to be reconstructed (modified body texture image 1642A), a texture image corresponding to the 3D shape of the face to be reconstructed (initial face texture image 1686), a texture map corresponding to the 3D shape of the body to be reconstructed (modified body texture map 1644A), and a texture map corresponding to the 3D shape of the face to be reconstructed (face texture map 1664).

[0099] Note that Figure 6(E) shows a deleted state as an example of invalidating the head portion image 1642H and head portion map 1644H, but they do not necessarily have to be deleted; they can simply be set so that they are not used in processing.

[0100] 7 is a schematic diagram for explaining the texture integration process in the initial model construction stage of information processing system 1 according to the present embodiment. Texture integration unit 158 ​​executes the following five processes.

[0101] (1) Alignment of the body 3D shape data 160 and the head 3D shape data 167 The texture integration unit 158 ​​aligns the two pieces of shape data by mapping the body 3D shape data 160 and the head 3D shape data 167 onto a common 3D space. Here, the body 3D shape data 160 and the head 3D shape data 167 represent 3D shapes reconstructed from the same person, and therefore are considered to have substantially the same topology.

[0102] The texture integration unit 158 ​​focuses on characteristic parts of the common face (eyes, nose, etc.) and maps each shape data onto a common 3D space so that the focused parts have the same coordinates. The process for achieving such alignment uses a coordinate transformation matrix including operations such as translation, rotation, and scale.

[0103] (2) Determining the correspondence between meshes Next, the texture integration unit 158 ​​determines the correspondence between the meshes between the two aligned shape data. That is, the texture integration unit 158 ​​determines the correspondence between the meshes (for example, a set of triangles defined by three vertices) included in the body 3D shape data 160 and the meshes included in the head 3D shape data 167.

[0104] More specifically, for each mesh included in the aligned body 3D shape data 160, the texture integration unit 158 ​​searches for the mesh that is closest to it among the meshes included in the aligned head 3D shape data 167. Finally, the texture integration unit 158 ​​determines the correspondence between the meshes (for example, an array indicating the correspondence between the index indicating each mesh included in the body 3D shape data 160 and the index indicating each mesh included in the head 3D shape data 167).

[0105] In this way, the texture integration unit 158 ​​determines the correspondence between the reconstructed 3D shape of the body (3D body shape data 160) and the reconstructed 3D shape of the head (3D head shape data 167).

[0106] (3) Determining the correspondence between texture maps Next, the texture synthesis unit 158 ​​determines a correspondence between the body texture map 1644 and the face texture map 1664 .

[0107] The correspondence (one-to-one) between the body 3D shape data 160 and the body texture map 1644 is known, and similarly, the correspondence (one-to-one) between the head 3D shape data 167 and the face texture map 1664 is also known. Since the correspondence (one-to-one) between the body 3D shape data 160 and the head 3D shape data 167 is determined by the above-mentioned processing, the texture integration unit 158 ​​determines the correspondence between the texture maps using this correspondence of the shape data.

[0108] In this way, the texture integration unit 158 ​​determines the correspondence between the texture map (body texture map 1644) corresponding to the 3D shape of the body (body 3D shape data 160) and the texture map (face texture map 1664) corresponding to the 3D shape of the head (head 3D shape data 167) based on the correspondence between the determined 3D shapes.

[0109] (4) Generation of initial facial texture image Next, the texture synthesis unit 158 ​​generates an initial facial texture image 1686 based on the correspondence between the body texture map 1644 and the facial texture map 1664 .

[0110] More specifically, the texture integration unit 158 ​​determines coordinates in the body texture map 1644 that correspond to each coordinate in the face texture map 1664, and applies the pixel values ​​of the body texture image 1642 at the determined coordinates in the body texture map 1644 as new pixel values ​​of the face texture image. That is, the body texture image 1642 is mapped based on the correspondence between the body texture map 1644 and the face texture map 1664, thereby generating an initial face texture image 1686, which is a new face texture image.

[0111] In this way, the texture integration unit 158 ​​generates a texture image (initial face texture image 1686) corresponding to the 3D shape of the head (head 3D shape data 167) from a texture image (body texture image 1642) corresponding to the 3D shape of the body (body 3D shape data 160) based on the correspondence of the determined texture maps.

[0112] (5) Data integration Finally, the texture synthesis unit 158 ​​reconstructs the initial texture data 168 (initial texture image 1682 and initial texture map 1684).

[0113] More specifically, the texture integration unit 158 ​​invalidates the head partial image 1642H corresponding to the head in the body texture image 1642, and then combines it with the generated initial face texture image 1686. The initial texture image 1682 corresponds to the modified body texture map 1644A and the initial face texture image 1686, adjusted to the same scale, and arranged adjacent to each other.

[0114] Furthermore, the texture integration unit 158 ​​invalidates a head portion map 1644H corresponding to the head of the body texture map 1644, and then combines it with the face texture map 1664. The initial texture map 1684 corresponds to the modified body texture image 1642A and the face texture map 1664, adjusted to the same scale, and arranged adjacent to each other.

[0115] Texture data in the SMPL-X format can be reformatted into the FLAME format by applying a specific scaling factor. That is, since there is a one-to-one correspondence between texture maps in the SMPL-X format and those in the FLAME format, the magnification factor when expanding a texture image can be uniquely determined based on the correspondence between the formats.

[0116] In this way, the texture integration unit 158 ​​integrates the texture image (body texture image 1642) corresponding to the 3D shape of the body (body 3D shape data 160) with the texture image (initial face texture image 1686) corresponding to the 3D shape of the head (head 3D shape data 167), and also integrates the texture map (modified body texture map 1644A) corresponding to the 3D shape of the body with the texture map (face texture map 1664) corresponding to the 3D shape of the head, thereby generating initial texture data 168.

[0117] The initial texture data 168 (initial texture image 1682 and initial texture map 1684) consists of a portion for the head, including the face, and a portion for the body other than the head. By preparing more textures for the head, including the face, it is possible to improve the reproducibility of facial expressions and movements (gestures) even when capturing images using a single camera.

[0118] The above processing completes the initial model construction process.

[0119] [E. Details of the process in the 3D model reproduction stage] In the 3D model reproduction stage of information processing system 1 according to the present embodiment, a 3D model is reproduced from 2D video (one frame) of a person captured by one camera 140. By updating the 3D model for each frame of the 2D video, the movements and facial expressions of the person can be reproduced as a video.

[0120] Fig. 8 is a schematic diagram showing an example of a functional configuration for realizing a 3D model reproduction stage of information processing system 1 according to the present embodiment. Fig. 9 is a diagram showing an example of data generated in the 3D model reproduction stage of information processing system 1 according to the present embodiment.

[0121] 8 are typically realized by the processor of the information processing device 100 executing a program (the 3D model reproduction program 124 shown in FIG. 2). Note that some of the processing may be handled by the information processing device 200.

[0122] Referring to FIG. 8, the information processing device 100 includes a stabilization unit 170, a facial texture reconstruction unit 156, a texture image blending unit 172, a facial shape reconstruction unit 174, a pose estimation unit 176, a shape integration unit 178, and a 3D model generation unit 180.

[0123] (e1: Stabilization part 170) The stabilization unit 170 detects a facial region included in the 2D image captured by the camera 140 and stabilizes the detected facial region over time. The stabilization unit 170 outputs a partial image corresponding to the temporally stabilized facial region to the facial texture reconstruction unit 156 and the facial shape reconstruction unit 174. That is, a partial image corresponding to a window set in the 2D image of a person captured by the camera 140 is input to the facial texture reconstruction unit 156 and the facial shape reconstruction unit 174.

[0124] The stabilization unit 170 temporally smooths the position and size of the face region 163 (window) extracted from the 2D video 146. Fig. 9(A) shows an example of the process of extracting the face regions 163A and 163B from the 2D video 146. The ranges of the face regions 163A and 163B can be determined by known image recognition processing.

[0125] Let us consider a case where a publicly known algorithm such as "DECA" is used to reproduce a person's face for each frame. "DECA" can reproduce a face for each frame, but if the size and position of the face area 163 are determined for each frame, fluctuations and discontinuities may occur in the reproduced face when viewed between frames.

[0126] Generally, the positions of facial key points (e.g., eyes) detected from a 2D image (one frame) may change from frame to frame, and therefore the position and size of the window determined based on the detected key points may also change from frame to frame.

[0127] Therefore, stabilization unit 170 stabilizes the reproduced face by temporally smoothing the position and size of the window. That is, stabilization unit 170 sets a window by temporally smoothing the position of the person in the 2D video.

[0128] More specifically, the stabilization unit 170 employs a window having a certain size that can cover the entire face of a person, and sets the window at a position based on a specific key point, for example, the window can be set with the tip of the nose as the center.

[0129] For example, if a person moves within the window set in the previous frame in the next frame, the stabilization unit 170 sets the position of the window in the next frame based on the average position of specific key points detected from the past n frames. Also, if a person moves closer to or farther away from the camera 140 in the next frame, the stabilization unit 170 changes the size of the window accordingly based on the moving average of the window sizes in the past n frames.

[0130] By performing such processing, it is possible to reduce the degree to which discontinuity occurs between frames when the window is made to follow the movement of a person in the 2D video image 146.

[0131] If the person moves too quickly and moves outside the window, the size and position of the window are reset and set again. In this case, discontinuity may occur in the reproduced face, so additional processing may be performed to reduce the sense of incongruity.

[0132] By adopting the above-described processing, the position and size of the sequentially extracted face area 163 (window) do not change significantly between frames, thereby reducing discontinuities that occur in the reconstructed face shape.

[0133] (e2: Face texture reconstruction unit 156) The facial texture reconstruction unit 156 reconstructs the facial texture based on the image of the facial region 163 extracted from the 2D image 146. More specifically, the facial texture reconstruction unit 156 reconstructs a facial texture image 1666. The facial texture reconstruction unit 156 is substantially the same as the facial texture reconstruction unit 156 shown in Fig. 5, and therefore detailed description will not be repeated. Fig. 9(B) shows an example of the reconstructed facial texture image 1666.

[0134] It should be noted that the facial texture reconstruction unit 156 also reconstructs a facial texture map, but this is not necessarily required by the texture image blending unit 172 and may therefore be discarded.

[0135] (e3: Texture image blending part 172) The texture image blending unit 172 blends the initial texture image 1682 reconstructed in the initial model construction stage and the facial texture image 1666 reconstructed by the facial texture reconstruction unit 156 to reconstruct a blended facial texture image 1824. That is, the texture image blending unit 172 blends the texture image (initial texture image 1682) of the reconstructed face with the texture image (initial texture image 1682) included in the texture data (initial texture data 168) to reconstruct a texture image (blended facial texture image 1824) of the person captured by the camera 140.

[0136] FIG. 9(C) shows an example of a reconstructed blended face texture image 1824.

[0137] Fig. 10 is a schematic diagram for explaining the blending process in information processing system 1 according to the present embodiment. Referring to Fig. 10, texture image blending unit 172 generates modified face texture image 1686A by blending face texture image 1666 with initial face texture image 1686 of initial texture image 1682 (initial face texture image 1686 and modified body texture image 1642A) using mask 1826.

[0138] That is, blending processing is performed on the initial texture image 1682 for the initial face texture image 1686 to generate a blended face texture image 1824. At this time, the texture image blending unit 172 superimposes the result of transmitting the mask 1826 from the face texture image 1666 onto the initial face texture image 1686. In this way, the texture image blending unit 172 superimposes the result of transmitting the mask from the texture image (blended face texture image 1824) of the person captured by the camera 140 onto the initial face texture image 1686 included in the initial texture data 168.

[0139] The mask 1826 may be generated, for example, by assigning the reliability of each pixel of the facial texture data 166 reconstructed by the facial texture reconstruction unit 156 as an intensity (transparency).

[0140] Alternatively, the mask 1826 may be generated based on the facial texture image 1666. More specifically, among the pixels included in the facial texture image 1666, pixels whose pixel values ​​exceed a predetermined threshold are assigned a value of "1" (transmitted), and other pixels are assigned a value of "0" (blocked). Next, a minimization filter is applied using a square window, and a blurring filter (e.g., a Gaussian filter or a box filter) is applied to the edges.

[0141] By using such a mask 1826, it is possible to achieve blending that gradually changes the periphery of the face texture image 1666 that is overlaid on the initial face texture image 1686. In other words, the mask 1826 is configured so that its transparency changes continuously.

[0142] By using such blending, facial expressions that reflect the image of the current frame can be reproduced in real time, while hairstyles and the like can be reproduced stably using the initial texture image 1682.

[0143] That is, in the 3D model reproduction stage, information such as facial expressions reconstructed from the image of each frame is used and reflected in the 3D model in real time, while for the texture of areas other than the facial area of ​​the head, which may not necessarily be reconstructed from the image of the frame, information from the initial facial texture image 1686 is reflected in the 3D model.

[0144] (e4: face shape reconstruction unit 174) The facial shape reconstruction unit 174 reconstructs parameters (facial expression parameters 184) indicating facial expression, movement, and 3D shape based on the image of the face region 163 extracted from the 2D image 146. In other words, the facial shape reconstruction unit 174 corresponds to a facial shape reconstruction unit that reconstructs the 3D shape of the face from the 2D image of the person captured by the camera 140. The facial shape reconstruction unit 174 may employ a known algorithm such as "DECA" as described above.

[0145] FIG. 9(D) shows an example of a visual representation of parameters (facial expression parameters 184) indicating the expression, movement, and 3D shape of the reconstructed face.

[0146] (e5: Pose estimation unit 176) Pose estimation unit 176 estimates a body pose (posture) for each frame from 2D video 146. That is, pose estimation unit 176 estimates the pose of a person from a 2D video of the person captured by camera 140. Body pose data 186 is output from pose estimation unit 176 for each frame. Typically, body pose data 186 includes information such as the angle of each joint. Note that a known pose estimation algorithm can be used for pose estimation unit 176.

[0147] FIG. 9(E) shows an example of a visual representation of the pose estimation process and estimated body pose data 186.

[0148] (e6: Shape integration part 178) The shape integration unit 178 reconstructs a 3D shape of the body corresponding to the captured 2D image 146 by inputting body pose data 186 and facial expression parameters 184 into the initial 3D shape data 162 reconstructed in the initial model construction stage.

[0149] More specifically, shape integration unit 178 reconstructs a 3D shape of the body corresponding to the pose specified by body pose data 186 and the expression defined by facial expression parameters 184, based on initial 3D shape data 162. In this way, integrated 3D shape data 188 is reconstructed, which indicates a 3D shape obtained by integrating the 3D shape of the body and the 3D shape of the face.

[0150] Furthermore, the shape integration unit 178 may reconstruct the integrated 3D shape data 188 using not only the initial 3D shape data 162 but also 3D shape data in which the head 3D shape data 167 is incorporated into the initial 3D shape data 162.

[0151] Furthermore, the shape integration unit 178 may determine the correspondence based on the initial texture map 1684 (modified body texture map 1644A and face texture map 1664), and then integrate the 3D shape of the body and the 3D shape of the face.

[0152] In this way, the shape integration unit 178 reconstructs a 3D shape of the body corresponding to the pose estimated based on the initial 3D shape data 162 (3D shape data), and integrates the reconstructed 3D shape of the body with the 3D shape of the face reconstructed based on the facial expression parameters 184 to reconstruct a 3D shape (integrated 3D shape data 188) of the person captured by the camera 140. Furthermore, the shape integration unit 178 can improve the reproduction accuracy by integrating the reconstructed 3D shape of the body and the reconstructed 3D shape of the face based on the texture map (initial texture map 1684) included in the initial texture data 168.

[0153] FIG. 9(F) shows an example of a visual representation of the reconstructed integrated 3D shape data 188.

[0154] (e7: 3D model generation unit 180) The 3D model generation unit 180 integrates a 3D shape based on the integrated 3D shape data 188 with the blended face texture image 1824. The 3D model generation unit 180 also outputs a 3D model 190 seen from a specified viewpoint.

[0155] In this way, the 3D model generation unit 180 generates a 3D model 190 of the person captured by the camera 140 based on the 3D shape of the person captured by the camera 140 (integrated 3D shape data 188) and the texture image of the person captured by the camera 140 (blended face texture image 1824).

[0156] Note that the initial texture map 1684 may be referenced when combining the integrated 3D shape data 188 and the blended face texture image 1824 (mapping the texture image). That is, the 3D model generation unit 180 may integrate the 3D shape of the person (integrated 3D shape data 188) and the texture image of the person (blended face texture image 1824) based on the initial texture map 1684 included in the initial texture data 168 (texture data).

[0157] Fig. 9(F) shows an example of a visual representation of the state of the 3D model 190 viewed from multiple viewpoints. Note that the 3D model generation unit 180 does not have to simultaneously display the 3D model viewed from multiple viewpoints as shown in Fig. 9(F), and outputs the 3D model viewed from one specified viewpoint.

[0158] [F. Variations] Although the configuration has been exemplified in which the process of constructing an initial model (initial model construction stage) and the process of generating a 3D model (3D model reproduction stage) are executed by the same information processing device 100, some of the processes may be executed by different information processing devices.

[0159] Furthermore, the initial model (initial 3D shape data 162 and initial texture data 168) may be constructed in advance and used appropriately when it is necessary to reproduce the 3D model.

[0160] 11 is a schematic diagram showing another example of the system configuration of information processing system 1 according to the present embodiment. Referring to Fig. 11, for example, server device 300 stores initial 3D shape data 162 and initial texture data 168 for each user in advance.

[0161] In response to requests from the information processing devices 100-3 and 100-4, the server device 300 provides the specified initial 3D shape data 162 and initial texture data 168. Each of the information processing devices 100-3 and 100-4 executes a process of generating a 3D model (3D model reproduction stage) using the initial 3D shape data 162 and initial texture data 168 provided by the server device 300.

[0162] It should be noted that the initial 3D shape data 162 and the initial texture data 168 do not necessarily have to be created based on 2D video images of a user using the information processing device 100. As described above, in the 3D model reproduction stage, texture images reconstructed by imaging a person are blended, so that even if the initial 3D shape data 162 and initial texture data 168 generated from another person are used, a 3D model of the person can be reproduced.

[0163] 1 may cooperate with the information processing device 200 to execute the process of constructing an initial model (initial model construction stage) and the process of generating a 3D model (3D model reproduction stage). The processes to be handled by each of the information processing devices may be designed arbitrarily.

[0164] [G. Summary] When reproducing a 3D model, the information processing system 1 according to the present embodiment can generate a 3D model of a person from one frame of 2D video, rather than from multiple 2D videos captured by multiple cameras. By reconstructing the shape and texture of the body and face, facial expressions and gestures can be reproduced with higher accuracy.

[0165] Furthermore, when reproducing a 3D model, the 3D model can be generated from one frame of 2D video captured by a camera, which reduces the processing load compared to using multiple 2D videos captured by multiple cameras, thereby enabling the 3D model to be reproduced in real time.

[0166] The embodiments disclosed herein should be considered to be illustrative in all respects and not restrictive. The scope of the present invention is defined by the claims, not by the description of the above embodiments, and is intended to include all modifications within the meaning and scope of the claims. [Explanation of symbols]

[0167] 1 Information processing system, 2 Network, 10 Person, 20,190 3D model, 100,200 Information processing device, 102 CPU, 104 GPU, 106 Main memory, 108 Display, 110 Network interface, 112 Input device, 114 Optical drive, 116 Optical disk, 118 Camera interface, 120 Storage, 122 Initial model construction program, 124 Reproduction program, 126,162 3D shape data, 140 Camera, 142 Image acquisition unit, 144 Multiple viewpoint image, 146 2D image, 150 Body 3D shape reconstruction unit, 152 3D shape correction unit, 154 Body texture reconstruction unit, 156 Face texture reconstruction unit, 157 Head 3D shape reconstruction unit, 158 Texture integration unit, 160 Body 3D shape data, 163,163A,163B Face region, 164 body texture data, 166 face texture data, 167 head 3D shape data, 168 initial texture data, 170 stabilization unit, 172 texture image blending unit, 174 face shape reconstruction unit, 176 pose estimation unit, 178 shape integration unit, 180 3D model generation unit, 184 facial expression parameters, 186 body pose data, 188 integrated 3D shape data, 300 server device, 1642 body texture image, 1642A corrected body texture image, 1642H head part image, 1644 body texture map, 1644A corrected body texture map, 1644H head part map, 1662, 1666 face texture image, 1664 face texture map, 1682 initial texture image, 1684 initial texture map, 1686 initial face texture image, 1686A Corrected face texture images, 1824 blended face texture images, 1826 masks.

Claims

1. A camera and a storage unit that stores 3D shape data indicating a 3D shape of a body and texture data indicating a texture of the body, both of which have been created in advance; a facial texture reconstruction unit that reconstructs a facial texture image from the 2D image of the person captured by the camera; a face shape reconstruction unit that reconstructs a 3D shape of a face from a 2D image of the person captured by the camera; a pose estimation unit that estimates a pose of a person from a 2D image of the person captured by the camera; a shape integration unit that reconstructs a 3D shape of a body corresponding to the estimated pose based on the 3D shape data, and integrates the reconstructed 3D shape of the body and the reconstructed 3D shape of the face to reconstruct a 3D shape of the person captured by the camera; a texture reconstruction unit that reconstructs a texture image of the person captured by the camera by blending the reconstructed facial texture image with a texture image included in the texture data; and a model generation unit that generates a 3D model of the person captured by the camera based on a 3D shape of the person captured by the camera and a texture image of the person captured by the camera.

2. The texture data is a texture image corresponding to the reconstructed 3D shape of the body, and a texture image corresponding to the reconstructed 3D shape of the face; a texture map corresponding to the reconstructed 3D shape of the body, and a texture map corresponding to the reconstructed 3D shape of the face; a body shape reconstruction unit that reconstructs a 3D shape of a body from a plurality of 2D images of the person captured by the camera; a body texture reconstruction unit that reconstructs a body texture from a plurality of 2D images of the person captured by the camera; a head shape reconstruction unit that reconstructs a 3D shape of a head from a plurality of 2D images of a person captured by the camera; 2. The information processing system according to claim 1, further comprising a texture integration unit that determines a correspondence relationship between the reconstructed 3D shape of the body and the reconstructed 3D shape of the head, determines a correspondence relationship between a texture map corresponding to the 3D shape of the body and a texture map corresponding to the 3D shape of the head based on the determined correspondence relationship of the 3D shapes, and generates a texture image corresponding to the 3D shape of the head from a texture image corresponding to the 3D shape of the body based on the determined correspondence relationship of the texture maps.

3. The information processing system according to claim 2 , wherein the shape integration unit integrates the reconstructed 3D shape of the body and the reconstructed 3D shape of the face based on a texture map included in the texture data.

4. The information processing system according to claim 2 , wherein the model generation unit integrates the 3D shape of the person with a texture image of the person based on a texture map included in the texture data.

5. 5. The information processing system according to claim 1, wherein the texture reconstruction unit superimposes a texture image of the person captured by the camera that is transmitted through a mask onto a texture image included in the texture data.

6. a partial image corresponding to a window set in the 2D image of the person captured by the camera is input to the face texture reconstruction unit and the face shape reconstruction unit; The information processing system according to claim 1 , further comprising a stabilization unit that sets the window by temporally smoothing the position of the person in the 2D video.

7. Reconstructing a facial texture image from a 2D image of the person captured by a camera; Reconstructing a 3D shape of a face from a 2D image of the person captured by the camera; estimating a pose of a person from a 2D image of the person captured by the camera; Reconstructing a 3D shape of the body corresponding to the estimated pose based on 3D shape data indicating a 3D shape of the body that has been created in advance; Reconstructing a 3D shape of the person captured by the camera by integrating the reconstructed 3D shape of the body and the reconstructed 3D shape of the face; a step of blending the reconstructed facial texture image with a texture image included in texture data indicating a body texture that has been created in advance, thereby reconstructing a texture image of the person captured by the camera; generating a 3D model of the person captured by the camera based on a 3D shape of the person captured by the camera and a texture image of the person captured by the camera.

8. An information processing program for causing a computer to execute the method according to claim 7.

Citation Information

Patent Citations

  • Animation creating system

    JP2002269580A

  • Avatar generation method, program, avatar generation system, and avatar display method

    WO2021261188A1