Information processing device, information processing method, and program
The information processing device enhances three-dimensional reconstruction by optimizing clothing movements using physical simulation and external force variables, addressing the limitations of conventional techniques in representing clothing and hair on natural persons.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-04-08
AI Technical Summary
Conventional three-dimensional reconstruction techniques struggle to accurately reproduce the movements of clothing and hair on natural persons due to reliance on body posture, failing to account for environmental factors like wind, leading to inaccurate representations.
An information processing device that utilizes time-series images to optimize a three-dimensional model by incorporating physical simulation of clothing and external force variables, using deep neural networks to generate a three-dimensional model that accurately represents the movement of clothing and environmental influences.
The method enables the generation of three-dimensional objects that realistically depict the movements of clothing and hair on natural persons, improving accuracy and naturalness in three-dimensional representations.
Smart Images

Figure 2026060577000001_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to three-dimensional reconstruction technology.
Background Art
[0002] There is a technology for estimating spatial information expressed by a Radiance Field or the like in an imaging space by using a plurality of imaging images obtained by imaging from a plurality of directions. According to this technology, it is possible to reconstruct a virtual three-dimensional object corresponding to an object existing in the imaging space in a virtual space. The technology of reconstructing an actual object as a virtual three-dimensional object is also called three-dimensional reconstruction technology.
[0003] Non-Patent Document 1 discloses a three-dimensional reconstruction technology called NeRF (Neural Radiance Fields) that utilizes deep learning. In the NeRF disclosed in Non-Patent Document 1, a plurality of imaging images obtained by imaging an object to be three-dimensionally reconstructed from a plurality of directions are input. The parameters included in NeRF are optimized for each imaging image by volume rendering using a radiance field estimated during learning. According to NeRF, since transmissive representation in three-dimensional space is possible, the appreciation quality regarding virtual three-dimensional objects can be improved. However, while the type of object is not limited in NeRF, in order to obtain a highly accurate optimization result, it is necessary to image the object using an enormous number of calibrated imaging devices, such as dozens or more.
[0004] Non-Patent Document 2 discloses a technique (hereinafter referred to as "prior art") for three-dimensional reconstruction of a natural person from multiple frames obtained by capturing a natural person as a moving image using a single imaging device, with the target of three-dimensional reconstruction being a natural person. Specifically, in the prior art, first, a three-dimensional space (observation space) and a three-dimensional space of the T-pose (normal space) are defined for the state of each frame. Subsequently, in the prior art, warping is performed between the observation space and the normal space by utilizing human body posture information. According to the prior art, by performing warping between spaces using human body posture information, optimization such as NeRF disclosed in Non-Patent Document 1 can be performed using multiple frames captured as moving images of the natural person to be reconstructed in three dimensions. [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] Ben Mildenhall, et al., “Nerf: Representing scenes as neural radiance fields for view synthesis”, 2020, European Conference on Computer Vision (ECCV), [online], [Retrieved August 30, 2024], Internet<https: / / www.ecva.net / papers / eccv_2020 / papers_ECCV / papers / 123460392.pdf> [Non-Patent Document 2] Chung-Yi Weng, et al., “Humannerf: Free viewpoint rendering of moving people from monocular video”, 2022, Proceedings of the IEEE / CVF conference on computer vision and pattern recognition (CVPR), [online], [Retrieved August 30, 2024], Internet <https: / / openaccess.thecvf.com / content / CVPR2022 / papers / Weng_HumanNeRF_Free-Viewpoint_Rendering_of_Moving_People_From_Monocular_Video_CVPR_2022_paper.pdf> [Overview of the project] [Problems that the invention aims to solve]
[0006] Conventional techniques use warping for normalization, which depends on the posture of the human body. Therefore, they do not take into account the movement of clothing or hair, which cannot be represented by posture alone. As a result, the accuracy of reproducing natural people wearing clothing that moves due to factors such as wind may decrease.
[0007] This disclosure aims to provide a technology that can improve the accuracy of generating three-dimensional objects corresponding to natural persons wearing clothing, etc. [Means for solving the problem]
[0008] The information processing device according to this disclosure includes: an image acquisition means for acquiring a time-series image obtained by imaging an object including a first object and a second object associated with the first object; an optimization means for optimizing a three-dimensional model corresponding to the object, parameters relating to the physical simulation of the second object, and variables relating to external forces that may affect the second object, based on the time-series image; and a model generation means for generating the three-dimensional model based on the optimization results obtained by the optimization. [Effects of the Invention]
[0009] According to this disclosure, it is possible to generate three-dimensional objects that accurately reproduce the movements of objects attached to animals, such as clothing worn by natural people. [Brief explanation of the drawing]
[0010] [Figure 1] This figure illustrates an example of the difference between a three-dimensional human model generated by conventional technology and a three-dimensional human model generated by the technology disclosed herein. [Figure 2] This is a block diagram showing an example of the hardware configuration of an information processing device according to Embodiment 1. [Figure 3] This is a block diagram showing an example of the functional configuration of the information processing device according to Embodiment 1. [Figure 4] This flowchart shows an example of the processing flow of the information processing device according to Embodiment 1. [Figure 5] This is a diagram illustrating the process of generating a three-dimensional human model according to Embodiment 1. [Figure 6] This is a diagram illustrating an example of a learning method according to Embodiment 1. [Figure 7] This is a diagram illustrating an example of a learning method according to Embodiment 1. [Figure 8] This is a diagram illustrating an example of the deformation process according to Embodiment 1. [Figure 9] This is a diagram illustrating an example of the deformation process according to Embodiment 1. [Figure 10] This is a block diagram showing an example of the internal configuration of the clothing simulation DNN according to Embodiment 1. [Figure 11] This figure shows an example of animal posture information according to Embodiment 1. [Figure 12] This figure shows an example of the editing screen for a three-dimensional human model according to Embodiment 2. [Figure 13] This figure shows an example of the process of producing video content according to Embodiment 2. [Figure 14] This is a sequence diagram showing an example of the processing flow in the production of video content according to Embodiment 2. [Figure 15] This is a sequence diagram showing an example of the processing flow in the production of video content according to Embodiment 2. [Figure 16] This figure shows an example of the editing screen according to Embodiment 2. [Figure 17] This figure shows an example of the editing screen according to Embodiment 2. [Figure 18] This figure shows an example of a three-dimensional human model placed in a virtual space according to Embodiment 2. [Modes for carrying out the invention]
[0011] The embodiments relating to this disclosure will be described below with reference to the drawings. The following embodiments are not intended to limit the technology of this disclosure, and not all combinations of features described in these embodiments are essential to the solutions of this disclosure. The configuration of the embodiments may be modified or changed as appropriate depending on the specifications of the device to which the technology of this disclosure is applied, or various conditions such as usage conditions or usage environment. In addition, in the following embodiments, the same or similar components are denoted by the same reference numerals to omit redundant explanations.
[0012] One of the technologies related to the metaverse is the technology to generate three-dimensional human avatars (hereinafter referred to as "three-dimensional human avatars") that realistically reproduce the appearance and movements of natural people. One method of generating three-dimensional human avatars is to use imaging information, including images obtained by photographing real natural people, to generate a three-dimensional human avatar that closely resembles that natural person. As the technology to three-dimensionally reconstruct virtual objects corresponding to natural people as three-dimensional human avatars develops further, natural people will be able to experience another life in the metaverse space that is different from the real world.
[0013] Various studies have been conducted on three-dimensional reconstruction techniques in the fields of computer graphics and computer vision. Three-dimensional reconstruction techniques include geometric methods using Shape-from-X techniques, such as the multi-view stereo method and the viewing volume cross-section method. These geometric methods are proposed as solutions to the inverse problem of poorly configured images by mathematically formalizing the projection process from three dimensions to two dimensions. To obtain three-dimensional reconstruction results with good viewing quality using the aforementioned methods, it was necessary to set up numerous imaging devices, and each imaging device needed to be precisely calibrated. Furthermore, these methods have limitations in reconstructing shapes or textures in many scenes. For example, in three-dimensional reconstruction using the viewing volume cross-section method, concave shapes may not be reproducible. Also, in three-dimensional reconstruction using the multi-view stereo method, the accuracy of the three-dimensional reconstruction could decrease if textures that failed to match between the stereoscopic images were input.
[0014] On the other hand, with the development of deep learning technology, techniques such as NeRF, which generate three-dimensional reconstructions using fewer input images compared to the geometric methods described above, have been proposed. It is expected that techniques utilizing deep learning technology will yield three-dimensional reconstruction results with a level of accuracy that was difficult to achieve with the geometric methods described above.
[0015] [Embodiment 1] This embodiment describes a method for generating an editable three-dimensional human model (hereinafter referred to as "three-dimensional human model") by taking moving images obtained by imaging a natural person as the target of the three-dimensional human avatar generation as input. Here, editing the three-dimensional human model means importing the generated three-dimensional human model into a three-dimensional CG (Computer Graphics) authoring tool or game engine, etc., and modifying or adding animation. In particular, this embodiment describes the generation of a three-dimensional human model that can reproduce with high accuracy the movement of loose clothing such as kimonos, hakama, or dresses, which is difficult to reproduce with conventional techniques.
[0016] Figure 1 illustrates an example of the difference between a three-dimensional human model 11 generated by conventional technology and a three-dimensional human model 12 generated by the technology of this disclosure. Specifically, Figure 1(a) shows an example of a three-dimensional human model 10. Figure 1(b) shows an example of a three-dimensional human model 11 when animation is added to the three-dimensional human model 10 generated by conventional technology. Furthermore, Figure 1(c) shows an example of a three-dimensional human model 12 when animation is added to the three-dimensional human model 10 generated by the technology of this disclosure.
[0017] When animating a three-dimensional human model 10 wearing a cape to depict the 3D human model 10 running, it is desirable to add an animation that makes the cape appear to float, as shown in the three-dimensional human model 12 in Figure 1(c). However, if the three-dimensional human model 10 is generated using conventional technology, it is difficult to reproduce movements in the three-dimensional human model 10 that do not depend on the posture of the human body as indicated by the human body posture information. Specifically, for example, in this case, as shown in the three-dimensional human model 11 in Figure 1(b), the cape sticks unnaturally to the parts of the human body in the three-dimensional human model 11, making it difficult to achieve an expression like that of the three-dimensional human model 12 in Figure 1(c). Also, for example, if the three-dimensional human model 10 is generated using conventional technology, it is difficult to reproduce the movement of clothing, etc., that sways due to the wind in a scene where a virtual wind is blowing, using animation.
[0018] In this embodiment, in addition to the cape shown in Figure 1, a method for obtaining a three-dimensional human model corresponding to a natural person wearing loose clothing, such as a kimono, hakama, or dress, that is, clothing that does not cling to the body line, will be described. Specifically, in this embodiment, a method for obtaining a three-dimensional human model corresponding to a natural person wearing loose clothing will be described using moving images obtained by imaging with at least one imaging device. More specifically, in this embodiment, a method for obtaining a three-dimensional human model corresponding to a natural person, a physical simulation model of the clothing worn by the natural person, and external force variables representing the influence of external environmental factors such as wind that change over time on the clothing will be described using the moving images. Furthermore, in this embodiment, an embodiment will be described in which the three-dimensional human model, the physical simulation model of the clothing, and the external force variables representing the influence of the external environment are obtained by optimizing parameters included in multiple DNNs (deep neural networks).
[0019] <Configuration of the information processing device> Figure 2 is a block diagram showing an example of the hardware configuration of the information processing device 200 according to Embodiment 1. Specifically, Figure 2(a) shows an example of the hardware configuration of the information processing device 200 when the information processing device 200 is composed of a PC (personal computer), smartphone, or tablet terminal. The information processing device 200 shown in Figure 2(a) has an imaging unit 201, a CPU 202, RAM 203, ROM 204, a storage unit 205, an operation unit 206, and a display unit 207 as its hardware configuration. The imaging unit 201, CPU 202, RAM 203, ROM 204, storage unit 205, operation unit 206, and display unit 207 of the information processing device 200 are connected to each other so as to be able to communicate with each other via a system bus 208.
[0020] The imaging unit 201 includes an image sensor and a processing circuit that generates an image by controlling the image sensor. The CPU 202 is a central processing unit that executes various processes using computer programs and data stored in the RAM 203 or ROM 204. The CPU 202 controls the entire information processing device 200 by executing computer programs. The RAM 203 is a memory that temporarily stores computer programs and data loaded from the ROM 204 or storage unit 205, as well as image data of images output from the imaging unit 201. The RAM 203 functions as a work area used for executing various processes by the CPU 202. The ROM 204 is a memory that stores setting data for the information processing device 200, computer programs and data related to the startup of the information processing device 200, and computer programs and data related to the basic operation of the information processing device 200.
[0021] The memory unit 205 is a storage device composed of an HDD (hard disk drive) or the like, and stores the OS (operating system) and various computer programs and data executed by the CPU 202. The data stored in the memory unit 205 includes data related to a DNN that performs three-dimensional reconstruction processing. The computer programs and data stored in the memory unit 205 are loaded into the RAM 203 as appropriate according to the control of the CPU 202 and become subject to processing by the CPU 202. The operation unit 206 is an input device composed of a keyboard, mouse, or touch panel, and inputs various instructions to the CPU 202 based on user operations (hereinafter referred to as "user operations"). The display unit 207 is a display device composed of a liquid crystal display or touch panel, and displays the processing results of the CPU 202 as images or characters. The display unit 207 may also be a projection device such as a projector that projects images or characters.
[0022] Figure 2(b) shows another example of the hardware configuration of the information processing device 200. The information processing device 200 does not necessarily have an imaging unit 201, as shown in Figure 2(b). The communication unit 209 is a communication interface that performs data communication with external devices such as the imaging device 220. The imaging device 220 is an imaging device composed of a digital still camera or digital video camera having an imaging unit 201, and outputs imaging information such as data of the image obtained by imaging and camera parameters related to said imaging. The information processing device 200 acquires the imaging information output from the imaging device 220 via the communication unit 209.
[0023] Figure 2(c) shows another example of the hardware configuration of the information processing device 200. The database 230 is composed of storage devices that store imaging information. The information processing device 200 may also be configured to acquire imaging information stored in the database 230 via the communication unit 209, as shown in Figure 2(c). Note that the above-described configuration of the information processing device 200 is just one example and is not limited to the configuration shown in Figures 2(a), (b), or (c).
[0024] Figure 3 is a block diagram showing an example of the functional configuration of the information processing device 200 according to Embodiment 1. The information processing device 200 has an image acquisition unit 301, an image determination unit 302, a reference model acquisition unit 303, a posture estimation unit 304, and an optimization processing unit 305 as its functional configuration. The optimization processing unit 305 has a deformation unit 310 and a learning unit 314 as its internal functional configuration, and the deformation unit 310 has a posture correction unit 311, a posture deformation unit 312, and a clothing deformation unit 313 as its internal functional configuration. Each part of the functional configuration of the information processing device 200 is realized, for example, by the CPU 202 executing a computer program stored in ROM 204, etc., using RAM 203 as work memory. It should be noted that not all of the processing of the functional configuration of the information processing device 200 necessarily has to be realized by the execution of a computer program by the CPU 202, and some or all of the processing may be configured to be executed by one or more processing circuits other than the CPU 202.
[0025] <Operation of the Information Processing Device> Figure 4 is a flowchart showing an example of the processing flow of the information processing device 200 according to Embodiment 1. The series of processing steps shown in the flowchart of Figure 4 are realized, for example, by the CPU 202 reading a predetermined computer program from ROM 204 or the like, loading it into RAM 203, and then executing it. In the following, each processing step will be represented by adding "S" to the beginning of its reference numeral. The flowchart shown in Figure 4(a) shows an example of the flow of a series of processes in the information processing device 200. The flowcharts shown in Figures 4(b) to (d) will be described later.
[0026] Figure 5 is a diagram illustrating the process of generating a three-dimensional human model by the information processing device 200 according to Embodiment 1. In this embodiment, as an example, a normal space 507 in which a normal three-dimensional human model exists and an observation space 513 are defined, as shown in Figure 5. Furthermore, in this embodiment, as an example, assuming that the target of the generation of the three-dimensional human model is a natural person, a method of transforming the normal space 507 into the observation space 513 and optimizing the two spaces with respect to observation information will be described.
[0027] First, in S401, the image acquisition unit 301 acquires the captured image 501. Specifically, for example, the image acquisition unit 301 acquires a series of frames as the captured image. More specifically, the image acquisition unit 301 acquires a two-dimensional moving image (hereinafter referred to as "2D moving image") as the captured image 501 obtained by the imaging unit 201 imaging the target natural person. In this way, the image acquisition unit 301 acquires multiple frames 502 to 504 that are continuous in time series and included in the captured image 501, i.e., the 2D moving image. The captured image 501 acquired by the image acquisition unit 301 may be obtained by the imaging unit 201 imaging the target natural person, or it may be one that was captured in the past and stored in the database 230 in advance. Hereinafter, the image acquisition unit 301 will be described as acquiring N (N is an integer of 2 or more) frames 502 to 504 that are continuous in time series as the captured image 501.
[0028] Next, in S402, the image determination unit 302 determines a reference image 505 from among the multiple frames 502 to 504 acquired in S401. The reference image 505 is used to define the reference person model 508, which is the initial value of the normal three-dimensional person model in the normal space 507 described later. For this reason, it is preferable that the reference image 505 has a larger amount of observed information about the natural person in the image (frame). Specifically, for example, the image determination unit 302 first extracts an image region (hereinafter referred to as "object region") containing the natural person as an image by performing segmentation processing such as MASK-R-CNN on each frame 502 to 504. Subsequently, the image determination unit 302 determines the frame containing the object region with the largest area among the multiple object regions obtained by the extraction as the reference image 505.
[0029] Furthermore, the method for extracting or identifying the object region is not limited to methods using trained models such as MASK-R-CNN obtained as a result of machine learning; the object region may be extracted or identified by methods such as background subtraction. Also, the method for determining the reference image 505 is not limited to methods that compare the size of the object region. For example, the reference image 505 may be determined as the frame in which the three-dimensional posture closest to a predetermined posture such as T-Pose or Y-Pose is estimated, based on the three-dimensional posture information obtained by estimation based on each of the frames 502 to 504 described later.
[0030] Next, in S403, the reference model acquisition unit 303 acquires a reference person model 508 as the initial value of a normalized three-dimensional person model corresponding to the target natural person, based on the reference image 505 determined in S402. Specifically, the reference model acquisition unit 303 acquires the reference person model 508 using a pre-trained three-dimensional person model reconstruction DNN 506, which estimates a three-dimensional person model corresponding to a natural person included as an image in a single input image. More specifically, the reference model acquisition unit 303 inputs a reference image to the three-dimensional person model reconstruction DNN 506 and acquires the normalized three-dimensional person model output by the three-dimensional person model reconstruction DNN 506 as the reference person model 508. The normalized three-dimensional person model estimated by the three-dimensional person model reconstruction DNN 506 is information about the three-dimensional shape of the target natural person when they are wearing clothes. As described above, the normal space 507 is a virtual space in which the reference person model 508, obtained as a result of estimation by the three-dimensional person model reconstruction DNN 506, exists.
[0031] Figure 6 is a diagram illustrating an example of a training method for the three-dimensional human model reconstruction DNN 506 according to Embodiment 1. Specifically, Figure 6(a) shows an example of a three-dimensional human model 601 created in advance using a CG authoring tool, etc., and multiple viewpoints (hereinafter referred to as "virtual viewpoints 602") arbitrarily placed in a virtual space. For training the three-dimensional human model reconstruction DNN 506, the three-dimensional human model 601 shown as an example in Figure 6(a) and images obtained by rendering the three-dimensional human model 601 based on each virtual viewpoint 602 are used as the training dataset. It is preferable to have multiple variations of clothing worn by the three-dimensional human model 601, such as casual wear, kimono, hakama, and dresses.
[0032] FIG. 6(b) shows an example of the input and output of the three-dimensional human model reconstruction DNN 604, which is the three-dimensional human model reconstruction DNN 506 during learning. The three-dimensional human model reconstruction DNN 604 is input with an image 603 obtained by rendering based on the virtual viewpoint 602. The three-dimensional human model reconstruction DNN 604 estimates and outputs a three-dimensional human model 605 based on the input image 603. The learning of the three-dimensional human model reconstruction DNN 604 is advanced by performing error backpropagation with the difference between the three-dimensional human model 605 obtained by the estimation and the three-dimensional human model 601 which is the correct answer (hereinafter referred to as "GT (Ground Truth)") as the loss.
[0033] Here, the coordinate transformation from the world coordinate system to the camera coordinate system, which is necessary for comparing the three-dimensional human model 605 obtained by the estimation by the three-dimensional human model reconstruction DNN 604 with the three-dimensional human model 601, will be described. When comparing two three-dimensional human models expressed in different coordinate systems, it is necessary to align their coordinate systems to a common coordinate system expression. Hereinafter, in FIG. 6(a), the ID of the virtual camera corresponding to each virtual viewpoint 602 arranged (hereinafter referred to as "camera ID") will be described as p (0 ≦ p < P (P is an integer of 1 or more)). Also, hereinafter, in the virtual camera with camera ID p, the pre-calibrated intrinsic internal parameters will be denoted as K , p , , , p ,
[0033] , p , world ,
[0035] , p , , p , , , ,
[0034] , and the pose of the virtual camera will be denoted as R p and t p and described. Here, the world coordinates are X world and the coordinates of the virtual camera with camera ID p are X p Then, the conversion from the world coordinate system to the camera coordinate system can be performed using, for example, Equation (1).
[0034] [Equation] Also, the pixel coordinates x p in the image obtained by imaging with the virtual camera with camera ID p can be calculated using, for example, Equation (2).
[0035] [Number] Hereinafter, assume that the data of the three-dimensional human model 601 which is GT is represented by a three-dimensional exclusive field point group based on binary values, and the points included in the three-dimensional human model are represented as 1 and the points not included are represented as 0. That is, the three-dimensional human model 601 which is GT is represented by binary correct values, and the voxel value f of the three-dimensional human model 601 corresponding to the voxel of coordinate X v * (X) can be expressed, for example, as in Equation (3).
[0036] [Number] On the other hand, in the estimation by the three-dimensional human model reconstruction DNN 604, continuous values between 0 and 1 are estimated as the values of each voxel coordinate in the three-dimensional exclusive field point group. Hereinafter, the information of the three-dimensional exclusive field point group output by the three-dimensional human model reconstruction DNN 604 will be described by referring to it as a three-dimensional voxel map. The estimated value of the three-dimensional human model reconstruction DNN 604 corresponding to the voxel coordinate X is determined by the image feature g(I(x)) obtained by the image encoder g based on the pixel coordinate x in the input image I. The pixel coordinate x can be calculated using Equation (2). A function for calculating the voxel value at the voxel coordinate X using the image feature g(I(x)) is f v Then, the voxel value can be expressed as f v (g(l(x)), z(x)). Here, z(x) represents the depth value from the coordinates of the virtual camera to the voxel coordinate X. If the number of sampled voxels is n v pieces, the loss L v can be calculated, for example, using Equation (4).
[0037] [Number] The reference human model 508, which is the initial value of the normalized three-dimensional human model, is obtained by converting the continuous three-dimensional human model output by the three-dimensional human model reconstruction DNN 506 obtained as a result of the learning described above into binary values of 0 or 1 by thresholding. The reference human model 508 may also be obtained by converting the three-dimensional human model output by the three-dimensional human model reconstruction DNN 506 into a mesh format using the Marching Cube method or the like. As described above, by using the three-dimensional human model reconstruction DNN 506 obtained as a result of prior learning, a normalized space 507 containing the normalized three-dimensional human model is estimated.
[0038] Following S403, in S404, the attitude estimation unit 304 determines the attitude θ1 to θ of the target natural person at the time corresponding to each frame 502 to 504, based on the continuous time series frames 502 to 504 acquired in S401. N Specifically, the posture estimation unit 304 estimates the posture of a natural person in each frame 502 to 504 using a pre-trained posture estimation DNN 509 that estimates the posture of a natural person contained as an image in a single input image.
[0039] Figure 7 is a diagram illustrating an example of a training method for the pose estimation DNN 509 according to Embodiment 1. Specifically, Figure 7(a) shows an example of the input and output of the pose estimation DNN 602, which is the pose estimation DNN 509 during training. The pose estimation DNN 702 receives an image 701 obtained by rendering based on a virtual viewpoint 602 as input. Based on the input image 701, the pose estimation DNN 702 estimates and outputs the pose 703 of a natural person included as an image in the image 701. The training of the pose estimation DNN 702 is carried out by backpropagation of the error, with the difference between the information on the pose 703 of a natural person obtained by the estimation and the information on the pose of the three-dimensional human model 601 as the GT being used as the loss.
[0040] Figures 7(b) to 7(d) show examples of the representation format of information regarding the postures 711 to 713 of a natural person, output by posture estimation DNNs 509 and 702. Note that the representation format of the posture of a natural person is not limited to these examples. For example, posture estimation DNNs 509 and 702 output a three-dimensional likelihood map representing the likelihood of existence in three-dimensional space for each joint of a natural person, as shown in posture 711 in Figure 7(b). In the three-dimensional likelihood map, the three-dimensional coordinate with the highest likelihood becomes the estimated position of each joint.
[0041] This section provides a detailed explanation of the training method for the pose estimation DNN702. As mentioned above, the training of the pose estimation DNN702 uses a three-dimensional human model 601 as the GT and an image 603 obtained by rendering the three-dimensional human model, similar to the training of the three-dimensional human model reconstruction DNN604. Specifically, the training dataset for the pose estimation DNN702 uses the pose of the three-dimensional human model 601 as the GT, i.e., information (coordinates) indicating the position of each joint of the three-dimensional human model 601, and the image 603 obtained by the rendering, as the training dataset. The loss backpropagated in the training of the pose estimation DNN702 can simply be the difference between the coordinates of each joint as the GT and the coordinates of each joint obtained as the result of estimation by the pose estimation DNN702. Alternatively, the loss L can be calculated by taking advantage of the characteristics of the human pose data structured as a tree. Bone The following design may be used. Below, the loss that is backpropagated in the training of the pose estimation DNN702 is a loss L that takes advantage of the characteristics of human pose data that is structured as a tree. Bone Let's explain the case where this is the case.
[0042] Each joint is represented by a dependency relationship with the pelvis as the root. Below, as shown in posture 712 in Figure 7(c), each joint is denoted as h (h=1,...,H (where H is the total number of joints)), and the position (coordinate) of each joint is denoted as J h This is how it is expressed and explained. In posture 713 shown in Figure 7(d), H=20. As in posture 713 shown in Figure 7(d), the branch connecting joint h and parent joint parent(h) is a three-dimensional vector (hereinafter referred to as "bone vector B"). hThis is called ". ) When replaced with bone vector B h This can be expressed, for example, as in equation (5).
[0043]
number
[0044]
number
[0045] Following S404, in S405, the optimization processing unit 305 selects an arbitrary frame from among the multiple frames 502 to 504 acquired in S401. Hereafter, the frame selected in S405 will be referred to as the "selected frame". Next, in S406, the optimization processing unit 305 performs optimization processing on the three-dimensional human model corresponding to the target natural person, the physical simulation model of clothing etc. associated with the natural person, and the external force variables, based on the selected frame selected in S405. Details of the optimization processing by the optimization processing unit 305 in S406 will be described later.
[0046] Next, in S407, the optimization processing unit 305 determines whether all of the multiple frames 502 to 504 acquired in S401 were selected in S405. If it is determined in S407 that at least some of the frames have not been selected, the information processing device 200 returns to the process of S405. Subsequently, the information processing device 200 repeatedly executes the processes from S405 to S407 until it is determined in S407 that all of the frames 502 to 504 have been selected. If it is determined in S407 that all of the frames 502 to 504 have been selected, in S408, the information processing device 200 generates a three-dimensional human model based on the result of the optimization processing in S406 repeated by the optimization processing unit 305. Specifically, the three-dimensional human model is generated by the deformation unit 310 after optimization processing deforming the reference human model using the posture estimation information estimated by the posture estimation unit 304 and the reference human model after optimization processing. After S408, the information processing device 200 completes the processing shown in the flowchart in Figure 4(a).
[0047] <About the optimization process> The flowchart shown in Figure 4(b) is an example of the optimization process flow in the optimization processing unit 305, and shows an example of a detailed processing flow of the optimization process in S406. The processing in the flowchart shown in Figure 4(b) is executed after S405. After S405, first, in S411, the deformation unit 310 deforms the reference person model 508 acquired in S403 based on the posture estimated in S403 (hereinafter referred to as the "selected frame posture") that corresponds to the selected frame selected in S405. In the iterative processing of S406, the deformation unit 310 deforms the reference person model 508 after deformation in S411 based on the new selected frame posture.
[0048] The flowchart shown in Figure 4(c) is an example of the deformation process of the reference human model 508 in the deformation unit 310, and shows an example of a detailed processing flow of the deformation process of the reference human model 508 by the deformation unit 310 in S411. As described above, the posture of a natural person obtained as a result of the estimation process by the posture estimation unit 304 in S403 is the result of independent estimation for each frame acquired by the image acquisition unit 301. Therefore, depending on the input image (frame), there may be cases where the error in the estimation result is large, or where there is variation in the length of each bone, which should be common between sequences. For this reason, first, in S421, the posture correction unit 311 corrects the error in the posture estimation result by the posture estimation unit 304 in S403. Specifically, the posture correction unit 311 corrects the error in the posture estimation result using the posture correction DNN 510.
[0049] Figure 8 is a diagram illustrating an example of the deformation process of the reference human model 508 by the deformation unit 310 according to Embodiment 1. Specifically, Figure 8(a) shows an example of the input and output of the attitude correction DNN 510 in the attitude correction unit 311. Figures 8(b) and (c) will be described later. Hereafter, the architecture of the attitude correction DNN 510 will be described assuming it is an MLP (Multi-layer Perceptron), but the architecture of the attitude correction DNN 510 is not limited to MLP. The output of the attitude correction DNN 510, i.e., Δθ, which is the error correction value in the attitude estimation result, can be expressed using, for example, formula (7).
[0050]
number
[0051]
number
[0052]
number
[0053]
number
[0054]
number
[0055]
number
[0056] Figure 9 is a diagram illustrating an example of deformation processing of parts corresponding to clothing in the clothing deformation unit 313 according to Embodiment 1. Figure 9(a) shows an example of a posture-corrected human model 901 obtained as a result of skinning processing by the posture deformation unit 312. Figure 9(b) shows an example of a clothing-deformed human model 902 obtained as a result of deformation processing of parts corresponding to clothing by the clothing deformation unit 313. As shown in Figure 9(a), loose clothing movements that cannot be reproduced by skinning processing using LBS, such as the movement of the sleeves or hem of a kimono when the wind blows, are represented by physical simulation by the clothing deformation unit 313, as shown in Figure 9(b).
[0057] Figure 9(c) shows point x in the normal space 507. c From point x in observation space 513 i This shows a series of deformation processes for the parts corresponding to clothing on the figure. Note that in Figure 9(c), for explanatory purposes, only the region corresponding to clothing is defined as normal space 507, but it is sufficient if the region corresponding to clothing in normal space 507 can be identified. The region corresponding to clothing in normal space 507 can be identified by the following method. For example, first, based on the estimated skinning weight, the influence of the joints of the human body w h (x cThe region corresponding to the human body in normal space 507 is identified by fitting a parametric model using the values of ). Subsequently, the region corresponding to clothing in normal space 507 is identified by dividing the region of normal space 507 using three-dimensional semantic segmentation, based on the difference between the normal three-dimensional human model and the region corresponding to the human body in normal space 507 identified. Note that the method for identifying the region corresponding to clothing in normal space 507 is not limited to the method described above.
[0058] Furthermore, in this embodiment, the region corresponding to the clothing in the normal space 507 is described as being identified as a preprocessing step for the optimization process, but the identification of the region corresponding to the clothing may also be performed during the optimization process. Below, the factors that cause deformation of clothing will be described as the gravitational force acting on the clothing due to the movement of a natural person wearing the clothing, the elastic force due to the stretching and contracting of the clothing, and external forces acting due to the external environment such as wind. The gravitational force acting on the clothing due to the movement of a natural person can be defined by Newton's second law. The gravitational force acting on the clothing in the i-th selected frame is F traction_i Therefore, F traction_i This can be expressed, for example, using formula (13).
[0059]
number
[0060]
number
[0061] Based on the above, the clothing simulation DNN512 will now be described. Figure 10 is a block diagram showing an example of the internal configuration of the clothing simulation DNN512 according to Embodiment 1. As shown in Figure 10, point x in the normal space 507 c From point x in observation space 513 i The deformation of the region corresponding to the clothing in ' is performed by executing a physical simulation process, including processing by the clothing simulation DNN512. Also, point x in normal space 507 c From point x in observation space 513 i The deformation of the normalized 3D human model is performed by executing a skinning process using LBS, which includes processing with the skinning DNN511.
[0062] The garment simulation DNN512 includes garment-specific MLP1001, mass MLP1002, and elastic MLP1003. The configuration of the garment simulation DNN512 and the architecture of each component included in the garment simulation DNN512 are not limited thereto.
[0063] The garment-specific MLP1001 is located at point x in the normal space 507. c The latent variable z represents the material properties of the clothing in the context of cloth The output is: Mass MLP1002 is the latent variable z output from garment-specific MLP1001. cloth Using as input, point x c It outputs the mass m in [location]. Additionally, the elastic MLP1003 outputs the latent variable z from the garment-specific MLP1001. cloth Using as input, point x cOutputs the elastic modulus k in the following: The functions corresponding to the garment-specific MLP1001, mass MLP1002, and elastic MLP1003 are output in order. Identity MLP m MLP k Therefore, the values output from each MLP can be expressed, for example, using formulas (15) to (17).
[0064]
number
[0065]
number
[0066] The flowchart shown in Figure 4(d) is an example of the learning process flow in the learning unit 314, specifically an example of the detailed processing flow in S412 shown in Figure 4(c). The processing in the flowchart shown in Figure 4(d) is executed after the processing in S411 shown in Figure 4(c). First, in S431, the learning unit 314 generates a rendered image 514 by rendering the three-dimensional human model (deformed human model) existing in the observation space 513 obtained as a result of the deformation process in S411 using the imaging camera parameters.
[0067] Figure 9(e) shows the deformation process in S411. As shown in Figure 9(e), each sample point on the camera ray r in the observation space 513 is warped from the point corresponding to that sample point in the normal space 507 by the deformation process in S411. The RGB values C(r) of the pixels corresponding to the camera ray r in the rendered image 514 can be calculated, for example, using formulas (19) and (20). In formulas (19) and (20), D (where D is an integer greater than or equal to 2) is the number of sample points on the camera ray, d (where d is an integer less than or equal to D) is the ID of each sample point, and Δs i This represents the sampling width, which corresponds to the distance between sampling points.
[0068]
number
[0069]
number
[0070] After S432, in S433, the learning unit 314 performs backpropagation of the loss LTotal to optimize the CNN, which is the target of optimization. Bone MLP Identity MLP m MLP k , external force variable F external_i , and latent variable z skinning ,z cloth The learning unit 314 performs the learning process. Specifically, in the learning process, the learning unit 314 updates the parameters of the posture correction DNN 510, the skinning DNN 511, and the clothing simulation DNN 512, as well as variables such as the external force variables included in the clothing simulation DNN 512. Next, in S434, the learning unit 314 performs the learning of the normal space 507, that is, updates the reference person model 508 that exists in the normal space 507. After S434, the learning unit 314 completes the process of the flowchart shown in Figure 4(d), that is, the process of S412 shown in Figure 4(c). After the process of S412, the information processing device 200 completes the process of the flowchart shown in Figure 4(b), that is, the process of S406 shown in Figure 4(a).
[0071] In addition, while the above describes an embodiment for optimizing the parameters of the clothing simulation DNN, physical simulation is not limited to simulations related to clothing. For example, the physical simulation according to this embodiment can also be applied to simulations of objects associated with natural people that are different from clothing, such as hair. Below, an embodiment for optimizing the physical simulation DNN related to hair will be described. First, a region corresponding to the hair of a natural person in the posture deformation person model (hereinafter referred to as the "hair-corresponding region") is determined using semantic segmentation or the like. Next, in the hair-corresponding region, parameters for physical simulation corresponding to mass and elastic modulus, etc., are defined, similar to the mass and elastic modulus, etc., described above using clothing as an example, and these parameters are optimized. With this configuration, a three-dimensional person model can be generated that corresponds to the case where the clothes and hair of performer 1302 flutter when wind blows on performer 1302, as shown in Figure 13, which will be described later.
[0072] Furthermore, although this embodiment describes the case where a natural person is imaged using a single imaging device, the imaging device is not limited to one; multiple devices may be used. When a natural person is imaged using multiple imaging devices, the optimization of the parameters and external force variables of each DNN is performed using images obtained from imaging by multiple imaging devices capturing the same scene. Therefore, in this case, a more accurate three-dimensional human model can be generated.
[0073] Furthermore, although the above explanation assumed that the target of the three-dimensional model generation was a natural human being, the target of the three-dimensional model generation is not limited to natural humans, but may also be animals other than natural humans, such as dogs. Figure 11 shows an example of posture information for an animal other than a natural human being, such as a dog. The posture information of the animal shown as an example in Figure 11 can be defined by an animal parametric model such as the SMAL model disclosed in Non-Patent Document 3 below, or by a technique called MMpose, etc.
[0074] <Non-Patent Document 3> Silvia Zuffi, et al., “3D Menagerie: Modeling the 3D Shape and Pose of Animals”, 2017, IEEE Conference on Computer Vision and Pattern Recognition (CVPR), [online], [Retrieved September 12, 2024], Internet <https: / / openaccess.thecvf.com / content_cvpr_2017 / papers / Zuffi_3D_Menagerie_Modeling_CVPR_2017_paper.pdf> Furthermore, although this embodiment describes the case where there is one natural person, the technology of this disclosure is also applicable when there are multiple natural people simultaneously. In this case, first, an object ID corresponding to each of the multiple objects is assigned to the region corresponding to the multiple objects in the video image, using a multi-object tracking task such as the one disclosed in Non-Patent Document 4 shown below. Subsequently, a normal space and an observation space are assigned to each object ID, and the process described above in this embodiment is executed.
[0075] <Non-patent document 4> Paul Voigtlaender, et al., “MOTS: Multi-Object Tracking and Segmentation”, 2019, In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition (CVPR), [online], [Retrieved September 12, 2024], Internet <https: / / openaccess.thecvf.com / content_CVPR_2019 / papers / Voigtlaender_MOTS_Multi-Object_Tracking_and_Segmentation_CVPR_2019_paper.pdf> Furthermore, in this embodiment, the normal space 507 is defined as a virtual space implicitly defined by the MLP, but is not limited to this. For example, a three-dimensional human model may be extracted from the implicitly defined virtual space by the method disclosed in Non-Patent Document 5 shown below. Alternatively, the normal space 507 may be represented by the method disclosed in Non-Patent Document 6 or Non-Patent Document 7 shown below.
[0076] <Non-Patent Document 5> Michael Oechsle, et al., “UNISURF: Unifying Neural Implicit Surfaces and Radiance Fields for Multi-View Reconstruction”, 2021, In International Conference on Computer Vision (ICCV), [online], [Retrieved September 12, 2024], Internet <https: / / openaccess.thecvf.com / content / ICCV2021 / papers / Oechsle_UNISURF_Unifying_Neural_Implicit_Surfaces_and_Radiance_Fields_for_Multi-View_ICCV_2021_paper.pdf> <Non-patent document 6> Anpei Chen, et al., “TensoRF: Tensorial Radiance Fields”, 2022, In European conference on computer vision (ECCV), [online], [Retrieved September 12, 2024], Internet<https: / / www.ecva.net / papers / eccv_2022 / papers_ECCV / papers / 136920332.pdf> <Non-Patent Document 7> Thomas Mu ¨ller, 4 outside, "Instant neural graphics primitives with a multiresolution hash encoding", 2022, ACM Trans. Graph. 41(4), 102:1-102:15 (Jul 2022), [online], [searched on September 12, 2024], Internet <https: / / dl.acm.org / doi / pdf / 10.1145 / 3528223.<3530127> According to the information processing apparatus 200 configured as described above, it is possible to generate a three-dimensional object in which the movement of an object associated with an animal such as clothes worn by a natural person is reproduced with high accuracy.
[0077] [Embodiment 2] In Embodiment 1, an aspect of accurately reproducing the movement of clothes accompanying the movement of a natural person and the movement of clothes independent of the movement of a natural person in a three-dimensional human model was described. Also, in Embodiment 1, in addition to the reproduction, an aspect of optimizing parameters related to the physical simulation of clothes worn by a natural person and the hair of a natural person (hereinafter referred to as "physical simulation parameters") and variables related to external forces applied to clothes and the like was described. In Embodiment 2, a method for facilitating re-editing by the user for the three-dimensional human model generated using the physical simulation parameters of clothes and the like and variables related to external forces, which were optimized and obtained by the method described above in Embodiment 1, will be described.
[0078] Figure 12 shows examples of editing screens 1200 and 1210 for a three-dimensional human model according to Embodiment 2. Editing screen 1200 or 1210 includes images 1201 and 1211 of a three-dimensional human model wearing a cape. Here, images 1201 and 1211 are images corresponding to the three-dimensional human model (hereinafter referred to as the "generated three-dimensional human model") generated by the optimization process in the information processing device 200 according to Embodiment 1. The information processing device 200 according to Embodiment 1 outputs the generated three-dimensional human model, which includes physical simulation parameters related to clothing worn by a natural person, and variables related to external forces applied to the clothing (external force variables). Therefore, using this information output from the information processing device 200, the generated three-dimensional human model can also reproduce the movement of clothing due to the influence of wind, etc., which is independent of the movement of a natural person.
[0079] The following describes an embodiment in which the data of the generated three-dimensional human model is imported into a three-dimensional authoring tool that operates on a user terminal such as a personal computer, and operations such as editing the animation of the generated three-dimensional human model are performed. As described above, the information processing device 200 according to Embodiment 1 outputs a generated three-dimensional human model including physical simulation parameters and external force variables that indicate the degree to which wind, etc., affects clothing, etc. Therefore, by using this information output by the information processing device 200, for example, if wind is generated in a virtual space, it becomes possible to generate an animation in which a cape flutters due to the effect of the wind, as shown in image 1211 in Figure 12(b).
[0080] The production of games, advertising videos, and other video content sometimes involves using three-dimensional human models that correspond to real people. Previously, animations that corresponded to the movement of clothing such as cloaks or skirts, or hair, which moved significantly due to factors other than body movement, required manually setting animation parameters. In contrast, by using the generated three-dimensional human model and external force variables output by the information processing device 200, users can easily add or edit animations to the three-dimensional human model.
[0081] In this embodiment, as an example, the production of video content in a shooting studio used for producing video content for advertising, etc., will be described. Figure 13 is a diagram showing an example of the video content production process in the shooting studio 1300 according to Embodiment 2. Figures 13(a) and (b) show an example of the imaging process in the shooting studio 1300, and Figure 13(c) shows an example of the video content production work in a location inside the shooting studio 1300 or a location away from the shooting studio 1300.
[0082] In the shooting studio 1300, the cameraman 1301 uses an imaging device 220 fixed to a tripod or the like to image the performer 1302, who is the subject of the image. In this imaging, various scenes are captured, such as having the performer 1302 strike multiple poses and move, and using a blower 1303 to blow wind 1304 onto the performer 1302 to create a sense of dynamism. For example, while the performer 1302 is being imaged, the information processing device 200 according to Embodiment 1 generates a three-dimensional human model in parallel. The worker 1322, who is producing the video content, uses a user terminal 1321 for video content production to edit the animation of the three-dimensional human model (generated three-dimensional human model) generated by the information processing device 200, in parallel with the imaging of the performer 1302. The display device of the user terminal 1321 displays the editing screens 1200 and 1210 shown in Figure 12. Furthermore, since the hardware configuration of the user terminal 1321 is the same as that of the information processing device 200 shown in Figure 2, a detailed explanation will be omitted.
[0083] Operator 1322 performs video content creation work on a three-dimensional human model generated based on captured images obtained by imaging performer 1302, via editing screens 1200 and 1210 displayed on user terminal 1321. Video content creation work includes, for example, correcting the movement of natural people, clothing, and hair, generating images corresponding to views from arbitrary virtual viewpoints (virtual viewpoint images), and compositing backgrounds or effects using CG. By performing such production work, video content that cannot be obtained from captured images alone can be created. In this shooting studio 1300, the use of artificial wind to create scenes allows for the capture of the movement of performer 1302's clothing and hair. Therefore, the information processing device 200 according to Embodiment 1 can acquire observational information regarding the movement of clothing, etc., due to the influence of wind, which is effective for generating a three-dimensional human model.
[0084] Figures 14 and 15 are sequence diagrams showing an example of the processing flow in the production of video content according to Embodiment 2. Specifically, the sequence diagram shown in Figure 14 shows an example of the processing flow when, in parallel with imaging by the imaging device 220, a three-dimensional human model is generated by the information processing device 200, and the generated three-dimensional human model is re-edited by the user terminal 1321. The sequence diagram shown in Figure 15 shows an example of the processing flow when the information processing device 200 generates a three-dimensional human model using imaging images stored in the database 230, etc., and in parallel with the re-editing of the generated three-dimensional human model by the user terminal 1321. The series of processing steps shown in the sequence diagrams of Figures 14 and 15 are realized, for example, by the CPU (Central Processing Unit) of each device loading a predetermined computer program into memory and executing it.
[0085] First, in S1401, the imaging device 220 performs imaging processing based on the operation of the cameraman 1301. Next, in S1402, the imaging device 220 transmits the image (frame) generated by the imaging processing to the information processing device 200, which receives it. The processes in S1401 and S1402 are repeatedly executed, for example, until the cameraman 1301 performs an operation to terminate the imaging processing. Next, in S1403, the information processing device 200 transmits the received image (frame) to the user terminal 1321, which receives it. The process in S1403 is repeatedly executed, for example, each time the information processing device 200 receives a frame in the video. Next, in S1404, the user terminal 1321 displays the received image (frame) as a live view on the display device of the user terminal 1321. The process in S1404 is repeatedly executed, for example, each time the user terminal 1321 receives a frame in the video.
[0086] Next, in S1405, the information processing device 200 generates a three-dimensional human model based on the received captured image (frame). The process in S1405 is the same as the series of processes in the flowchart shown in Figure 4, so its explanation is omitted. That is, in the process of S1405, the information processing device 200 executes the series of processes in the flowchart shown in Figure 4. The process in S1405 is executed repeatedly, for example, each time the information processing device 200 receives a frame of a moving image from the imaging device 220. The process in S1405 may also be executed repeatedly, for example, each time the information processing device 200 receives a predetermined number of frames.
[0087] Each time the information processing device 200 receives a frame, the number of frames stored in the information processing device 200 increases. As the number of stored frames increases, the information processing device 200 increases the number of frames used in the optimization process S406 executed in the S1405 process. The performer 1302, which is the object being imaged, remains the same. Therefore, the observation information increases step by step as frames are stored. The information processing device 200 can perform efficient fine tuning by using the result of the previous optimization process as an initial value and performing optimization processing on the new observation information. In addition, the optimization process S406 executed in the S1405 process may use all stored frames, only newly received frames, or multiple frames newly stored within a predetermined period. Furthermore, the frames used in the optimization process may be updated at predetermined timings, such as when setting the loss threshold in learning.
[0088] After S1405, in S1406, the information processing device 200 sends the three-dimensional human model (generated three-dimensional human model) and external force variables generated in S1405 to the user terminal 1321, which receives them. As mentioned above, the generated three-dimensional human model includes physical simulation parameters related to clothing worn by a natural person. Next, in S1407, the user terminal 1321 imports the received generated three-dimensional human model and external force variables into the three-dimensional authoring tool. The process in S1407 is repeated, for example, each time the user terminal 1321 receives a new generated three-dimensional human model. In this case, the user terminal 1321 may import only the difference between the generated three-dimensional human model already imported and the newly received generated three-dimensional human model into the three-dimensional authoring tool.
[0089] Next, in S1408, the user terminal 1321 displays the generated 3D human model imported into the 3D authoring tool on the display device of the user terminal 1321. Specifically, the user terminal 1321 displays an editing screen including the image of the generated 3D human model on the display device. Next, in S1409, the user terminal 1321 obtains information regarding editing operations such as animation performed by the worker 1322 in each frame of the video content to be generated. Next, in S1410, the user terminal 1321 performs editing processing such as adding or modifying animation to the frames of the video content based on the editing operation information obtained in S1409.
[0090] As is clear from the above explanation, the processes from S1401 to S1410 are repeatedly executed while the imaging process in the imaging device 220 continues, that is, while the information processing device 200 continues to receive frames. Note that the series of processes in S1409 and S1410 do not depend on the processes in S1407 and S1408. Therefore, the processes in S1409 and S1410 do not have to be executed at all during the period from when a new generated 3D human model is imported in S1407 until when the next generated 3D human model is imported in S1407, or they may be executed multiple times.
[0091] Figure 16 shows an example of an editing screen 1600 according to Embodiment 2. The editing screen 1600 has an image display area 1610, a spatial display area 1602, and a settings area 1603. The image display area 1610 displays, for example, frames 1611 that are continuous in the time series received by the user terminal 1321 up to the time the editing screen 1600 is displayed. The display content of the image display area 1610 is updated, for example, each time the user terminal 1321 receives a new frame. The spatial display area 1602 displays an image of the generated three-dimensional human model that has been imported into the three-dimensional authoring tool at the time the editing screen 1600 is displayed. The display content of the spatial display area 1602 is updated, for example, each time a new generated three-dimensional human model is imported into the three-dimensional authoring tool.
[0092] The settings area 1603 displays a GUI (Graphical User Interface) for accepting user input. The settings area 1603 contains GUI components such as multiple buttons or input boxes, and the operator 1322 can perform various settings by pressing the buttons or by entering numbers or characters into the input boxes. These settings include settings necessary for generating video content, such as setting the time or frame, settings to adjust the shape of the generated 3D human model, settings to set the position and direction of the viewpoint used for rendering the generated 3D human model, and settings for the background and effects. The coordinate axes 1605 displayed in the image display area 1610 represent the world coordinate system of the virtual space displayed in the spatial display area 1602.
[0093] Figure 17 shows an example of an editing screen according to Embodiment 2 when various operations are received from the operator 1322 in the setting area 1603. Specifically, Figure 17(a) shows an example of an editing screen 1700 when the user terminal 1321 receives a user operation on GUI component A, which is located in the setting area 1603 and is used to select a time or frame. When GUI component A is selected and frame 1611 is selected, the user terminal 1321 displays an image in the spatial display area 1602 that includes an image of a three-dimensional human model representing the time corresponding to frame 1611.
[0094] Figure 17(b) shows an example of the editing screen 1720 when the user terminal 1321 receives a user operation on GUI component B, which is located in the setting area 1603 and is used to adjust the shape of a three-dimensional human model. When GUI component B is selected and an editing operation is performed on the image of the three-dimensional human model displayed in the spatial display area 1602, the user terminal 1321 changes the shape of the three-dimensional human model by updating the bones of the three-dimensional human model based on the editing operation. For example, if an operation is performed to change the position of the parts corresponding to both arms of the image of the three-dimensional human model displayed in the spatial display area 1602 upwards, the user terminal 1321 updates the three-dimensional human model so that the positions of the parts corresponding to both arms of the three-dimensional human model are upwards.
[0095] Furthermore, since the 3D human model has physical simulation parameters for clothing and other elements, the parts corresponding to clothing are updated with high accuracy as the shape of the 3D human model is updated. With such a user terminal 1321, it is possible to edit the animation of the 3D human model with high accuracy using information about the pose of the 3D human model and physical simulation parameters for clothing and other elements. For example, it is possible to make highly accurate edits such as fine-tuning the movement of parts of the 3D human model that correspond to clothing or parts that correspond to the human body.
[0096] Figure 17(c) shows an example of the editing screen 1740 when the user terminal 1321 receives a user operation on GUI component C for setting a virtual viewpoint for rendering a three-dimensional human model. When GUI component C is selected, and the external and internal parameters of the virtual camera corresponding to the virtual viewpoint for rendering the three-dimensional human model are set, the user terminal 1321 renders the three-dimensional human model based on the setting operation. The image generated by this rendering is displayed in the spatial display area 1602 as an image of the three-dimensional human model. In this way, by changing the position of the virtual viewpoint and the direction of the line of sight at that virtual viewpoint, the operator 1322 can check how the performer 1302 looks from viewpoints that are not actually captured, thereby expanding the range of video expression in the generation of video content.
[0097] Figure 17(d) shows an example of the editing screen 1760 when the user terminal 1321 receives a user operation on GUI component D, which is located in the setting area 1603 and is used to set the background and effects. When GUI component D is selected, and an editing operation is performed to add a pre-prepared 3D CG background or effect to the 3D authoring tool, the user terminal 1321 composites the background or effect into the virtual space where the 3D human model exists. The editing functions for the 3D human model or the virtual space containing the 3D human model in the 3D authoring tool are not limited to those described above. For example, the 3D authoring tool may have a function to generate simulated wind in the virtual space where the 3D human model exists, and generate an animation in which the clothing changes due to the effect of the wind. This function is realized using physical simulation parameters of the clothing and external force variables indicating the degree of influence of wind, etc., on the clothing, which are obtained by the optimization process of the information processing device 200.
[0098] After imaging by the imaging device 220 is completed based on the operation by the cameraman 1301, the user terminal 1321 executes processes S1411 and S1412 until information regarding the completion of video content editing is obtained in S1421. Here, the process of S1411 is the same as the process of S1409, and the process of S1412 is the same as the process of S1410. Specifically, for example, processes S1411 and S1412 correspond to the finishing work related to the production of video content that is performed after imaging by the imaging device 220 is completed. After information regarding the completion of video content editing is obtained in S1421, the user terminal 1321 generates the video content in S1422. Specifically, when the operator 1322 presses the "Export" button located in the setting area 1603, the user terminal 1321 obtains information regarding the completion of video content editing in S1421. Subsequently, in S1422, the user terminal 1321 performs the process of generating video content, including the results of animation editing for the frames up to that point.
[0099] Furthermore, the editing screens shown in Figures 16 and 17 are examples only, and the editing screens in the 3D authoring tool are not limited to these. Also, although the information processing device 200 and the user terminal 1321 were described as being implemented as separate devices in this embodiment, the information processing device 200 and the user terminal 1321 may be implemented as a single device. Also, although the information processing device 200 and the imaging device 220 were described as being implemented as separate devices in this embodiment, the information processing device 200 and the imaging device 220 may be implemented as a single device.
[0100] Referring to Figure 15, the processing flow when the information processing device 200 generates a three-dimensional human model using captured images stored in the database 230, etc., and simultaneously performs re-editing of the generated three-dimensional human model by the user terminal 1321 will be described. In this explanation of the processing flow, processing steps that perform the same processing as those shown in the sequence diagram of Figure 14 will be denoted by the same reference numerals as in Figure 14, and their explanation will be omitted. First, in S1501, the information processing device 200 acquires time-series captured images (frames) stored in the database 230 based on instructions from the user of the information processing device 200. These captured images are the same as the captured images transmitted from the imaging device 220 in S1402 shown in Figure 14.
[0101] After S1501, in S1503, the information processing device 200 transmits the captured image received in S1501 to the user terminal 1321, which then receives it. After S1503, the user terminal 1321 executes the process in S1404. Also after S1503, in S1505, the information processing device 200 generates a three-dimensional human model based on the captured image received in S1501. The process in S1505 is the same as the series of processes from S402 to S408 in the flowchart shown in Figure 4, so its explanation is omitted. That is, in the process of S1505, the information processing device 200 executes the series of processes from S402 to S408 in the flowchart shown in Figure 4. After S1505, the information processing device 200 executes the process in S1406. After S1406, the user terminal 1321 executes the processes from S1407 to S1422.
[0102] With the user terminal 1321 configured as described above, video content can be produced using a three-dimensional human model generated by the information processing device 200 based on time-series captured images obtained by imaging the performer 1302, who is the subject of imaging in the shooting studio 1300, etc. Conventional methods for generating three-dimensional human models have difficulty reproducing the movement of clothing and other elements that do not depend on the movement of a natural person, resulting in unnatural movement of parts of the three-dimensional human model corresponding to clothing and other elements. However, with the user terminal 1321 according to this embodiment, by using the three-dimensional human model generated by the information processing device 200 according to Embodiment 1, it is possible to generate video content that reproduces the movement of clothing and other elements with high accuracy. Furthermore, with the user terminal 1321 according to this embodiment, by using the three-dimensional human model generated by the information processing device 200 based on captured images obtained by imaging a real natural person, the following animation editing can be performed.
[0103] For example, if the information processing device 200 and the user terminal 1321 have ample computing power per unit time, the information processing device 200 can generate a three-dimensional human model of a natural person being captured in real time, as shown in Figure 13, in accordance with the frame rate of the capture. In this case, the user terminal 1321 can import the three-dimensional human model generated by the information processing device 200 into a three-dimensional authoring tool in accordance with the frame rate of the generation. Furthermore, the user terminal 1321 can perform real-time animation editing of parts of the three-dimensional human model corresponding to the human body and clothing, as well as rendering from a virtual viewpoint, to generate video content.
[0104] Furthermore, for example, the three-dimensional human model generated by the information processing device 200 includes not only information about the shape corresponding to the natural person being imaged, but also physical simulation parameters for clothing and other objects. Therefore, the user terminal 1321 can easily edit the movement of parts corresponding to objects attached to the natural person, such as clothing, and add or modify animations related to those parts, by using the three-dimensional human model generated by the information processing device 200.
[0105] Furthermore, for example, the information processing device 200 can archive and store three-dimensional human models it has previously generated in a database 230, and the user terminal 1321 can generate video content using these three-dimensional human models. For example, the user terminal 1321 places a first three-dimensional human model generated by the information processing device 200 in real time, for example, based on captured images obtained from imaging of the performer 1302, and a second three-dimensional human model generated by the information processing device 200 in the past, in the same virtual space. Figure 18 shows an example of a first three-dimensional human model 1801 and a second three-dimensional human model 1802 placed in the same virtual space according to Embodiment 2. The user terminal 1321 adds or modifies animations to the human body and clothing of the first three-dimensional human model 1801 and the second three-dimensional human model 1802, respectively. With the user terminal 1321 configured in this way, it becomes possible to generate video content using multiple three-dimensional human models generated based on captured images taken at different time periods.
[0106] [Other embodiments] The technology disclosed herein can also be implemented by supplying a program that implements one or more of the functions of the embodiments described above to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be implemented by a circuit such as an ASIC (Application Specific Integrated Circuit) that implements one or more functions.
[0107] Within the scope of this disclosure, it is possible to freely combine the embodiments, modify any component of each embodiment, or omit any component in each embodiment.
[0108] [Technical Features of This Disclosure] This disclosure includes the following configurations, methods, and programs.
[0109] <Configuration 1> Image acquisition means for acquiring a time-series image obtained by imaging an object including a first object and a second object associated with the first object, An optimization means that optimizes a three-dimensional model corresponding to the object, parameters related to the physical simulation of the second object, and variables related to external forces that may affect the second object, based on the aforementioned time-series images. A model generation means for generating the three-dimensional model based on the optimization results obtained by the above optimization, An information processing device characterized by having the following features.
[0110] <Configuration 2> The three-dimensional model includes shape information relating to the object and orientation information relating to the first object. An information processing device as described in Configuration 1, characterized by the above.
[0111] <Structure 3> The second object is at least one of the clothing worn by the first object and the hair of the first object. An information processing device according to configuration 1 or 2, characterized by the above.
[0112] <Structure 4> The parameters relating to the physical simulation include at least one of the following: information relating to the mass of the second object and information relating to the elastic modulus of the second object. An information processing device according to any one of configurations 1 to 3 characterized by the above.
[0113] <Composition 5> The optimization means performs the optimization of the parameters relating to the physical simulation based on at least one of the changes in attitude information relating to the first object and information on external forces that may affect the shape of the second object. An information processing device according to any one of configurations 1 to 4 characterized by the above.
[0114] <Composition 6> The optimization means performs the optimization of the parameters relating to the physical simulation based on latent variables representing the characteristics of the second object. An information processing device according to any one of configurations 1 to 5 characterized by the above.
[0115] <Composition 7> Image determination means for determining a reference image from the aforementioned time-series images, A model acquisition means that acquires a reference three-dimensional model, which is the reference three-dimensional model, based on the aforementioned reference image, It further possesses, The optimization means performs the optimization with respect to the reference three-dimensional model. An information processing device according to any one of configurations 1 to 6 characterized by the above.
[0116] <Structure 8> The image determination means determines the reference image based on the size of the region containing the image of the object in each of the time-series images. An information processing device according to configuration 7, characterized by the above.
[0117] <Composition 9> The image determination means determines the reference image based on the pose information relating to the first object estimated based on each of the time-series images. An information processing apparatus according to configuration 7 or 8, characterized by the above.
[0118] <Composition 10> A display means for displaying the three-dimensional model generated by the model generation means, An operation acquisition means for acquiring operation information related to the user's animation editing operations, Content generation means for generating video content based on the aforementioned operation information, Having further, An information processing device according to any one of configurations 1 to 9 characterized by the above.
[0119] <Composition 11> The operation information includes selection operation information relating to the operation in which the user selects any image from the time-series images, The content generation means generates the video content by changing the state of the three-dimensional model to match the state of the image selected by the user based on the selection operation information. An information processing device according to configuration 10, characterized by the above.
[0120] <Composition 12> The operation information includes editing operation information relating to an operation in which the user edits an animation relating to a part of the three-dimensional model that corresponds to the second object, The content generation means generates the video content by adding animation to the part of the three-dimensional model corresponding to the second object so as to reproduce the movement of the second object based on the editing operation information. An information processing apparatus according to configuration 10 or 11, characterized by the above.
[0121] <Composition 13> The operation information includes viewpoint operation information relating to the operation of setting a viewpoint when the user renders the three-dimensional space in which the three-dimensional model exists. The content generation means generates the video content by rendering the three-dimensional space based on the viewpoint manipulation information. An information processing device according to any one of configurations 10 to 12 characterized by the above.
[0122] <Method> An image acquisition step of acquiring a time-series image obtained by imaging an object including a first object and a second object associated with the first object, An optimization step is performed to optimize the three-dimensional model corresponding to the object, the parameters for the physical simulation of the second object, and the variables for external forces that may affect the second object, based on the aforementioned time-series images. A model generation step is performed to generate the three-dimensional model based on the optimization results obtained by the above optimization, An information processing method characterized by including
[0123] <Program> A program to cause a computer to function as an information processing device described in any one of configurations 1 to 13. [Explanation of Symbols]
[0124] 200 Information Processing Devices 305 Optimization Processing Unit 310 Deformed part
Claims
1. Image acquisition means for acquiring a time-series image obtained by imaging an object including a first object and a second object associated with the first object, An optimization means that optimizes a three-dimensional model corresponding to the object, parameters related to the physical simulation of the second object, and variables related to external forces that may affect the second object, based on the aforementioned time-series images. A model generation means for generating the three-dimensional model based on the optimization results obtained by the above optimization, An information processing device characterized by having the following features.
2. The three-dimensional model includes shape information relating to the object and orientation information relating to the first object. The information processing apparatus according to claim 1, characterized by the following:
3. The second object is at least one of the clothing worn by the first object and the hair of the first object. The information processing apparatus according to claim 1, characterized by the following:
4. The parameters relating to the physical simulation include at least one of the following: information relating to the mass of the second object and information relating to the elastic modulus of the second object. The information processing apparatus according to claim 1, characterized by the following:
5. The optimization means performs the optimization of the parameters relating to the physical simulation based on at least one of the changes in attitude information relating to the first object and information on external forces that may affect the shape of the second object. The information processing apparatus according to claim 1, characterized by the following:
6. The optimization means performs the optimization of the parameters relating to the physical simulation based on latent variables representing the characteristics of the second object. The information processing apparatus according to claim 1, characterized by the following:
7. Image determination means for determining a reference image from the aforementioned time-series images, A model acquisition means that acquires a reference three-dimensional model, which is the reference three-dimensional model, based on the aforementioned reference image, It further possesses, The optimization means performs the optimization with respect to the reference three-dimensional model. The information processing apparatus according to claim 1, characterized by the following:
8. The image determination means determines the reference image based on the size of the region containing the image of the object in each of the time-series images. The information processing apparatus according to claim 7, characterized by the following:
9. The image determination means determines the reference image based on the pose information relating to the first object estimated based on each of the time-series images. The information processing apparatus according to claim 7, characterized by the following:
10. A display means for displaying the three-dimensional model generated by the model generation means, An operation acquisition means for acquiring operation information related to the user's animation editing operations, Content generation means for generating video content based on the aforementioned operation information, Having further, The information processing apparatus according to claim 1, characterized by the following:
11. The operation information includes selection operation information relating to the operation in which the user selects any image from the time-series images, The content generation means generates the video content by changing the state of the three-dimensional model to match the state of the image selected by the user based on the selection operation information. The information processing apparatus according to claim 10, characterized by the above.
12. The operation information includes editing operation information relating to an operation in which the user edits an animation relating to a part of the three-dimensional model that corresponds to the second object, The content generation means generates the video content by adding animation to the part of the three-dimensional model corresponding to the second object so as to reproduce the movement of the second object based on the editing operation information. The information processing apparatus according to claim 10, characterized by the above.
13. The operation information includes viewpoint operation information relating to the operation of setting a viewpoint when the user renders the three-dimensional space in which the three-dimensional model exists. The content generation means generates the video content by rendering the three-dimensional space based on the viewpoint manipulation information. The information processing apparatus according to claim 10, characterized by the above.
14. An image acquisition step of acquiring a time-series image obtained by imaging an object including a first object and a second object associated with the first object, An optimization step is performed to optimize the three-dimensional model corresponding to the object, the parameters for the physical simulation of the second object, and the variables for external forces that may affect the second object, based on the aforementioned time-series images. A model generation step is performed to generate the three-dimensional model based on the optimization results obtained by the above optimization, An information processing method characterized by including
15. A program for causing a computer to function as an information processing device according to any one of claims 1 to 13.