Information processing device, information processing method, and program
The information processing device addresses the limitation of existing methods by determining object priorities and generating 3D models from multiple images, allowing for the creation of three-dimensional objects from any scene in a moving image.
Patent Information
- Application Number
- JP2025088199
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-08-13
- Estimated Expiration
- 2041-02-26
AI Technical Summary
Existing methods for creating three-dimensional objects from moving images are limited to selecting objects from highlight scenes, making it difficult to model objects not included in the list.
An information processing device determines the priority of objects, acquires 3D models from multiple captured images, and generates modeling data including the 3D models based on these models and priorities.
Enables the easy creation of three-dimensional objects from any scene of a moving image, overcoming the limitations of existing technologies.
Smart Images

Figure 2025119041000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a technology for generating modeling data for an object from a moving image. [Background technology]
[0002] In recent years, it has become possible to use modeling devices such as 3D printers to create figures of objects based on three-dimensional models (hereinafter referred to as 3D models), which are data representing the three-dimensional shape of the object. Objects include not only characters that appear in games and anime, but also real people. By inputting a 3D model obtained by photographing or scanning a real person into a 3D printer, a figure can be created that is one-tenth the size of the person.
[0003] Patent Document 1 discloses a method for creating a doll of a desired object by a user selecting a desired scene or an object contained in the scene from a list of highlight video scenes created by filming a sports game. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] JP 2020-62322 A Summary of the Invention [Problem to be solved by the invention]
[0005] However, in Patent Document 1, an object to be modeled can only be selected from highlight scenes in the list, making it difficult to model a doll of an object included in a scene that is not in the list.
[0006] An object of the present disclosure is to easily create a three-dimensional object of an object in any scene of a moving image. [Means for solving the problem]
[0007] An information processing device according to one embodiment of the present disclosure is characterized by having an object for which modeling data is to be generated, a determination means for determining a priority for the object, an acquisition means for acquiring a 3D model of the object generated based on a plurality of captured images, and a generation means for generating modeling data including the 3D model based on the 3D model and the priority. [Effects of the Invention]
[0008] According to the present disclosure, it is possible to easily create a three-dimensional object of an object in any scene of a moving image. [Brief explanation of the drawings]
[0009] [Figure 1] A diagram showing an example of the configuration of an information processing system. [Figure 2] FIG. 1 is a diagram showing an example of the configuration of an image generating device. [Figure 3] A diagram showing a virtual camera and its operation screen. [Figure 4] Flowchart showing the flow of modeling data generation processing [Figure 5] A diagram showing the range of modeling data generation and examples of generation [Figure 6] An example of information managed by a database [Figure 7] Flowchart showing the flow of modeling data generation processing [Figure 8] A diagram showing an example of specifying and generating the drawing ranges of multiple virtual cameras. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the following embodiments do not limit the scope of the present disclosure according to the claims, and not all of the combinations of features described in the embodiments are necessarily essential to the solutions of the present disclosure. Note that the same components are given the same reference numerals, and their description will be omitted.
[0011] [Embodiment 1] In this embodiment, a system for generating modeling data will be described, which uses three-dimensional shape data (hereinafter referred to as a 3D model) representing the three-dimensional shape of an object, obtained from images captured from multiple viewpoints, for generating a virtual viewpoint image. A virtual viewpoint image is an image generated by an end user and / or a designated operator specifying the position and line of sight of a virtual viewpoint, and is also called a free viewpoint image or an arbitrary viewpoint image. A virtual viewpoint image may be a video or a still image, but this embodiment will describe a video image as an example. In the following description, the virtual viewpoint will be essentially replaced with a virtual camera (virtual camera). In the following description, the position of the virtual viewpoint corresponds to the position of the virtual camera, and the line of sight from the virtual viewpoint corresponds to the posture of the virtual camera.
[0012] (System Configuration) FIG. 1 is a diagram showing an example of the configuration of an information processing system (virtual viewpoint image generation system) that generates modeling data for an object in order to model a three-dimensional object of the object in a virtual viewpoint image. FIG. 1(a) shows an example of the configuration of an information processing system 100, and FIG. 1(b) shows an example of the installation of a sensor system included in the information processing system. The information processing system 100 includes n sensor systems 101a-101n, an image recording device 102, a database 103, an image generating device 104, and a tablet 105. Each of the sensor systems 101a-101n includes at least one camera, which is an imaging device. In the following, unless otherwise specified, the n sensor systems from sensor system 101a to sensor system 101n will be referred to as a multiple sensor system 101 without distinction.
[0013] An example of the installation of the multi-sensor system 101 and virtual cameras will be described with reference to Fig. 1(b). As shown in Fig. 1(b), the multi-sensor system 101 is installed to surround an area 120, which is a target area for imaging, and the cameras of the multi-sensor system 101 each capture images of the area 120 from different directions. The virtual camera 110 captures images of the area 120 from a different direction than the cameras of the multi-sensor system 101. Details of the virtual camera 110 will be described later.
[0014] If the imaging target is a professional sports match such as rugby or soccer, the area 120 is the field (ground) of a stadium, and n (e.g., 100) multi-sensor systems 101 are installed to surround the field. The imaging target area 120 may include not only people on the field but also a ball or other objects. The imaging target is not limited to a stadium field, but may also be a live music concert held in an arena or the like, or a commercial shoot held in a studio, as long as multiple sensor systems 101 can be installed. There is no limit to the number of sensor systems 101 to be installed. The multiple sensor systems 101 do not have to be installed around the entire perimeter of the area 120, and may be installed only in a portion of the perimeter of the area 120 depending on installation location restrictions, etc. The multiple cameras in the multi-sensor system 101 may include imaging devices with different functions, such as a telephoto camera and a wide-angle camera.
[0015] The multiple cameras included in the multi-sensor system 101 capture images synchronously to acquire multiple images. Each of the multiple images may be a captured image, or may be an image obtained by performing image processing on the captured image, such as a process of extracting a predetermined area.
[0016] Each of the sensor systems 101a-101n may have a microphone (not shown) in addition to a camera. The microphones of the multiple sensor systems 101 synchronously collect audio. Based on this collected audio, an acoustic signal can be generated that is played back along with the display of an image on the image generation device 104. For the sake of simplicity, a description of audio will be omitted below, but it is assumed that images and audio are basically processed together.
[0017] The image recording device 102 acquires multiple images from the multiple sensor system 101, and stores the acquired multiple images together with the time code used for capturing the images in a database 103. The time code is time information expressed as an absolute value for uniquely identifying the time at which an image was captured, and can be specified in a format such as day:hour:minute:second.frame number.
[0018] The database 103 manages event information, 3D model information, and the like. The event information includes data indicating the storage location of 3D model information for each object associated with all time codes of the event to be imaged. The objects may include people and objects that the user wants to model, or may include people and objects that are not the target of modeling. The 3D model information includes information about the 3D model of the object.
[0019] The image generating device 104 receives as input an image corresponding to a time code from the database 103 and information about a virtual camera 110 set by a user operation on a tablet 105. The virtual camera 110 is set in a virtual space associated with the area 120, and can view the area 120 from a viewpoint different from that of any of the cameras in the multi-sensor system 101. The virtual camera 110 and its operation and behavior will be described in detail below with reference to the accompanying drawings.
[0020] The image generating device 104 generates a 3D model for generating a virtual viewpoint image based on images for each time code acquired from the database 103, and generates a virtual viewpoint image using the generated 3D model for generating a virtual viewpoint image and information related to the viewpoint of the virtual camera. The viewpoint information of the virtual camera includes information indicating the position and orientation of the virtual viewpoint. Specifically, the viewpoint information includes parameters representing the three-dimensional position of the virtual viewpoint and parameters representing the orientation of the virtual viewpoint in the pan direction, tilt direction, and roll direction. The virtual viewpoint image is an image representing the view from the virtual camera 110, and is also called a free viewpoint image. The virtual viewpoint image generated by the image generating device 104 is displayed on a touch panel such as a tablet 105.
[0021] In this embodiment, the image generation device 104 generates modeling data (shape data of an object) to be used by the modeling device 106 from a 3D model corresponding to an object present in the drawing range, based on the drawing range of at least one virtual camera. That is, the image generation device 104 sets conditions for identifying an object to be modeled for the virtual viewpoint image, acquires shape data representing the three-dimensional shape of the object based on the set conditions, and generates modeling data based on the acquired shape data. Details of the modeling data generation process will be described later with reference to the drawings. Note that the format of the modeling data generated by the image generation device 104 may be any format that can be processed by the modeling device 106, such as a general-purpose polygon mesh format for handling general 3D models, or a format unique to the modeling device 106. That is, the modeling device 106 can process shape data for generating a virtual viewpoint image, and the image generation device 104 identifies the shape data for generating the virtual viewpoint image as modeling data. On the other hand, the modeling device 106 cannot process shape data for generating a virtual viewpoint image, and the image generation device 104 generates modeling data using shape data for generating a virtual viewpoint image.
[0022] The tablet 105 is a portable device having a touch panel that functions as both a display unit for displaying images and an input unit for accepting user operations. The tablet 105 may also be a portable device with other functions, such as a smartphone. The tablet 105 accepts a user operation for setting information about the virtual camera. The tablet 105 also displays a virtual viewpoint image generated by the image generation device 104 or a virtual viewpoint image stored in the database 103. The tablet 105 then accepts a user operation for setting conditions for identifying an object to be modeled for the virtual viewpoint image. This user operation sets an arbitrary spatial range and a temporal range for generating modeling data. Details of this operation method will be described later with reference to the accompanying drawings. Note that the operation of the virtual camera is not limited to a user operation on the touch panel of the tablet 105, but may also be a user operation on an operating device such as a three-axis controller.
[0023] The image generation device 104 and the tablet 105 may be configured as separate devices as shown in Fig. 1, or may be configured as an integrated device. In the case of an integrated device, the image generation device 104 has a touch panel or the like, receives operations of the virtual camera, and displays a virtual viewpoint image generated by the image generation device 104 on the touch panel. Note that the virtual viewpoint image may be displayed on a liquid crystal screen of a device other than the tablet 105 or the image generation device 104.
[0024] The modeling device 106 is, for example, a 3D printer, and receives modeling data generated by the image generating device 104 as input to model a corresponding three-dimensional object, such as a doll (3D model figure) or a relief. The modeling method of the modeling device 106 is not limited to optical modeling, inkjet printing, powder fixing, etc., as long as it can model a three-dimensional object. The modeling device 106 is not limited to a device that outputs a three-dimensional object such as a doll, and may also be a device that prints on a plate or paper.
[0025] The image generation device 104 and the modeling device 106 may be configured as separate devices as shown in FIG. 1, or may be configured as an integrated device.
[0026] 1(a), in which the tablet 105 and the modeling device 106 are connected one-to-one to the image generation device 104. For example, the information processing system 100 may be an information processing system in which a plurality of tablets 105 and a plurality of modeling devices 106 are connected to the image generation device 104.
[0027] (Configuration of image generating device) An example of the configuration of the image generation device 104 will be described with reference to the drawings. Fig. 2 is a diagram showing an example of the configuration of the image generation device 104, in which Fig. 2(a) shows an example of the functional configuration of the image generation device 104 and Fig. 2(b) shows an example of the hardware configuration of the image generation device 104.
[0028] 2(a), the image generating device 104 includes a virtual camera control unit 201, a 3D model generating unit 202, an image generating unit 203, and a modeling data generating unit 204. The image generating device 104 uses the above-mentioned functional units to generate modeling data based on the drawing range of at least one virtual camera. An overview of each function will be described here, and details of the processing will be described later.
[0029] The virtual camera control unit 201 receives virtual camera operation information from the tablet 105 or the like. The virtual camera operation information includes at least the position and attitude of the virtual camera, and a time code. Details of the virtual camera operation information will be described later using figures. Note that if the image generation device 104 has a touch panel or the like and is configured to be able to receive virtual camera operation information, the virtual camera control unit 201 receives the virtual camera operation information from the image generation device 104.
[0030] The 3D model generation unit 202 generates a 3D model representing the three-dimensional shape of an object in the area 120 based on a plurality of captured images. Specifically, the 3D model generation unit 202 acquires, from the images, a foreground image in which a foreground area corresponding to an object such as a person or a ball is extracted, and a background image in which a background area other than the foreground area is extracted. Then, the 3D model generation unit 202 generates a foreground 3D model for each object based on the plurality of foreground images.
[0031] These 3D models are generated using a shape estimation method such as Visual Hull and are composed of point clouds. However, the data format of the 3D models representing the shapes of each object is not limited to this. Note that the background 3D model may be acquired in advance by an external device.
[0032] The 3D model generation unit 202 records the generated 3D model together with a time code in the database 103.
[0033] The 3D model generation unit 202 may be configured to be included in the image recording device 102, rather than in the image generation device 104. In that case, the image generation device 104 simply reads the 3D model generated by the image recording device 102 from the database 103 via the 3D model generation unit 202.
[0034] The image generation unit 203 acquires a 3D model from the database 103 and generates a virtual viewpoint image based on the acquired 3D model. Specifically, the image generation unit 203 acquires appropriate pixel values from the image for each point constituting the 3D model and performs coloring processing. The image generation unit 203 then places the colored 3D model in a three-dimensional virtual space, projects it onto a virtual camera (virtual viewpoint), and renders it to generate a virtual viewpoint image.
[0035] However, the method for generating the virtual viewpoint image is not limited to this, and various methods may be used, such as a method for generating the virtual viewpoint image by projective transformation of a captured image without using a 3D model.
[0036] The modeling data generation unit 204 calculates and determines the drawing range or projection range of the virtual camera using the position and orientation of the virtual camera and the time code.The modeling data generation unit 204 then determines the generation range of modeling data based on the determined drawing range of the virtual camera, and generates modeling data from a 3D model included in that range.These processes will be described in detail later with reference to the drawings.
[0037] (Hardware configuration of image generation device) Next, the hardware configuration of the image generation device 104 will be described with reference to Fig. 2(b). As shown in Fig. 2(b), the image generation device 104 has a CPU 211, a RAM 212, a ROM 213, an operation input unit 214, a display unit 215, and a communication I / F (interface) unit 216.
[0038] A CPU (Central Processing Unit) 211 performs processing using programs and data stored in a RAM (Random Access Memory) 212 and a ROM (Read Only Memory) 213 .
[0039] The CPU 211 controls the overall operation of the image generating device 104 and executes processing to realize each function shown in Fig. 2(a). Note that the image generating device 104 may have one or more pieces of dedicated hardware different from the CPU 211, and at least a part of the processing by the CPU 211 may be executed by the dedicated hardware. Examples of the dedicated hardware include an ASIC (application specific integrated circuit), an FPGA (field programmable gate array), and a DSP (digital signal processor).
[0040] The ROM 213 holds programs and data. The RAM 212 has a work area for temporarily storing programs and data read from the ROM 213. The RAM 212 also provides a work area used by the CPU 211 when it executes various processes.
[0041] The operation input unit 214 is, for example, a touch panel, and receives input operations from a user and acquires information input by the received user operation. Examples of the input information include information about the virtual camera and information about the time code of the virtual viewpoint image to be generated. The operation input unit 214 may be connected to an external controller and receive input information from the user regarding operations. The external controller is, for example, an operation device such as a three-axis controller such as a joystick or a mouse. The external controller is not limited to these.
[0042] The display unit 215 is a touch panel or a screen that displays the generated virtual viewpoint image. In the case of a touch panel, the operation input unit 214 and the display unit 215 are integrated into one unit.
[0043] The communication I / F unit 216 transmits and receives information to and from the database 103, the tablet 105, the modeling device 106, and the like via, for example, a LAN. The communication I / F unit 216 may also transmit information to an external screen via an image output port that complies with the following communication standards. Examples of image output ports that comply with communication standards include HDMI (registered trademark) (High-Definition Multimedia Interface) and SDI (Serial Digital Interface). The communication I / F unit 216 may also transmit image data and modeling data via, for example, Ethernet.
[0044] (Virtual camera (virtual viewpoint) and its operation screen) Next, we will explain the virtual camera and the operation screen for setting it up, using an example in which a rugby match is captured on a rugby field and a captured image is acquired. Figure 3 is a diagram showing the virtual camera and its operation screen. Figure 3(a) shows a coordinate system, Figure 3(b) shows an example of a field to which the coordinate system of Figure 3(a) is applied, Figures 3(c) and 3(d) show examples of the virtual camera's drawing range, and Figure 3(e) shows an example of virtual camera movement. Figure 3(f) shows an example of a virtual viewpoint image as seen from the virtual camera.
[0045] First, a coordinate system representing the three-dimensional space of the object to be imaged, which serves as a reference for setting the virtual viewpoint, will be described. As shown in FIG. 3(a), this embodiment uses a Cartesian coordinate system in which the three-dimensional space is represented by three axes: X, Y, and Z. This Cartesian coordinate system is set for each object shown in FIG. 3(b), i.e., the rugby field 391, the ball 392 on the field, and the players 393. Furthermore, the Cartesian coordinate system may also be set for facilities (structures) within the rugby field, such as spectator seats and signs around the field 391. Specifically, the origin (0, 0, 0) is set at the center of the field 391. The X axis is set in the direction of the long side of the field 391, the Y axis is set in the direction of the short side of the field 391, and the Z axis is set in the direction perpendicular to the field 391. Note that the directions of the axes are not limited to these. Using such a coordinate system, the position and orientation of the virtual camera 110 are specified.
[0046] Next, the rendering range of the virtual camera will be explained using diagrams. In a quadrangular pyramid 300 shown in FIG. 3(c), vertex 301 represents the position of the virtual camera 110, and a line-of-sight vector 302 with vertex 301 as the base point represents the orientation of the virtual camera 110. Note that vector 302 is also called the optical axis vector of the virtual camera. The position of the virtual camera is expressed by the components (x, y, z) of each axis, and the orientation of the virtual camera 110 is expressed by a unit vector with the components of each axis as scalars. Vector 302 representing the orientation of the virtual camera 110 passes through the center points of the front clip plane 303 and the back clip plane 304. The viewing frustum of the virtual camera, which is the projection range (rendering range) of the 3D model, is a space 305 sandwiched between the front clip plane 303 and the back clip plane 304.
[0047] Next, components indicating the rendering range of the virtual camera will be explained using diagrams. FIG. 3(d) is a view of the virtual viewpoint in FIG. 3(c) as seen from above (Z axis). The rendering range is determined by the following values. The values of the distance 311 from the vertex 301 to the front clip plane 303, the distance 312 from the vertex 301 to the back clip plane 304, and the angle of view 313 of the virtual camera 110 may be predetermined values (default values) set in advance, or may be set values changed by user operation. The angle of view 313 may also be a value calculated based on a separate variable, the focal length of the virtual camera 110. The relationship between the angle of view and focal length is a common technique, and a description thereof will be omitted.
[0048] Next, a change in the position of the virtual camera 110 (movement of the virtual viewpoint) and a change in the attitude (rotation) of the virtual camera 110 will be described. The virtual viewpoint can be moved and rotated within a space expressed in three-dimensional coordinates. FIG. 3(e) is a diagram illustrating the movement and rotation of the virtual camera. In FIG. 3(e), a dashed-dotted arrow 306 represents the movement of the virtual camera (virtual viewpoint), and a dashed-dotted arrow 307 represents the rotation of the moved virtual camera (virtual viewpoint). The movement of the virtual camera is expressed by the components of each axis (x, y, z), and the rotation of the virtual camera is expressed by yaw, which is the rotation around the Z axis, pitch, which is the rotation around the X axis, and roll, which is the rotation around the Y axis. In this way, since the virtual camera can freely move and rotate in the three-dimensional space of the subject (field), a virtual viewpoint image can be generated in which any region of the subject becomes the drawing range.
[0049] Next, an operation screen for setting the position and attitude of the virtual camera (virtual viewpoint) will be described. Fig. 3(f) is a diagram for explaining an example of an operation screen for the virtual camera (virtual viewpoint).
[0050] In this embodiment, the generation range of the modeling data is determined based on the drawing range of at least one virtual camera. Therefore, the virtual camera operation screen 320 shown in Fig. 3(f) can also be said to be an operation screen for determining the generation range of the modeling data. In Fig. 3(f), the virtual camera operation screen 320 is displayed on the touch panel of the tablet 105. Note that the display destination of the virtual camera operation screen 320 is not limited to this, and it may also be the touch panel of the image generation device 104, etc.
[0051] On the operation screen 320, the drawing range of the virtual camera (the drawing range related to capturing the virtual viewpoint image) is displayed as a virtual viewpoint image in accordance with the screen frame of the operation screen 320. With such a display, the user can visually set the conditions for specifying the object to be modeled.
[0052] The operation screen 320 has a virtual camera operation area 322 that accepts user operations for setting the position and orientation of the virtual camera 110, and a time code operation area 323 that accepts user operations for setting a time code. First, the virtual camera operation area 322 will be described. Because the operation screen 320 is displayed on a touch panel, the virtual camera operation area 322 accepts general touch operations 325, such as tapping, swiping, and pinching in and out, as user operations. This touch operation 325 adjusts the position and focal length (angle of view) of the virtual camera. The virtual camera operation area 322 also accepts touch operations 324, such as continuously pressing each axis of a Cartesian coordinate system, as user operations. This touch operation 324 rotates the virtual camera 110 around the X-axis, Y-axis, or Z-axis, adjusting the orientation of the virtual camera. In this way, by assigning movement, rotation, and scaling of the virtual camera to each touch operation on the operation screen 320, the virtual camera 110 can be freely operated. These operation methods are well known, and their description will be omitted.
[0053] Note that the operation regarding the position and orientation of the virtual camera is not limited to a touch operation on a touch panel, but may be an operation using an operating device such as a joystick.
[0054] Next, we will explain the time code operation area 323. The time code operation area 323 has multiple elements for operating the time code, including a main slider 332 and a knob 342, and a sub-slider 333 and a knob 343. The time code operation area 323 also has an add button 350 for a virtual camera and an output button 351.
[0055] The main slider 332 is an input element that can be used to set a desired time code from among all the time codes in the imaging data. When the position of the knob 342 is moved to the desired position by dragging or the like, the time code corresponding to the position of the knob 342 is specified. In other words, by adjusting the position of the knob 342 of the main slider 332, an arbitrary time code can be specified.
[0056] The sub-slider 333 is an input element that displays an enlarged portion of the total timecode, allowing for operations to set the timecode in detail for that enlarged portion. When the knob 343 is moved to the desired position by dragging or other operations, the timecode corresponding to the knob 343's position is specified. The main slider 332 and sub-slider 333 are the same length on the screen, but the range of timecode that can be selected differs. For example, the main slider 332 allows selection from three hours, the length of one game, while the sub-slider 333 allows selection from a portion of that, such as 30 seconds. In this way, the slider scales are different, and the sub-slider allows for more precise timecode specification, such as in seconds or frames.
[0057] The time code specified using the knob 342 of the main slider 332 or the knob 343 of the sub-slider 333 may be displayed as a number in the format of day:hour:minute:second.frame number. The sub-slider 333 may be displayed constantly on the operation screen 320, or may be displayed temporarily. For example, it may be displayed when an instruction to display the time code is received, or when an instruction for a specific operation such as pausing is received. The time code section that can be selected by the sub-slider 333 may be variable.
[0058] The virtual viewpoint image generated according to the position and attitude of the virtual camera set by the above operations and the time code is displayed in the virtual camera operation area 322. In Fig. 3(f), as an example, the subject is a rugby match, and a decisive pass scene that leads to a goal is displayed. As will be described in detail later, this decisive scene is generated or output as shape data of an object to be modeled.
[0059] FIG. 3(f) shows the case where the time code of the moment the ball is released is specified, but by operating the sub-slider 333 in the time code operation area 323, for example, it is also possible to easily specify the time code when the ball is in the air.
[0060] 3(f), the position and orientation of the virtual camera are specified so that three players are present in the drawing range, but this is not limiting. For example, by performing touch operations 324 and 325 on the virtual camera operation area 322, the space (drawing range of the virtual camera) 305 can be freely manipulated, and further manipulation can be performed so that surrounding players are present in the drawing range.
[0061] The output button 351 is a button that is operated when determining the drawing range of the virtual camera set by user operations in the virtual camera operation area 322 and the time code operation area 323, and outputting shape data of an object for modeling. When a time code indicating a time range, a position indicating a spatial range, and an attitude thereof are determined for the virtual camera, the drawing range of the virtual camera corresponding to these determinations is calculated. Then, based on the drawing range of the virtual camera that is the calculation result, a generation range of the modeling data is determined. Details of this process will be described later using figures.
[0062] The add virtual camera button 350 is a button used when using multiple virtual cameras to generate modeling data. Details of the process of generating modeling data using multiple virtual cameras will be described in the second embodiment, and therefore will not be described here.
[0063] 3(f), any operation screen for the virtual camera may be used as long as it allows for operations to set the position and orientation of the virtual camera and the time code. For example, the virtual camera operation area 322 and the time code operation area 323 do not need to be separated. For example, when an operation such as a double tap is performed on the virtual camera operation area 322, a pause or the like may be processed as an operation for setting the time code.
[0064] Furthermore, although the case where the operation input unit 214 is the tablet 105 has been described, the operation input unit 214 is not limited to this, and may be an operation device having a general display, a three-axis controller, or the like.
[0065] (Modeling data generation process) Next, the generation process of the modeling data in the image generating device 104 will be described with reference to the drawings. FIG. 4 is a flowchart showing the flow of the generation process of the modeling data. This series of processes is realized by the CPU 211 executing a predetermined program and operating each functional unit shown in FIG. 2(a). Hereinafter, steps will be abbreviated as "S". This also applies to the following explanation. The description will be made on the assumption that shape data representing the three-dimensional shape of an object used to generate a virtual viewpoint image is in a format different from that of the modeling data.
[0066] In S401, the modeling data generation unit 204 accepts the specification of the time code of the virtual viewpoint image via the virtual camera control unit 201. The time code may be specified, for example, by using a slider as shown in FIG. 3(f) or by directly inputting numbers.
[0067] In S402, the modeling data generation unit 204 receives operation information for the virtual camera via the virtual camera control unit 201. The operation information for the virtual camera includes at least information about the position and attitude of the virtual camera. The virtual camera may be operated using, for example, a tablet as shown in FIG. 3(f) or an operation device such as a joystick.
[0068] In S403, the modeling data generation unit 204 determines whether an output instruction has been received via the virtual camera control unit 201, i.e., whether the output button 351 has been operated. If the modeling data generation unit 204 obtains a determination result that the output instruction has been received (YES in S403), the process proceeds to S404. If the modeling data generation unit 204 obtains a determination result that the output instruction has not been received (NO in S403), the process returns to S401, and the processes of S401 and S402 are executed again. That is, the process of receiving user operations for the time code of the virtual viewpoint image and the position and attitude of the virtual camera continues until an output instruction is received.
[0069] In S404, the modeling data generation unit 204 determines the rendering range of the virtual camera using the time code of the virtual viewpoint image received in the process of S401 and the viewpoint information of the virtual camera, including the position and attitude of the virtual camera, received in the process of S402. The rendering range of the virtual camera may be determined, for example, by a user operation of the front clipping plane 303 and the rear clipping plane 304 shown in Figures 3(c) and 3(d), or may be determined based on the results of calculations using calculation formulas that are preset in the device.
[0070] In S405, the modeling data generation unit 204 determines the generation range of modeling data corresponding to the virtual camera based on the rendering range of the virtual camera determined in S404. A method for determining the generation range of modeling data will be described. FIG. 5 illustrates the generation range of modeling data and an example of its generation. FIG. 5(a) illustrates the generation range of modeling data in three-dimensional space, and FIG. 5(b) illustrates the yz plane of the generation range of the modeling data shown in FIG. 5(a). FIG. 5(c) illustrates an example of a 3D model within the generation range of modeling data in three-dimensional space. FIG. 5(d) illustrates the yz plane when spectator seats are present within the generation range of the modeling data. FIG. 5(e) illustrates an example of a 3D model figure corresponding to the 3D model shown in FIG. 5(c). FIG. 5(f) illustrates another example of a 3D model figure corresponding to the 3D model shown in FIG. 5(c).
[0071] 5(a), a virtual camera (position 301, space (view frustum) 305, etc.) is displayed in three-dimensional space, and a plane 500 of the image capture target at Z=0 (for example, a stadium field) is displayed in the same space. As shown in FIG. 5(a), a generation range 510 of modeling data is determined based on a plane (bottom) 501 of the plane 500 at Z=0 that is included in the space (view frustum) 305, which is the drawing range of the virtual camera.
[0072] The generation range 510 of the modeling data is a part of the space (view frustum) 305, which is the rendering range of the virtual camera. The generation range 510 of the modeling data is a space surrounded by the plane 501, a front plane (front surface) 513 located near the front clip plane 303, and a rear plane (rear surface) 514 located near the rear clip plane 304. The sides of the space that becomes the generation range 510 of the modeling data are set based on the position 301 of the virtual camera, its angle of view, and the plane 500. In addition, the top of the space that becomes the generation range 510 of the modeling data is set based on the position 301 of the virtual camera and its angle of view.
[0073] The plane 513 and the plane 514 will be described with reference to FIG. 5(b). FIG. 5(b) is a simplified side view of the virtual camera and the plane at Z=0 in FIG. 5(a). The plane 513 passes through the intersection 503 or 505, is located a predetermined distance from the front clipping plane 303 in the space 305, and is parallel to the front clipping plane 303. The plane 514 passes through the intersection 504 or 506, is located a predetermined distance from the rear clipping plane 304 in the space 305, and is parallel to the rear clipping plane 304. The predetermined distance is set in advance. The plane 501 is a rectangle with the above-mentioned intersections 503 to 506 as vertices.
[0074] The plane 514 may be the same as the rear clipping plane 304. The planes 513 and 514 do not have to be parallel to the front clipping plane and the rear clipping plane, respectively. For example, the planes 513 and 514 may be planes that pass through the vertex of the plane 501 and are perpendicular to the plane 501.
[0075] In addition to the above, the generation range 510 of the modeling data may be determined taking into account the spectator seats in the stadium, etc. This example will be described with reference to FIG. 5(d). Similar to FIG. 5(b), FIG. 5(d) shows the rendering range when the virtual camera in FIG. 5(a) and the plane at Z=0 are viewed from the side in a simplified manner. In FIG. 5(d), a plane 514 may be determined based on the position of the stadium spectator seats 507. Specifically, the plane 514 is a plane that passes through the intersection of the spectator seats 507 and the plane 501 and is parallel to the rear clip plane. This prevents the area behind the spectator seats 507 from being included in the generation range 510 of the modeling data.
[0076] The items that can be added to the conditions for determining the generation range of the modeling data are not limited to stadium seats, and the user may manually specify the object, etc. The plane 500 that intersects with the view frustum is not limited to Z=0, and may have any shape. For example, it may have irregularities like an actual field. If the field, which is the 3D model of the actual background, has irregularities, it may be corrected to a flat 3D model. The configuration may not have a plane or curved surface that intersects with the view frustum. In that case, the drawing range of the virtual camera may be used as the generation range of the modeling data as is.
[0077] In S406, the modeling data generation unit 204 acquires, from the database 103, a 3D model of an object included in the generation range of the modeling data determined in S405.
[0078] Examples of tables of event information and 3D model information managed by the database 103 will be described with reference to the drawings. Fig. 6 shows tables of information managed by the database 103, with Fig. 6(a) showing the table of event information and Fig. 6(b) showing the table of 3D model information. As shown in Fig. 6(a), the table 610 of event information managed by the database 103 indicates the storage location of 3D model information for each object for all time codes of the event to be imaged. For example, the table 610 of event information indicates that the storage location of 3D model information for object A when the time code is "16:14:24.041" is "DataA100."
[0079] As shown in FIG. 6B, the 3D model information table 620 managed by the database 103 stores data for each of the following items: "Total Point Cloud Coordinates," "Texture," "Average Coordinates," "Gravity Center Coordinates," and "Maximum and Minimum Coordinates." The "Total Point Cloud Coordinates" field stores data related to the coordinates of each point in the point cloud that constitutes the 3D model. The "Texture" field stores data related to the texture image to be applied to the 3D model. The "Average Coordinates" field stores data related to the coordinates of a point calculated by averaging all of the coordinates of the point cloud that constitutes the 3D model. The "Gravity Center Coordinates" field stores data related to the coordinates of the point that serves as the gravity center based on all of the coordinates of the point cloud that constitutes the 3D model. The "Maximum and Minimum Coordinates" field stores data related to the coordinates of the maximum and minimum points among the coordinates of the point cloud that constitutes the 3D model. Note that the data items stored in the 3D model information table 620 are not limited to "Total Point Cloud Coordinates," "Texture," "Average Coordinates," "Gravity Center Coordinates," and "Maximum and Minimum Coordinates." For example, the table may contain only "Total Point Cloud Coordinates" and "Texture," or other items may be added to these fields.
[0080] By using the information shown in Figure 6, when a certain time code is specified, it is possible to obtain 3D model information at the specified time code, such as the coordinates of all point groups for each object and the coordinates of the maximum and minimum values on each axis of the three-dimensional coordinate system.
[0081] 5(c) will be used to describe an example in which a 3D model included in the generation range of the modeling data is acquired from the database 103. The modeling data generation unit 204 refers to the 3D model information for each object associated with the time code specified in S401 in the database 103, and determines whether or not each object is included in the generation range of the modeling data determined in S406.
[0082] The determination method may be, for example, a method of determining whether all point cloud coordinates included in the 3D model information for each object, such as a person, are included in the generation range, or a method of determining whether only the average value of all point cloud coordinates or the maximum and minimum values of each axis of the three-dimensional coordinates are included in the generation range.
[0083] Fig. 5(c) shows an example in which, as a result of the above-mentioned determination, three 3D models are included in the modeling data generation range 510. As will be described in detail later, an example in which these three 3D models become modeling data is shown as 3D models 531-533 in Fig. 5(e).
[0084] The determination result as to whether each object is included in the generation range of the modeling data may be displayed on the tablet operation screen 320. As the determination result, for example, for an object that is on the boundary of the generation range 510 of the modeling data, a warning may be displayed to the effect that the entire object will not be output as modeling data because it is on the boundary.
[0085] The determination process in S406 may be executed at any time while accepting the time code of the virtual viewpoint image in S401 and the user operation of the position and attitude of the virtual camera in S402. When such determination process in S406 is executed at any time, a warning message may be displayed to notify that the target object is on the boundary.
[0086] The method for acquiring a 3D model of an object is not limited to the automatic acquisition method based on the above-described determination result, but may be a manual acquisition method. For example, an object for which a 3D model is to be acquired may be designated by accepting a user operation, such as a tap operation, on the tablet operation screen 310 for an object for which the user wants to acquire modeling data.
[0087] In addition, in the process of S406, 3D model information of the background field and spectator seats, which are objects included in the generation range of the modeling data, may be obtained from a database. In this case, the obtained 3D models of the field and spectator seats may be partially converted to form a pedestal. For example, even if the 3D model of the field has no thickness, it may be converted into a 3D model with a specified height added, thereby forming pedestal 540 in a rectangular parallelepiped shape, as shown in FIG. 5(e).
[0088] In S407, the modeling data generation unit 204 generates a 3D model that does not actually exist but is intended to support the 3D model acquired from the database. For example, the player in the 3D figure model 532 in the foreground of FIG. 5(e) is jumping and is in the air, not touching the field. Therefore, to display the 3D figure model, a support (supporting part) is required to support the 3D figure model and secure it to a pedestal. Therefore, in S407, the modeling data generation unit 204 first determines whether a supporting part is required using the 3D model information for each object acquired from the database 103. This determination includes determining whether the target 3D model exists in the air and whether the target 3D model can stand on the pedestal. The determination of whether the target 3D model exists in the air may be made using, for example, the coordinates or minimum coordinates of all point clouds included in the 3D model information. The determination of whether the target 3D model can stand on the pedestal may be made using, for example, the coordinates of all point clouds included in the 3D model information, or the center of gravity coordinates and maximum and minimum coordinates.
[0089] If the determination result indicates that an auxiliary part is necessary, the modeling data generation unit 204 generates a 3D model of an auxiliary part that secures the target 3D figure model to the pedestal based on the 3D model information. If the target 3D model exists in the air, the modeling data generation unit 204 may generate a 3D model of the auxiliary part based on the coordinates of the entire point cloud, or the minimum coordinate and its surrounding coordinates. If the target 3D model does not stand on the pedestal, the modeling data generation unit 204 may generate a 3D model of the auxiliary part based on the coordinates of the entire point cloud, or the center of gravity coordinates and the maximum and minimum coordinates. Note that the support pillar may be positioned to support the object vertically from the center of gravity coordinate included in the 3D model information for each object. Furthermore, a 3D model of a support pillar may be added even for objects that have contact points in the field.
[0090] On the other hand, if the determination result indicates that the auxiliary part is not necessary, the modeling data generation unit 204 proceeds to step S408 without generating a 3D model of the auxiliary part.
[0091] In S408, the modeling data generation unit 204 combines the set of 3D models generated in the processing up to S407, converts the combined set of 3D models into the format of the output destination, and outputs it as modeling data for a three-dimensional object. An example of the modeling data to be output will be described with reference to FIG. 5(e).
[0092] In Fig. 5(e), the modeling data includes foreground 3D models 531, 532, and 533, which are included in a generation range 510 of the modeling data, a support 534, and a pedestal 540. The coordinates of the original foreground 3D models of people, etc., and the background 3D model of the field, etc., are expressed by the single three-dimensional coordinate system shown in Fig. 3(a), and the positional relationship between these can accurately reproduce the positions and postures of players on the actual field.
[0093] 5(e) is output to the modeling device 106, the modeling device 106 can model a 3D model figure using this shape as is. For this reason, although FIG. 5(e) is shown as modeling data, it may also be regarded as a 3D model figure output by a 3D printer.
[0094] Note that, in order to reduce the risk of damage during delivery or transportation, the modeling data of FIG. 5(e) may be output separately as a foreground 3D model and a background 3D model in S408. For example, as shown in FIG. 5(f), a pedestal 540 corresponding to the background field and 3D models 531, 532, and 533 corresponding to the foreground people may be output separately, and small sub-pedestals 541, 542, and 543 may be attached to the 3D models 531, 532, and 533. In this case, the foreground 3D models 531, 532, and 533 can be attached to the pedestal 540 by providing recesses 551-553 in the pedestal 540 that correspond to the sizes of the sub-pedestals 541-543. Since the sub-pedestals 541, 542, and 543 are 3D models that do not actually exist, they may be added in S407 as support for the foreground 3D models, similar to supports.
[0095] The shapes of the sub-pedestals 541-543 and the recesses 551-553 may be polygonal, as well as rectangular, as shown in FIG. 5(f). The shapes of the sub-pedestals 541-543 and the recesses 551-553 may be different, like puzzle pieces. In this case, after the foreground 3D model and the background 3D model are separately output by the modeling device, the foreground 3D model can be properly attached without making mistakes in the attachment position or attachment direction. The positions of the sub-pedestals 541-543 for each object and the recesses 551-553 on the pedestal 540 may be changeable by user operation.
[0096] It is also possible to output modeling data in which information that does not actually exist is embedded in the pedestal 540. For example, it is also possible to output modeling data for the pedestal 540 in which the time code of the virtual camera used in the corresponding modeling data, or information regarding the position and attitude of the virtual camera, is embedded. It is also possible to output modeling data for the pedestal 540 in which three-dimensional coordinate values are embedded so that the generation range of the modeling data can be determined.
[0097] As described above, according to this embodiment, it is possible to generate modeling data for an object corresponding to the specification information specified by a user operation. That is, it is possible to generate modeling data for an object included in the drawing range of the virtual camera based on the drawing range of the virtual camera corresponding to the specification information.
[0098] For example, in field sports such as rugby or soccer played in a stadium, a decisive scene leading to a goal can be captured at any time and in any spatial range, and modeling data can be generated that preserves the actual positions and postures of the players. Furthermore, this data can be used to generate a 3D model figure of the scene by outputting it to a 3D printer.
[0099] Furthermore, if the shape data representing the three-dimensional shape of the object used to generate the virtual viewpoint image has the same format as the modeling data of the object, the following processing is performed in S408. That is, the modeling data generation unit 204 outputs data in which the 3D model of the auxiliary part is added to the 3D model of the object generated in the processing up to S407 (the previous step), as modeling data of the object. In this way, if the format of the shape data for generating the virtual viewpoint image is the same as the format of the modeling data, the shape data for generating the virtual viewpoint image can be output as is as modeling data.
[0100] [Embodiment 2] In this embodiment, a mode will be described in which a plurality of virtual cameras are used and modeling data is generated based on the rendering ranges of the cameras.
[0101] In this embodiment, the configuration of the information processing system is the same as that shown in Figure 1, and the configuration of the image generation device is the same as that shown in Figure 2, so their explanations will be omitted and only the differences will be explained. In this embodiment, the explanation will mainly focus on the operation method for using multiple virtual cameras on the tablet 105 or the image generation device 104, and the generation process of modeling data based on these multiple drawing ranges.
[0102] (Modeling data generation process) The process of generating shaping data according to this embodiment will be described with reference to the accompanying drawings. FIG. 7 is a flowchart showing the process of generating shaping data according to this embodiment. This process is realized by the CPU 211 executing a predetermined program and operating the respective functional units shown in FIG. 2(a). FIG. 8 is a diagram illustrating an example of an operation screen for a virtual viewpoint (virtual camera). FIG. 8(a) shows a case where virtual camera 1 is selected, FIG. 8(b) shows a case where virtual camera 2 is selected, FIG. 8(c) shows a case where the priority setting screen in FIG. 8(b) is displayed, and FIG. 8(d) shows an example of a 3D model. The operation screens shown in FIGS. 8(a) to 8(c) are screens with expanded functions from the operation screen shown in FIG. 3(f), and only the differences will be described here.
[0103] In S701, the modeling data generation unit 204 specifies a virtual camera that will accept operations. In this embodiment, since multiple virtual cameras are handled, in S702 and S703 following the processing of S701, a specification is made to identify which virtual camera will accept operations. Note that in the initial state, since there is only one virtual camera, the identifier of that single camera is specified, and if there are multiple virtual cameras, the identifier of one camera selected by the user is specified. A method for selecting a virtual camera on the operation screen will be described later.
[0104] In S702, the modeling data generation unit 204 accepts the specification of the time code of the virtual viewpoint image obtained by capturing an image using the virtual camera specified in S701 via the virtual camera control unit 201. The method of specifying the time code is the same as the method of specifying the time code described in S401, and therefore, the description thereof will be omitted.
[0105] In S703, the modeling data generation unit 204 receives operation information of the virtual camera specified in S701 via the virtual camera control unit 201. The method of operating the virtual camera is the same as the method of operating the virtual camera described in S402, and therefore, the description thereof will be omitted.
[0106] In S704, the modeling data generation unit 204 receives one of the following instructions from the user and determines which instruction the received instruction corresponds to. The user instruction received here is one of selecting a virtual camera, adding a virtual camera, and outputting modeling data. "Selecting a virtual camera" is an instruction to select a different virtual camera displayed on the tablet 105 than the virtual camera specified in S701. "Adding a virtual camera" is an instruction to add a different virtual camera not displayed on the tablet 105 than the virtual camera specified in S701. The instruction to add a virtual camera is issued, for example, when the add button 350 shown in FIG. 8(a) is pressed. "Outputting modeling data" is an instruction to output modeling data.
[0107] If the determination result indicates that the operation instruction is "add virtual camera," the modeling data generation unit 204 proceeds to step S705.
[0108] In S705, the modeling data generation unit 204 adds a virtual camera. The screen displayed on the tablet 105 switches from the operation screen shown in Fig. 8(a) to the operation screen shown in Fig. 8(b), and a new tab 802 is added to the operation screen 320, and the virtual camera is also added. As a result, on the added tab screen, it is possible to receive the time code of the virtual viewpoint image obtained by capturing an image with the added virtual camera, and operations on the position and attitude of the added virtual camera.
[0109] Figure 8 displays virtual viewpoint images of a volleyball player spiking the ball as an example of a sports player's playing scene. Figure 8(a) shows the scene where the volleyball player takes a crouched stance before jumping, and Figure 8(b) shows the scene where the time code is advanced from the scene in Figure 8(a) and the volleyball player takes a raised arm stance before jumping and hitting the ball.
[0110] 8(a) to 8(c) show the case where the number of tabs and the number of virtual cameras are two, but this is not limiting. The number of tabs and the number of virtual cameras may be three or more. As will be described later, the output example in FIG. 8(d) uses six virtual cameras, and captures the volleyball player's spike scenes in chronological order by manipulating the time codes of the virtual viewpoint images captured by each virtual camera, as well as the position and attitude of each virtual camera.
[0111] In addition, the time code of the virtual viewpoint image obtained by capturing an image using the added virtual camera, and the position and attitude of the added virtual camera may be set to the values of the virtual camera displayed on the screen when the additional instruction was received, as initial values.
[0112] If the determination result in S704 above indicates that the operation instruction is "selection of a virtual camera," the modeling data generation unit 204 returns the process to S701 and executes the series of processes from S701 to S703 for the selected virtual camera. The instruction to select a virtual camera is given, for example, by a user selecting tabs 801-802 shown in FIG. 8(a). When an image including tabs 801-802 is displayed on a touch panel, the user may perform a general touch operation.
[0113] In the spike scene in Figure 8, by manipulating the time codes of the virtual viewpoint images captured by each virtual camera in detail, it is possible to easily specify the scene where the volleyball player reaches the highest point of his jump or the moment when the volleyball player hits the ball.
[0114] If the determination result in S704 above indicates that the operation instruction is "output of modeling data," the modeling data generation unit 204 proceeds to S706. The output instruction for modeling data is given, for example, by a user operation on the output button 351 shown in FIG. 8(a). When the output button 351 is pressed by a user operation, the time codes of the virtual viewpoint images obtained by capturing images using all the virtual cameras specified up to the previous step, the positions and The attitude is confirmed, and the process proceeds to S706.
[0115] In steps S706 to S710, the modeling data generation unit 204 executes processing for all virtual cameras specified in step S701.
[0116] In S707, the modeling data generation unit 204 determines the rendering range of the virtual camera using the time code of the virtual viewpoint image received in the process of S702 and the viewpoint information of the virtual camera, including the position and orientation of the virtual camera, received in the process of S703. The method for determining the rendering range of the virtual camera is the same as the method for determining the rendering range of the virtual camera described in S404, and therefore, the description thereof will be omitted.
[0117] In S708, the modeling data generation unit 204 determines the generation range of the modeling data corresponding to the virtual camera based on the drawing range of the virtual camera determined in S707. The method for determining the generation range of the modeling data is the same as the method for determining the generation range of the modeling data described in S405, and therefore, the description thereof will be omitted.
[0118] In S709, the modeling data generation unit 204 acquires 3D models of objects included in the generation range of the modeling data of the virtual camera determined in S708 from the database 103. The method for acquiring the 3D models is the same as the method for acquiring the 3D models described in S406, and therefore, the description thereof will be omitted.
[0119] In S711, the modeling data generation unit 204 synthesizes 3D models included in the generation range of the modeling data determined for each virtual camera acquired up to the previous step, based on the priority of each object.
[0120] The priority of each object is a parameter set during the operation of the virtual camera in S702 or S703, and is referred to when generating modeling data. The priority of each object will be described with reference to FIG. 8(c). As shown in FIG. 8(c), for example, a long tap on an object is performed on the operation screen 320 to display a priority setting screen 820. The user selects a priority from among the selection items (e.g., high, medium, low) displayed on the priority setting screen 820, and the selected priority is set. Note that the number of priority selection items is not limited to three. If an initial value of the priority of each object is set in advance and no priority is set on the setting screen 820, the initial value may be set as the priority of each object.
[0121] Priority-based compositing generates modeling data that highlights high-priority objects, for example, by reducing the color density of medium- or low-priority objects compared to high-priority objects, thereby highlighting the high-priority objects. Multiple objects composited in this manner are described with reference to FIG. 8(d). Details of FIG. 8(d) will be described later. This example shows six 3D models 831-836 of objects composited using a virtual camera capturing six virtual viewpoint images with different time codes, and output as modeling data. In FIG. 8(d), 3D models 833 and 834 are set to a "high" priority and are composited in normal colors, while 3D models 831, 832, 835, and 836 are set to a "low" priority and are composited in pale colors. Using such priority settings, modeling data can be generated that highlights objects with noteworthy time codes.
[0122] In addition, when multiple objects overlap, the modeling data may be generated by combining objects with a "high" priority so that the objects with a "low" priority are in front of each other.
[0123] In S712, the modeling data generation unit 204 generates a 3D model that does not actually exist but is used to support the 3D model acquired from the database. The method for generating the 3D model of the support part is the same as the method for generating the 3D model of the support part described in S407, and therefore, the description thereof will be omitted.
[0124] In S713, the modeling data generation unit 204 generates modeling data by synthesizing the set of 3D models generated in the processing up to S712, which is the previous step. An example of the synthesized modeling data will be described with reference to FIG. 8(d).
[0125] 8(d), the modeling data is generated by retrieving 3D models included in the generation range of each modeling data from database 103 using six virtual viewpoint images with different time codes, positions, and postures, and a virtual camera, and then synthesizing them. Specifically, the modeling data includes 3D models 831-836 of people, 3D models 843 and 844 of supports, and a 3D model 850 of a pedestal.
[0126] The coordinates of the original 3D models of the foreground, such as people, and the 3D models of the background, such as the field, are expressed by a single three-dimensional coordinate system as shown in Figure 3(a), and the relative positions of these accurately reproduce the positions and postures of the players on the actual field.
[0127] Furthermore, even if the 3D model is of the same person (object), by using virtual viewpoint images captured by multiple virtual cameras with different time codes, the following data can be output: In other words, by calculating the generation range of modeling data based on these drawing areas, a series of plays by a sports player can be output as modeling data arranged in chronological order.
[0128] It should be noted that the number of target objects is not limited to one, and as shown in embodiment 1, it is of course possible to specify the time codes of different virtual viewpoint images, the position of the virtual camera, and the posture of the virtual camera for multiple objects.
[0129] As described above, according to this embodiment, it is possible to generate modeling data of objects according to multiple pieces of designation information specified by user operations. That is, it is possible to generate modeling data of multiple objects based on the drawing ranges of multiple virtual cameras according to the designation information.
[0130] For example, a series of plays (continuous movements) of a professional athlete can be synthesized into a single 3D model figure and output as modeling data. For example, modeling data can be output that shows a volleyball player spiking or a figure skater jumping in chronological order.
[0131] [Other embodiments] In the above-described embodiment, an example has been described in which modeling data is generated from a virtual viewpoint image generated using a 3D model generated based on a plurality of captured images. However, the present embodiment is not limited to this, and can also be applied to a moving image using a 3D model generated using computer graphics (CG) software, for example.
[0132] The present disclosure can also be realized by providing a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions. [Explanation of symbols]
[0133] 201 Virtual camera control unit 204 Modeling data generation unit
Claims
1. an object for which modeling data is to be generated, and a determination means for determining a priority of the object; an acquisition means for acquiring a 3D model of the object generated based on a plurality of captured images; a generating means for generating modeling data including the 3D model based on the 3D model and the priority; An information processing device comprising:
2. The modeling data includes the 3D model corrected according to the priority.
2. The information processing apparatus according to claim 1, wherein:
3. The modeling data includes the 3D model whose color information has been corrected according to the priority.
3. The information processing apparatus according to claim 2, wherein:
4. The object and the priority are determined based on a user operation.
2. The information processing apparatus according to claim 1, wherein:
5. an acquisition means for acquiring information indicating a position and an attitude of the virtual camera and a time based on a user operation; The object is determined based on the information.
2. The information processing apparatus according to claim 1, wherein:
6. The modeling data includes a support portion that supports the shape of the object.
6. The information processing apparatus according to claim 1, wherein the information processing apparatus is a computer.
7. The modeling data is data for modeling a model figure or a relief.
7. The information processing apparatus according to claim 1, wherein the information processing apparatus is a computer.
8. a determination step of determining an object for which modeling data is to be generated and a priority for the object; an acquisition step of acquiring a 3D model of the object generated based on a plurality of captured images; a generating step of generating modeling data including the 3D model based on the 3D model and the priority; An information processing method comprising:
9. A program for causing a computer to function as the information processing device according to any one of claims 1 to 7.
Citation Information
Patent Citations
Methods and apparati for implementing programmable pipeline for three-dimensional printing including multi-material applications
US20140324204A1
3D Printing Using 3D Video Data
US20180046167A1
Doll molding system, information processing method, and program
JP2020062322A