Information processing apparatus, system including the same, information processing method, and program

The information processing apparatus addresses the limitation of previous technologies by enabling the easy modeling of a three-dimensional object from an arbitrary scene of a moving image, using specifying, acquiring, and generating means to create modeling data from shape data.

JP7690301B2Active Publication Date: 2025-06-10CANON KK
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2021030905
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-02-26
Publication Date
2025-06-10
Estimated Expiration
2041-02-26

AI Technical Summary

Technical Problem

Existing methods, such as those described in Patent Document 1, are limited in their ability to model a doll of an object from scenes not included in a list of highlight video scenes.

Method used

An information processing apparatus that includes specifying means for selecting an object in a moving image, acquiring means for obtaining shape data representing the three-dimensional shape of the object, and generating means for creating modeling data based on the acquired shape data, allowing for easy modeling of a three-dimensional object from an arbitrary scene.

Benefits of technology

Enables the easy modeling of a solid object from an arbitrary scene of a moving image, overcoming the limitations of previous technologies that were restricted to modeling from specific highlight scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007690301000001
    Figure 0007690301000001
  • Figure 0007690301000002
    Figure 0007690301000002
  • Figure 0007690301000003
    Figure 0007690301000003
Patent Text Reader

Abstract

To allow for creating a three-dimensional representation of an object in any scene of a video image.SOLUTION: An information processing device provided herein is configured to: identify information for specifying a target for generating formation data for a three-dimensional representation in a video image generated using shape data representing a three-dimensional shape of an object; acquire shape data of the object in the video image corresponding to a specified time; and output formation data of a three-dimensional representation to be formed according to the acquired shape data.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a technique for generating object modeling data from a moving image.

Background Art

[0002] In recent years, it has become possible to model a figure of an object based on a three-dimensional model (hereinafter referred to as a 3D model), which is data representing the three-dimensional shape of the object, using a modeling device such as a 3D printer. The objects of interest include not only characters appearing in games and animations but also real people. By inputting a 3D model obtained by imaging or scanning a real person or the like into a 3D printer, a figure can be modeled at about one-tenth the size of the person or the like.

[0003] Patent Document 1 discloses a method of modeling a doll of a desired object by allowing a user to select a desired scene or an object included in the scene from a list of highlight video scenes created by imaging a sports game.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in Patent Document 1, it is only possible to select an object to be modeled from the highlight scenes in the list, and it is difficult to model a doll of an object included in a scene not in the list.

[0006] An object of the present disclosure is to enable easy modeling of a three-dimensional object of an object in an arbitrary scene of a moving image.

Means for Solving the Problems

[0007] An information processing apparatus according to one aspect of the present disclosure includes specifying means for specifying an object that is a target for generating modeling data in a moving image generated using shape data representing the three-dimensional shape of the object, and the object specified by the specifying means The above-mentioned acquiring means for acquiring the shape data of the object, and the shape data acquired by the acquiring means , a plurality of the shape data including generating means for generating modeling data, and is characterized by having the same.

Effect of the Invention

[0008] According to the present disclosure, a solid object of an object in an arbitrary scene of a moving image can be easily modeled.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Embodiment for Carrying Out the Invention

[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the following embodiments do not limit the present disclosure according to the claims, and not all combinations of the features described in the present embodiments are essential for the solution means of the present disclosure. The same reference numerals are assigned to the same components, and the description thereof will be omitted.

[0011] [Embodiment 1] In the present embodiment, a mode of generating modeling data will be described in a system that draws a virtual viewpoint image using three-dimensional shape data (hereinafter referred to as a 3D model) representing the three-dimensional shape of an object, which is for generating a virtual viewpoint image and obtained from captured images of a plurality of viewpoints. Note that a virtual viewpoint image is an image generated by an end user and / or a selected operator specifying the position and line-of-sight direction of the virtual viewpoint, and is also called a free viewpoint image, an arbitrary viewpoint image, or the like. The virtual viewpoint image may be a moving image or a still image, but in the present embodiment, the case of a moving image will be described as an example. In the following description, originally, the virtual viewpoint will be described by replacing it with a virtual camera. In the following description, the position of the virtual viewpoint corresponds to the position of the virtual camera, and the line-of-sight direction from the virtual viewpoint corresponds to the attitude of the virtual camera.

[0012] (System Configuration) FIG. 1 is a diagram showing a configuration example of an information processing system (virtual viewpoint image generation system) that generates modeling data of an object to model a three-dimensional object in a virtual viewpoint image. FIG. 1(a) shows a configuration example of the information processing system 100, and FIG. 1(b) shows an installation example of a sensor system included in the information processing system. The information processing system 100 includes n sensor systems 101a - 101n, an image recording device 102, a database 103, an image generation device 104, and a tablet 105. Each of the sensor systems 101a - 101n has at least one camera as an imaging device. In the following, unless otherwise specified, the n sensor systems from the sensor system 101a to the sensor system 101n will not be distinguished, and will be collectively referred to as a plurality of sensor systems 101.

[0013] An installation example of the multi-sensor system 101 and the virtual camera will be described with reference to FIG. 1(b). As shown in FIG. 1(b), the multi-sensor system 101 is installed so as to surround the region 120 which is the target region for imaging, and the cameras of the multi-sensor system 101 image the region 120 from different directions. The virtual camera 110 images the region 120 from a direction different from that of the cameras of the multi-sensor system 101. Details of the virtual camera 110 will be described later.

[0014] When the imaging target is a professional sports game such as rugby or soccer, the region 120 is the stadium field (ground), and n (for example, 100) multi-sensor systems 101 are installed so as to surround the field. Also, in the imaging target region 120, there may be not only people on the field but also balls and other objects. Note that the imaging target is not limited to the stadium field, and it may be a music live held in an arena or the like, or a CM shooting held in a studio, as long as the multi-sensor system 101 can be installed. Note that the number of sensor systems 101 to be installed is not limited. Also, the multi-sensor system 101 does not have to be installed over the entire circumference of the region 120, and depending on restrictions on the installation location and the like, it may be installed only in a part of the periphery of the region 120. Also, the plurality of cameras included in the multi-sensor system 101 may include imaging devices with different functions such as a telephoto camera and a wide-angle camera.

[0015] The plurality of cameras included in the multi-sensor system 101 perform imaging synchronously and acquire a plurality of images. Note that each of the plurality of images may be an imaged image, or may be an image obtained by performing image processing such as a process of extracting a predetermined region from the imaged image.

[0016] Note that each of the sensor systems 101a - 101n may have a microphone (not shown) in addition to the camera. The microphones of the plurality of sensor systems 101 each synchronously collect sound. Based on the collected sound, an acoustic signal can be generated to be reproduced together with the display of the image in the image generation device 104. Hereinafter, for simplicity of explanation, the description of the sound will be omitted, but basically both the image and the sound are processed together.

[0017] The image recording device 102 acquires a plurality of images from the plurality of sensor systems 101, and stores the acquired plurality of images together with the time code used for imaging in the database 103. The time code is time information represented by an absolute value for uniquely identifying the imaging time, and is, for example, time information that can be specified in a format such as day:hour:minute:second.frame number.

[0018] The database 103 manages event information, 3D model information, and the like. The event information includes data indicating the storage destination of the 3D model information for each object associated with the time codes of all events of the imaging target. The object may include a person or thing that the user wants to model, or a person or thing that is not a modeling target. The 3D model information includes information regarding the 3D model of the object.

[0019] The image generation device 104 receives, as inputs, an image corresponding to the time code from the database 103 and information regarding the virtual camera 110 set by a user operation from the tablet 105. The virtual camera 110 is set in a virtual space associated with the region 120, and can view the region 120 from a viewpoint different from that of any of the cameras of the plurality of sensor systems 101. Details of the virtual camera 110, its operation method, and its operation will be described later with reference to the drawings.

[0020] The image generation device 104 generates a 3D model for virtual viewpoint image generation based on the images for each time code acquired from the database 103, and generates a virtual viewpoint image using the generated 3D model for virtual viewpoint image generation and the information regarding the viewpoint of the virtual camera. The viewpoint information of the virtual camera includes information indicating the position and orientation of the virtual viewpoint. Specifically, the viewpoint information includes parameters representing the three-dimensional position of the virtual viewpoint and parameters representing the orientation of the virtual viewpoint in the pan direction, tilt direction, and roll direction. Note that the virtual viewpoint image is an image representing what is seen from the virtual camera 110 and is also called a free viewpoint video. The virtual viewpoint image generated by the image generation device 104 is displayed on a touch panel such as the tablet 105.

[0021] In the present embodiment, the image generation device 104 generates modeling data (object shape data) used by the modeling device 106 from the 3D model corresponding to the object existing in the drawing range based on the drawing range of at least one virtual camera. That is, the image generation device 104 sets conditions for specifying an object to be modeled for the virtual viewpoint image, acquires shape data representing the three-dimensional shape of the object based on the set conditions, and generates modeling data based on the acquired shape data. Details of the generation process of the modeling data will be described later with reference to the drawings. Note that the format of the modeling data generated by the image generation device 104 may be any format that can be processed by the modeling device 106, such as a general polygon mesh format for handling 3D models or the unique format of the modeling device 106. That is, when the modeling device 106 can process the shape data for virtual viewpoint image generation, the image generation device 104 specifies the shape data for virtual viewpoint image generation as the modeling data. On the other hand, when the modeling device 106 cannot process the shape data for virtual viewpoint image generation, the image generation device 104 generates modeling data using the shape data for virtual viewpoint image generation.

[0022] The tablet 105 is a portable device having a touch panel that has the functions of both a display unit for displaying images and an input unit for receiving user operations. The tablet 105 may be a portable device equipped with other functions such as a smartphone. The tablet 105 receives a user operation for setting information regarding the virtual camera. Further, the tablet 105 displays the virtual viewpoint image generated by the image generation device 104 or the virtual viewpoint image stored in the database 103. Then, the tablet 105 receives a user operation for setting conditions for specifying an object to be modeled with respect to the virtual viewpoint image. By this user operation, an arbitrary spatial range and time range for generating the modeling data are set. Details of this operation method will be described later with reference to the drawings. Note that the operation of the virtual camera is not limited to a user operation on the touch panel of the tablet 105, and may be a user operation on an operation device such as a three-axis controller.

[0023] As shown in FIG. 1, the image generation device 104 and the tablet 105 may be configured as separate devices or as an integrated unit. In the case of an integrated unit, the image generation device 104 has a touch panel or the like, receives an operation of the virtual camera, and displays the virtual viewpoint image generated by the image generation device 104 on the touch panel. Note that the virtual viewpoint image may be displayed on the liquid crystal screen of a device other than the tablet 105 or the image generation device 104.

[0024] The modeling device 106 is, for example, a 3D printer or the like, and takes the modeling data generated by the image generation device 104 as an input, and models a three-dimensional object of a corresponding object such as a doll (3D model figure) or a relief. Note that the modeling method of the modeling device 106 is not limited to stereolithography, inkjet, powder adhesion method, etc., and any method may be used as long as it can model a three-dimensional object. Note that the modeling device 106 is not limited to a device that outputs a three-dimensional object such as a doll, and may be a device that prints on a plate or paper.

[0025] As shown in FIG. 1, the image generation device 104 and the modeling device 106 may be configured as separate devices or as an integrated unit.

[0026] Note that the configuration of the information processing system 100 is not limited to the one in which the tablet 105 and the modeling device 106 are connected to the image generation device 104 on a one-to-one basis as shown in FIG. 1(a). For example, an information processing system in which a plurality of tablets 105 and a plurality of modeling devices 106 are connected to the image generation device 104 may be used.

[0027] (Configuration of the Image Generation Device) A configuration example of the image generation device 104 will be described with reference to the drawings. FIG. 2 is a diagram showing a configuration example of the image generation device 104. FIG. 2(a) shows a functional configuration example of the image generation device 104, and FIG. 2(b) shows a hardware configuration example of the image generation device 104.

[0028] As shown in FIG. 2(a), the image generation device 104 includes a virtual camera control unit 201, a 3D model generation unit 202, an image generation unit 203, and a modeling data generation unit 204. The image generation device 104 uses the above-described functional units to generate modeling data based on the drawing range of at least one virtual camera. Here, an overview of each function will be described, and details of the processing will be described later.

[0029] The virtual camera control unit 201 receives operation information of the virtual camera from the tablet 105 or the like. The operation information of the virtual camera includes at least the position and orientation of the virtual camera and the time code. Details of these operation information of the virtual camera will be described later with reference to the drawings. Note that when the image generation device 104 has a touch panel or the like and is configured to be able to receive operation information of the virtual camera, the virtual camera control unit 201 receives the operation information of the virtual camera from the image generation device 104.

[0030] Based on a plurality of captured images, the 3D model generation unit 202 generates a 3D model representing the three-dimensional shape of the object within the area 120. Specifically, the 3D model generation unit 202 obtains, from the images, a foreground image obtained by extracting a foreground region corresponding to an object such as a person or a ball, and a background image obtained by extracting a background region other than the foreground region. Then, the 3D model generation unit 202 generates a foreground 3D model for each object based on the plurality of foreground images.

[0031] These 3D models are generated by a shape estimation method such as Visual Hull and are composed of point clouds. However, the data format of the 3D model representing the shape for each object is not limited to this. Note that the background 3D model may be obtained in advance by an external device.

[0032] The 3D model generation unit 202 records the generated 3D model in the database 103 together with the time code.

[0033] Note that the 3D model generation unit 202 may be a configuration included in the image recording device 102 instead of the image generation device 104. In that case, the image generation device 104 only needs to read out the 3D model generated by the image recording device 102 from the database 103 via the 3D model generation unit 202.

[0034] The image generation unit 203 obtains a 3D model from the database 103 and generates a virtual viewpoint image based on the obtained 3D model. Specifically, for each point constituting the 3D model, the image generation unit 203 obtains an appropriate pixel value from the image and performs a coloring process. Then, the image generation unit 203 arranges the colored 3D model in a three-dimensional virtual space, projects it onto a virtual camera (virtual viewpoint), and renders it to generate a virtual viewpoint image.

[0035] However, the method for generating the virtual viewpoint image is not limited to this, and various methods such as a method of generating a virtual viewpoint image by projective transformation of the captured image without using a 3D model may be used.

[0036] The modeling data generation unit 204 calculates and determines the drawing range or projection range of the virtual camera using the position and orientation of the virtual camera and the time code. Then, the modeling data generation unit 204 determines the generation range of the modeling data based on the determined drawing range of the virtual camera, and generates the modeling data from the 3D models included in that range. Details of these processes will be described later with reference to the figures.

[0037] (Hardware Configuration of the Image Generation Device) Next, the hardware configuration of the image generation device 104 will be described with reference to FIG. 2(b). As shown in FIG. 2(b), the image generation device 104 includes a CPU 211, a RAM 212, a ROM 213, an operation input unit 214, a display unit 215, and a communication I / F (interface) unit 216.

[0038] The CPU (Central Processing Unit) 211 performs processing using programs and data stored in the RAM (Random Access Memory) 212 and the ROM (Read Only Memory) 213.

[0039] The CPU 211 controls the overall operation of the image generation device 104 and executes processes for realizing each function shown in FIG. 2(a). Note that the image generation device 104 may have one or more dedicated hardware different from the CPU 211, and at least a part of the processes by the CPU 211 may be executed by the dedicated hardware. Examples of the dedicated hardware include an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), and a DSP (Digital Signal Processor).

[0040] The ROM 213 holds programs and data. The RAM 212 has a work area for temporarily storing programs and data read from the ROM 213. Also, the RAM 212 provides a work area used when the CPU 211 executes each process.

[0041] The operation input unit 214 is, for example, a touch panel, which receives input operations by the user and acquires information input by the received user operations. Examples of the input information include information regarding a virtual camera and information regarding the time code of a virtual viewpoint image to be generated. Note that the operation input unit 214 may be connected to an external controller and receive input information from the user regarding operations. The external controller is, for example, a three-axis controller such as a joystick or an operation device such as a mouse. Note that the external controller is not limited to these.

[0042] The display unit 215 is a touch panel or a screen, which displays the generated virtual viewpoint image. In the case of a touch panel, the operation input unit 214 and the display unit 215 are integrated.

[0043] The communication I / F unit 216 performs transmission and reception of information with the database 103, the tablet 105, the modeling device 106, etc. via, for example, a LAN. Further, the communication I / F unit 216 may transmit information to an external screen via an image output port corresponding to the following communication standards. Examples of the image output port corresponding to the communication standards include HDMI (registered trademark) (High-Definition Multimedia Interface) and SDI (Serial Digital Interface). Further, the communication I / F unit 216 may transmit image data and modeling data via, for example, Ethernet.

[0044] (Virtual camera (virtual viewpoint) and its operation screen) Subsequently, taking as an example the case where a rugby game at a rugby field is imaged to acquire an imaged image, an operation screen for setting a virtual camera and its settings will be described. FIG. 3 is a diagram showing a virtual camera and its operation screen. FIG. 3(a) shows a coordinate system, FIG. 3(b) shows an example of a field to which the coordinate system of FIG. 3(a) is applied, FIGS. 3(c) and 3(d) show examples of the drawing range of the virtual camera, and FIG. 3(e) shows an example of the movement of the virtual camera. FIG. 3(f) shows an example of the display of a virtual viewpoint image seen from the virtual camera.

[0045] First, a coordinate system representing the three-dimensional space of the imaging target, which serves as a reference when setting the virtual viewpoint, will be described. As shown in FIG. 3(a), in this embodiment, a rectangular coordinate system representing the three-dimensional space with three axes of the X-axis, Y-axis, and Z-axis is used. This rectangular coordinate system is set for each object shown in FIG. 3(b), that is, the rugby field 391, the ball 392 existing thereon, the player 393, and the like. Further, it may be set for facilities (structures) in the rugby field such as the spectator seats and billboards around the field 391. Specifically, the origin (0, 0, 0) is set at the center of the field 391. Then, the X-axis is set in the long side direction of the field 391, the Y-axis is set in the short side direction of the field 391, and the Z-axis is set in the vertical direction with respect to the field 391. Note that the directions of the respective axes are not limited to these. Using such a coordinate system, the position and orientation of the virtual camera 110 are specified.

[0046] Subsequently, the drawing range of the virtual camera will be described with reference to the drawings. In the quadrangular pyramid 300 shown in FIG. 3(c), the vertex 301 represents the position of the virtual camera 110, and the vector 302 in the line-of-sight direction with the vertex 301 as the base point represents the orientation of the virtual camera 110. Note that the vector 302 is also called the optical axis vector of the virtual camera. The position of the virtual camera is represented by the components (x, y, z) of each axis, and the orientation of the virtual camera 110 is represented by a unit vector with the components of each axis as scalars. The vector 302 representing the orientation of the virtual camera 110 is assumed to pass through the center points of the front clip plane 303 and the rear clip plane 304. The frustum of the virtual camera, which is the projection range (drawing range) of the 3D model, is the space 305 sandwiched between the front clip plane 303 and the rear clip plane 304.

[0047] Next, the components indicating the drawing range of the virtual camera will be described with reference to the figures. Fig. 3(d) is a view of the virtual viewpoint in Fig. 3(c) as seen from above (the Z-axis). The drawing range is determined by the following values. Each value of the distance 311 from the vertex 301 to the front clip plane 303, the distance 312 from the vertex 301 to the rear clip plane 304, and the viewing angle 313 of the virtual camera 110 may be a preset predetermined value (specified value), or a set value in which the predetermined value is changed by a user operation. Also, the viewing angle 313 may separately be a value obtained based on a variable focal length of the virtual camera 110, with this variable as a basis. Note that the relationship between the viewing angle and the focal length is a general technique and its explanation will be omitted.

[0048] Next, the change in the position of the virtual camera 110 (movement of the virtual viewpoint) and the change in the orientation of the virtual camera 110 (rotation) will be described. The virtual viewpoint can be moved and rotated within a space represented in three-dimensional coordinates. Fig. 3(e) is a diagram for explaining the movement and rotation of the virtual camera. In Fig. 3(e), the dashed arrow 306 represents the movement of the virtual camera (virtual viewpoint), and the dashed arrow 307 represents the rotation of the moved virtual camera (virtual viewpoint). The movement of the virtual camera is represented by the components (x, y, z) of each axis, and the rotation of the virtual camera is represented by yaw, which is rotation around the Z-axis, pitch, which is rotation around the X-axis, and roll, which is rotation around the Y-axis. Since the virtual camera can thus freely move and rotate in the three-dimensional space of the subject (field), a virtual viewpoint image in which an arbitrary area of the subject becomes the drawing range can be generated.

[0049] Next, an operation screen for setting the position and orientation of the virtual camera (virtual viewpoint) will be described. Fig. 3(f) is a diagram for explaining an example of an operation screen for the virtual camera (virtual viewpoint).

[0050] In this embodiment, since the generation range of the modeling data is determined based on the drawing range of at least one virtual camera, the operation screen 320 of the virtual camera shown in FIG. 3(f) can also be said to be an operation screen for determining the generation range of the modeling data. In FIG. 3(f), the operation screen 320 of the virtual camera is displayed on the touch panel of the tablet 105. Note that the display destination of the operation screen 320 of the virtual camera is not limited to this, and it may be a touch panel of the image generation device 104 or the like.

[0051] On the operation screen 320, the drawing range of the virtual camera (the drawing range related to the imaging of the virtual viewpoint image) is displayed as a virtual viewpoint image in accordance with the screen frame of the operation screen 320. With such a display, the user can set the conditions for specifying the object to be modeled while visually observing them.

[0052] The operation screen 320 has a virtual camera operation area 322 that receives a user operation for setting the position and orientation of the virtual camera 110, and a time code operation area 323 that receives a user operation for setting a time code. First, the virtual camera operation area 322 will be described. Since the operation screen 320 is displayed on the touch panel, in the virtual camera operation area 322, general touch operations 325 such as taps, swipes, pinch-ins, and pinch-outs are received as user operations. By this touch operation 325, the position and focal length (angle of view) of the virtual camera are adjusted. Also, in the virtual camera operation area 322, touch operations 324 such as continuously pressing each axis of the orthogonal coordinate system are received as user operations. By this touch operation 324, the virtual camera 110 rotates around the X-axis, Y-axis, or Z-axis, and the orientation of the virtual camera is adjusted. By assigning movement, rotation, and scaling of the virtual camera to each touch operation on the operation screen 320 in this way, the virtual camera 110 can be freely operated. These operation methods are well-known and their descriptions are omitted.

[0053] Note that the operations related to the position and orientation of the virtual camera are not limited to touch operations on the touch panel, and operations using an operation device such as a joystick may also be used.

[0054] Next, the time code operation area 323 will be described. The time code operation area 323 has a main slider 332 and a knob 342, and a sub-slider 333 and a knob 343, and has a plurality of elements for operating the time code. The time code operation area 323 has an additional button 350 for a virtual camera and an output button 351.

[0055] The main slider 332 is an input element capable of performing an operation of setting to a desired time code among all the time codes of the imaging data. When the position of the knob 342 is moved to a desired position by a drag operation or the like, the time code corresponding to the position of the knob 342 is specified. That is, by adjusting the position of the knob 342 of the main slider 332, an arbitrary time code is specified.

[0056] The sub-slider 333 is an input element capable of performing an operation of enlarging and displaying a part of all the time codes and setting the time code in detail for the enlarged part. When the position of the knob 343 is moved to a desired position by a drag operation or the like, the time code corresponding to the position of the knob 343 is specified. The main slider 332 and the sub-slider 333 have the same length on the screen, but the widths of the selectable time codes are different. For example, the main slider 332 can be selected from 3 hours, which is the length of one game, while the sub-slider 333 can be selected from 30 seconds, which is a part of it. In this way, the scales of the sliders are different, and the sub-slider can specify a more detailed time code such as seconds or frames.

[0057] In addition, the time code specified using the grip 342 of the main slider 332 or the grip 343 of the sub-slider 333 may be displayed as a numerical value in the format of day:hour:minute:second.frame number. The sub-slider 333 may be always displayed on the operation screen 320 or may be temporarily displayed. For example, it may be displayed when an instruction to display the time code is received, or when an instruction for a specific operation such as a pause is received. The time code section selectable by the sub-slider 333 may be variable.

[0058] According to the position and orientation of the virtual camera set by the above operations and the time code, the generated virtual viewpoint image is displayed in the operation area 322 of the virtual camera. In FIG. 3(f), as an example, the subject is a rugby game, and a decisive pass scene leading to a score is displayed. Although details will be described later, this decisive scene will be generated or output as shape data of an object for modeling.

[0059] FIG. 3(f) shows a case where the time code at the moment of releasing the ball is specified. However, for example, by operating the sub-slider 333 in the time code operation area 323, the time code when the ball is in the air can also be easily specified.

[0060] In FIG. 3(f), the position and orientation of the virtual camera are specified so that three players exist in the drawing range, but it is not limited to this. For example, by performing touch operations 324 and 325 on the virtual camera operation area 322, the space (the drawing range of the virtual camera) 305 can be freely operated, and further, the operation may be performed so that the surrounding players exist in the drawing range.

[0061] The output button 351 is a button that is operated when determining the drawing range of the virtual camera set by user operations on the virtual camera operation area 322 and the time code operation area 323, and outputting the shape data of the object for modeling. When the time code indicating the time range and the position and its orientation indicating the space range are determined for the virtual camera, the drawing range of the virtual camera corresponding to these determinations is calculated. Then, based on the drawing range of the virtual camera which is the calculation result, the generation range of the modeling data is determined. The details of this process will be described later with reference to the figures.

[0062] The additional button 350 for the virtual camera is a button used when using a plurality of virtual cameras for generating the modeling data. The details of the generation process of the modeling data using a plurality of virtual cameras will be described in Embodiment 2, so the description is omitted here.

[0063] Note that the operation screen of the virtual camera is not limited to the operation screen 320 shown in FIG. 3(f), and any operation for setting the position and its orientation and the time code of the virtual camera is acceptable. For example, the virtual camera operation area 322 and the time code operation area 323 do not have to be separated. For example, when an operation such as a double-tap is performed on the virtual camera operation area 322, a pause or the like may be processed as an operation for setting the time code.

[0064] Also, although the case where the operation input unit 214 is the tablet 105 has been described, the operation input unit 214 is not limited to this, and it may be an operation device having a general display and a three-axis controller or the like.

[0065] (Generation Process of Modeling Data) Next, the generation process of the modeling data in the image generation device 104 will be described with reference to the drawings. FIG. 4 is a flowchart showing the flow of the generation process of the modeling data. This series of processes is realized by the CPU 211 executing a predetermined program to operate each functional unit shown in FIG. 2(a). Hereinafter, steps will be denoted as "S". The same applies to the following description. It should be noted that the description will be made on the premise that the shape data representing the three-dimensional shape of the object used for generating the virtual viewpoint image is in a different format from the modeling data.

[0066] In S401, the modeling data generation unit 204 receives the specification of the time code of the virtual viewpoint image via the virtual camera control unit 201. The method of specifying the time code may be, for example, the method using the slider shown in FIG. 3(f), or the method of directly inputting numbers.

[0067] In S402, the modeling data generation unit 204 receives the operation information of the virtual camera via the virtual camera control unit 201. The operation information of the virtual camera includes at least information regarding the position and orientation of the virtual camera. The method of operating the virtual camera may be, for example, the method using the tablet shown in FIG. 3(f), or the method using an operating device such as a joystick.

[0068] In S403, the modeling data generation unit 204 determines, via the virtual camera control unit 201, whether an output instruction has been received, that is, whether the output button 351 has been operated. When the modeling data generation unit 204 obtains a determination result that an output instruction has been received (YES in S403), the process proceeds to S404. When the modeling data generation unit 204 obtains a determination result that no output instruction has been received (NO in S403), the process returns to S401, and the processes of S401 and S402 are executed again. That is, until an output instruction is received, the process of receiving user operations of the time code of the virtual viewpoint image and the position and orientation of the virtual camera continues.

[0069] In S404, the modeling data generation unit 204 determines the drawing range of the virtual camera using the time code of the virtual viewpoint image received in the process of S401 and the viewpoint information of the virtual camera including the position and orientation of the virtual camera received in the process of S402. As for the method of determining the drawing range of the virtual camera, for example, the front clipping plane 303 and the rear clipping plane 304 shown in FIGS. 3(c) and 3(d) may be determined by user operation, or may be determined based on the calculation result by an arithmetic formula set in the device in advance.

[0070] In S405, the modeling data generation unit 204 determines the generation range of the modeling data corresponding to this virtual camera based on the drawing range of the virtual camera determined in S404. The method of determining the generation range of the modeling data will be described. FIG. 5 is a diagram showing the generation range and generation example of the modeling data. FIG. 5(a) shows the generation range of the modeling data in the three-dimensional space, and FIG. 5(b) shows the yz plane of the generation range of the modeling data shown in FIG. 5(a). FIG. 5(c) shows an example of a 3D model within the generation range of the modeling data in the three-dimensional space. FIG. 5(d) shows the yz plane when an auditorium exists within the generation range of the modeling data. FIG. 5(e) shows an example of a 3D model figure corresponding to the 3D model shown in FIG. 5(c). FIG. 5(f) shows another example of the 3D model figure corresponding to the 3D model shown in FIG. 5(c).

[0071] In FIG. 5(a), a virtual camera (position 301, space (frustum) 305, etc.) is displayed in the three-dimensional space, and a plane (e.g., the field of a stadium) 500 with Z = 0 of the imaging target is displayed in the same space. As shown in FIG. 5(a), the generation range 510 of the modeling data is determined based on the plane (bottom surface) 501 included in the space (frustum) 305, which is the drawing range of the virtual camera, within the plane 500 with Z = 0.

[0072] The generation range 510 of the shaping data is a part of the space (frustum) 305 which is the drawing range of the virtual camera. The generation range 510 of the shaping data is a space surrounded by the plane 501, the front plane (front face) 513 located near the front clip plane 303, and the rear plane (rear face) 514 located near the rear clip plane 304. The sides of the space that becomes the generation range 510 of the shaping data are set based on the position 301 of the virtual camera, its field angle, and the plane 500. Also, the upper part of the space that becomes the generation range 510 of the shaping data is set based on the position 301 of the virtual camera and its field angle.

[0073] The planes 513 and 514 will be described with reference to FIG. 5(b). FIG. 5(b) is a simplified view from the side of the virtual camera and the plane of Z = 0 in FIG. 5(a). The plane 513 is a plane that passes through the intersection points 503 or 505, is located at a predetermined distance from the front clip plane 303 within the space 305, and is parallel to the front clip plane 303. The plane 514 is a plane that passes through the intersection points 504 or 506, is located at a predetermined distance from the rear clip plane 304 within the space 305, and is parallel to the rear clip plane 304. The predetermined distance is set in advance. The plane 501 is a rectangle with the above-mentioned intersection points 503 to 506 as vertices.

[0074] Note that the plane 514 may be the same as the rear clip plane 304. Note that the planes 513 and 514 do not have to be parallel to the front clip plane and the rear clip plane, respectively. For example, it may be a plane that passes through the vertices of the plane 501 and is perpendicular to the plane 501.

[0075] In addition to the above, the generation range 510 of the modeling data may be determined in consideration of the spectator seats in the stadium or the like. This example will be described with reference to FIG. 5(d). Similar to FIG. 5(b), FIG. 5(d) shows the drawing range when the virtual camera in FIG. 5(a) and the plane of Z = 0 are viewed from the side in a simplified manner. In FIG. 5(d), the plane 514 may be determined based on the position of the spectator seat 507 in the stadium. Specifically, the plane 514 is a plane parallel to the rear clipping plane passing through the intersection of the spectator seat 507 and the plane 501. Thereby, the area behind the spectator seat 507 can be excluded from the generation range 510 of the modeling data.

[0076] What can be added to the determination conditions of the generation range of the modeling data is not limited to the spectator seats in the stadium, and may be configured such that the user manually specifies the object or the like. The plane 500 having an intersection with the frustum is not limited to Z = 0 and may have any shape. For example, it may have unevenness like an actual field. When the actual background 3D model of the field has unevenness, it may be corrected to a 3D model of a plane. It may also be configured not to have a plane or a curved surface having an intersection with the frustum. In that case, the drawing range of the virtual camera may be used as the generation range of the modeling data as it is.

[0077] In S406, the modeling data generation unit 204 acquires the 3D model of the object included in the generation range of the modeling data determined in S405 from the database 103.

[0078] An example of a table of event information and 3D model information managed by the database 103 will be described with reference to the drawings. FIG. 6 is a diagram showing a table of information managed by the database 103. FIG. 6(a) shows a table of event information, and FIG. 6(b) shows a table of 3D model information. As shown in FIG. 6(a), the table 610 of event information managed by the database 103 indicates the storage destination of 3D model information for each object for all time codes of the event to be imaged. In the table 610 of event information, for example, it is shown that the storage destination of the 3D model information with the time code "16:14:24.041" and the object being object A is "DataA100".

[0079] As shown in FIG. 6(b), the table 620 of 3D model information managed by the database 103 stores data for each item of "all point cloud coordinates", "texture", "average coordinates", "center of gravity coordinates", and "maximum / minimum coordinates". In the "all point cloud coordinates", data regarding the coordinates of each point of the point cloud constituting the 3D model is stored. In the "texture", data regarding the texture image applied to the 3D model is stored. In the "average coordinates", data regarding the coordinates of the point obtained by averaging all the coordinates of the point cloud constituting the 3D model is stored. In the "center of gravity coordinates", data regarding the coordinates of the point that becomes the center of gravity based on all the coordinates of the point cloud constituting the 3D model is stored. In the "maximum / minimum coordinates", data regarding the coordinates of the point with the maximum / minimum among the coordinates of the point cloud constituting the 3D model is stored. Note that the items of data stored in the table 620 of 3D model information are not limited to all of "all point cloud coordinates", "texture", "average coordinates", "center of gravity coordinates", and "maximum / minimum coordinates". For example, only "all point cloud coordinates" and "texture" may be used, or other items may be added to these items.

[0080] By using the information shown in FIG. 6, when a certain time code is specified, it is possible to obtain the coordinates of all point clouds for each object, the coordinates of the maximum / minimum values on each axis of the three-dimensional coordinates, etc. as the 3D model information at the specified time code.

[0081] Using FIG. 5(c), an example of obtaining a 3D model included in the generation range of the modeling data from the database 103 will be described. The modeling data generation unit 204 refers to the 3D model information for each object associated with the time code specified in S401 in the database 103, and determines for each object whether it is included in the generation range of the modeling data determined in S406.

[0082] As a determination method, for example, a method of determining whether all the point cloud coordinates included in the 3D model information for each object such as a person are included in the generation range may be used, or a method of determining whether only the average value of all the point cloud coordinates or the maximum / minimum values of each axis of the three-dimensional coordinates are included in the generation range may also be used.

[0083] FIG. 5(c) shows an example in which, as a result of the above-described determination, three 3D models are included in the generation range 510 of the modeling data. Although details will be described later, an example in which these three 3D models become the modeling data is as shown by the 3D models 531 - 533 in FIG. 5(e).

[0084] Note that the determination result as to whether each object is included in the generation range of the modeling data may be displayed on the operation screen 320 of the tablet. As the determination result, for example, for an object on the boundary of the generation range 510 of the modeling data, a warning or the like indicating that since it is on the boundary, the entire object will not be output as the modeling data may be displayed.

[0085] The determination process in S406 may be executed at any time while accepting the time code of the virtual viewpoint image in S401 and the user operations of the position and orientation of the virtual camera in S402. When such a determination process in S406 is executed at any time, a warning display for notifying that the target object is on the boundary may be performed.

[0086] The method for obtaining the 3D model of the object is not limited to the method of automatically obtaining based on the above-described determination result, and a manual obtaining method may be used. For example, on the operation screen 310 of the tablet, by receiving a user operation such as a tap operation on the object that the user wants to be the target of the modeling data, the object to be the acquisition target of the 3D model may be specified.

[0087] Also, in the process of S406, among the objects included in the generation range of the modeling data, the 3D model information of the field or the auditorium, which is the background, may be obtained from the database. In that case, the obtained 3D model of the field or the auditorium may be partially converted so as to serve as a pedestal. For example, even when the 3D model of the field has no thickness, by converting it into a 3D model with a specified height added, as shown in FIG. 5(e), the pedestal 540 may have a rectangular parallelepiped shape.

[0088] In S407, the modeling data generation unit 204 generates a 3D model for assisting the 3D model obtained from the database, although it does not actually exist. For example, the athlete in the 3D figure model 532 in the foreground of FIG. 5(e) is jumping and is not grounded on the field and exists in the air. Therefore, in order to display the 3D figure model, a support column (support part) for supporting the 3D figure model and fixing it to the pedestal is required as an auxiliary. Therefore, in S407, the modeling data generation unit 204 first determines whether an auxiliary part is necessary using the 3D model information for each object obtained from the database 103. This determination includes a determination as to whether the target 3D model exists in the air and a determination as to whether the target 3D model can stand on its own on the pedestal. In the determination as to whether the target 3D model exists in the air, for example, it may be determined using the coordinates of all the point clouds or the minimum coordinates included in the 3D model information. In the determination as to whether the target 3D model can stand on its own on the pedestal, for example, it may be determined using the coordinates of all the point clouds included in the 3D model information, or the center of gravity coordinates and the maximum / minimum coordinates.

[0089] When a determination result indicating that an auxiliary part is necessary is obtained, based on the 3D model information, the modeling data generation unit 204 generates a 3D model of the auxiliary part for fixing the target 3D figure model to the pedestal. When the target 3D model exists in the air, the modeling data generation unit 204 may generate a 3D model of the auxiliary part based on the coordinates of all the point clouds or the minimum coordinates and their surrounding coordinates. When the target 3D model cannot stand independently on the pedestal, the modeling data generation unit 204 may generate a 3D model of the auxiliary part based on the coordinates of all the point clouds or the center-of-gravity coordinates and the maximum / minimum coordinates. Note that the support column may be positioned to vertically support from the center-of-gravity coordinates included in the 3D model information for each object. Also, even for an object having a contact point in the field, a 3D model of the support column may be added.

[0090] On the other hand, when a determination result indicating that an auxiliary part is not necessary is obtained, the modeling data generation unit 204 skips generating the 3D model of the auxiliary part and proceeds to S408.

[0091] In S408, the modeling data generation unit 204 synthesizes the 3D models generated in the processes up to the previous step S407, converts the synthesized 3D models into the format of the output destination, and outputs them as the modeling data for the three-dimensional object. An example of the modeling data to be output will be described with reference to FIG. 5(e).

[0092] In FIG. 5(e), the modeling data includes foreground 3D models 531, 532, and 533 included in the generation range 510 of the modeling data, a support column 534, and a pedestal 540. The coordinates of the foreground 3D models such as people and the background 3D models such as the field that are the sources of these are represented by a single three-dimensional coordinate shown in FIG. 3(a), and their positional relationships can accurately reproduce the positional relationships and postures of the players on the actual field.

[0093] When the modeling data in FIG. 5(e) is output to the modeling device 106, the modeling device 106 can model the 3D model figure in this shape. Therefore, although FIG. 5(e) is described as the modeling data, it may be regarded as the 3D model figure output by the 3D printer.

[0094] Note that in order to reduce the risk of damage during delivery or transportation, in S408, the 3D models of the foreground and the background may be output separately. For example, as shown in FIG. 5(f), a pedestal 540 corresponding to the background field and 3D models 531, 532, 533 corresponding to the people in the foreground may be output separately, and small-shaped sub-pedestals 541, 542, 543 may be provided for the 3D models 531, 532, 533. In this case, by providing depressions 551-553 corresponding to the sizes of the sub-pedestals 541-543 in the pedestal 540, the 3D models 531, 532, 533 of the foreground can be mounted on the pedestal 540. Since the sub-pedestals 541, 542, 543 are actually non-existent 3D models, in S407, they may be added as auxiliary components for the 3D models of the foreground and the like, similar to the support columns.

[0095] The shapes of the sub-pedestals 541-543 and the depressions 551-553 may be not only quadrilaterals as shown in FIG. 5(f), but also other polygons. Also, the shapes of the sub-pedestals 541-543 and the depressions 551-553 may be different from each other like puzzle pieces. In this case, after the 3D models of the foreground and the background are output separately by the shaping device, the 3D model of the foreground can be properly mounted without misidentifying the mounting position and direction. Note that the positions of the sub-pedestals 541-543 for each object and the depressions 551-553 in the pedestal 540 may be changeable by user operation.

[0096] It is also possible to output the shaping data with information that does not actually exist embedded in the pedestal 540. For example, it is possible to output the shaping data of the pedestal 540 with the time code of the virtual camera used in the corresponding shaping data, and information regarding the position and orientation of the virtual camera embedded therein. Further, it is possible to output the shaping data of the pedestal 540 with the values of the three-dimensional coordinates embedded therein so that the generation range of the shaping data can be understood.

[0097] As described above, according to the present embodiment, it is possible to generate modeling data of an object according to the specified information specified by a user operation. That is, based on the drawing range of the virtual camera according to the specified information, it is possible to generate modeling data of an object included in the drawing range of the virtual camera.

[0098] For example, in field sports such as rugby and soccer held in a stadium, for a decisive scene related to scoring, by cutting out an arbitrary time code and an arbitrary spatial range, it is possible to generate modeling data while keeping the positional relationship and posture of the actual players as they are. Also, by outputting the data to a 3D printer, it is possible to generate a 3D model figure of the scene.

[0099] Also, when the shape data representing the three-dimensional shape of the object used for generating the virtual viewpoint image is in the same format as the modeling data of the object, in S408, the following processing will be performed. That is, the modeling data generation unit 204 outputs, as the modeling data of the object, the data obtained by adding the 3D model of the auxiliary part to the 3D model of the object generated in the processing up to the previous step S407. In this way, when the format of the shape data for generating the virtual viewpoint image is the same as the format of the modeling data, the shape data for generating the virtual viewpoint image can be output as it is as the modeling data.

[0100] [Embodiment 2] In the present embodiment, a mode of using a plurality of virtual cameras and generating modeling data based on their drawing ranges will be described.

[0101] In the present embodiment, since the configuration of the information processing system is the same as that in FIG. 1 and the configuration of the image generation device is the same as that in FIG. 2, the description thereof will be omitted, and the differences will be described. In the present embodiment, mainly, in the tablet 105 or the image generation device 104, an operation method for using a plurality of virtual cameras and a generation process of modeling data based on their plurality of drawing ranges will be described.

[0102] (Generation Process of Modeling Data) The generation process of modeling data according to this embodiment will be described with reference to the drawings. FIG. 7 is a flowchart showing the flow of the generation process of modeling data according to this embodiment. This series of processes is realized by the CPU 211 executing a predetermined program to operate each functional unit shown in FIG. 2(a). FIG. 8 is a diagram for explaining an example of an operation screen of a virtual viewpoint (virtual camera). FIG. 8(a) shows the case where the virtual camera 1 is selected, FIG. 8(b) shows the case where the virtual camera 2 is selected, FIG. 8(c) shows the case where the priority setting screen is displayed in FIG. 8(b), and FIG. 8(d) shows an example of a 3D model. Note that the operation screens shown in FIGS. 8(a) to 8(c) are screens obtained by expanding the functions of the operation screen shown in FIG. 3(f), and only the differences will be described here.

[0103] In S701, the modeling data generation unit 204 designates a virtual camera that accepts operations. In this embodiment, since a plurality of virtual cameras are handled, in S702 and S703 following the process of S701, a designation is made to identify which virtual camera to accept operations for. In the initial state, since there is one virtual camera, the identifier of that one virtual camera is designated, and when there are a plurality of virtual cameras, the identifier of one virtual camera selected by the user is designated. The method of selecting a virtual camera on the operation screen will be described later.

[0104] In S702, the modeling data generation unit 204 accepts the designation of the time code of the virtual viewpoint image obtained by imaging with the virtual camera designated in S701 via the virtual camera control unit 201. The method of designating the time code is the same as the method of designating the time code described in S401, and the description thereof will be omitted.

[0105] In S703, the modeling data generation unit 204 accepts the operation information of the virtual camera designated in S701 via the virtual camera control unit 201. The operation method of the virtual camera is the same as the operation method of the virtual camera described in S402, and the description thereof will be omitted.

[0106] In S704, the modeling data generation unit 204 receives an instruction from the user as listed below and determines which of the instruction contents the received instruction content corresponds to. Here, the user instruction received is either the selection of a virtual camera, the addition of a virtual camera, or the output of modeling data. "Selection of a virtual camera" is an instruction to select a different virtual camera displayed on the tablet 105 than the virtual camera specified in S701. "Addition of a virtual camera" is an instruction to add a different virtual camera not displayed on the tablet 105 than the virtual camera specified in S701. The instruction to add a virtual camera is performed, for example, when the add button 350 shown in Fig. 8(a) is pressed. "Output of modeling data" is an instruction to output modeling data.

[0107] When obtaining the determination result that the operation instruction is "addition of a virtual camera", the modeling data generation unit 204 shifts the process to S705.

[0108] In S705, the modeling data generation unit 204 adds a virtual camera. The screen displayed on the tablet 105 switches from the operation screen shown in Fig. 8(a) to the operation screen shown in Fig. 8(b), a new tab 802 is added to the operation screen 320, and a virtual camera is also added. On the tab screen added thereby, it is possible to receive the time code of the virtual viewpoint image obtained by imaging with the added virtual camera, and the operations of the position and orientation of the added virtual camera.

[0109] In Fig. 8, as an example of a sports player's play scene, a virtual viewpoint image of a spike scene by a volleyball player is displayed. Fig. 8(a) shows a scene where the volleyball player is in a stance of stepping in and bending before jumping, and Fig. 8(b) shows a scene where the time code is advanced from the scene in Fig. 8(a) and the volleyball player is in a stance of raising the arm before jumping and hitting the ball.

[0110] Note that, in FIGS. 8(a) to 8(c), the case where the number of tabs and the number of virtual cameras are 2 is shown, but the present invention is not limited to this. The number of tabs and the number of virtual cameras may be 3 or more. Further, as will be described later, in the output example of FIG. 8(d), six virtual cameras are used, and by operating the time code of the virtual viewpoint images obtained by imaging with each virtual camera and the positions and postures of the respective virtual cameras, the spike scene of the volleyball player is captured in time series.

[0111] Further, regarding the time code of the virtual viewpoint image obtained by imaging with the added virtual camera and the position and posture of the added virtual camera, the values of the virtual cameras displayed on the screen when the addition instruction is received may be used as initial values.

[0112] When it is determined in S704 described above that the operation instruction is "selection of virtual camera", the modeling data generation unit 204 returns the process to S701 and executes a series of processes from S701 to S703 for the selected virtual camera. The selection instruction of the virtual camera is performed, for example, by a user selection operation on the tabs 801-802 shown in FIG. 8(a). When an image including the tabs 801-802 is displayed on the touch panel, the user operation may be a general touch operation or the like.

[0113] In the spike scene of FIG. 8, by operating in detail the time code etc. of the virtual viewpoint images obtained by imaging with each virtual camera, scenes such as when the jump of the volleyball player reaches the highest point and when the volleyball player hits the ball can be easily specified.

[0114] When it is determined in S704 described above that the operation instruction is "output of modeling data", the modeling data generation unit 204 shifts the process to S706. The output instruction of the modeling data is performed, for example, by a user operation on the output button 351 shown in FIG. 8(a). When the output button 351 is pressed by a user operation, the time code of the virtual viewpoint images obtained by imaging with all the virtual cameras specified up to the previous step and the positions and The posture is determined, and the process proceeds to S706.

[0115] From S706 to S710, the modeling data generation unit 204 executes processing for all the virtual cameras specified in S701.

[0116] In S707, the modeling data generation unit 204 determines the drawing range of the virtual camera using the time code of the virtual viewpoint image received in the process of S702 and the viewpoint information of the virtual camera including the position and posture of the virtual camera received in the process of S703. The method for determining the drawing range of the virtual camera is the same as the method for determining the drawing range of the virtual camera described in S404, and the description thereof is omitted.

[0117] In S708, the modeling data generation unit 204 determines the generation range of the modeling data corresponding to this virtual camera based on the drawing range of the virtual camera determined in S707. The method for determining the generation range of the modeling data is the same as the method for determining the generation range of the modeling data described in S405, and the description thereof is omitted.

[0118] In S709, the modeling data generation unit 204 acquires the 3D model of the object included in the generation range of the modeling data of the virtual camera determined in S708 from the database 103. The method for acquiring the 3D model is the same as the method for acquiring the 3D model described in S406, and the description thereof is omitted.

[0119] In S711, the modeling data generation unit 204 synthesizes the 3D models included in the generation range of the modeling data determined for each virtual camera acquired up to the previous step based on the priority for each object.

[0120] The priority for each object is a parameter set during operations on the virtual cameras in S702 and S703, and is referenced when generating modeling data. The priority for each object will be described with reference to FIG. 8(c). As shown in FIG. 8(c), on the operation screen 320, for example, when a long tap is performed on an object, a priority setting screen 820 is displayed. By selecting the priority from among the selection items (e.g., high, medium, low) displayed on the priority setting screen 820 through user operation, the selected priority is set. Note that the selection items for the priority are not limited to three. An initial value for the priority for each object is set in advance, and if the priority is not set on the setting screen 820, the initial value may be set as the priority for each object.

[0121] In the composition based on the priority, for example, compared with an object with a high priority, processing such as reducing the color density and making the color lighter is performed on an object with a medium or low priority, and modeling data is generated in which processing is performed to make the object with a high priority stand out. The plurality of objects combined in this way will be described with reference to FIG. 8(d). Although the details of FIG. 8(d) will be described later, this is an example in which 3D models 831 - 836 of six objects are combined using virtual cameras related to the imaging of six virtual viewpoint images with different time codes and output as modeling data. In FIG. 8(d), the 3D model 833 and the 3D model 834 have a priority set to "high" and are combined in normal colors, and the 3D model 831, the 3D model 832, the 3D model 835, and the 3D model 836 have a priority set to "low" and are combined in light colors. By using such a priority setting, it is possible to generate modeling data that emphasizes the object at the time code to be noted.

[0122] Note that when a plurality of objects overlap, modeling data may be generated by combining the object with a priority set to "high" in front of the object with a priority set to "low".

[0123] In S712, the modeling data generation unit 204 generates a 3D model for assisting the 3D model obtained from the database, although it does not actually exist. The method for generating the 3D model of the auxiliary unit is the same as the method for generating the 3D model of the auxiliary unit described in S407, and the description thereof is omitted.

[0124] In S713, the modeling data generation unit 204 synthesizes the set of 3D models generated in the processes up to the previous step S712 to generate modeling data. An example of the synthesized modeling data will be described with reference to FIG. 8(d).

[0125] In FIG. 8(d), the modeling data is generated by acquiring, from the database 103, the 3D models included in the generation range of each modeling data using the time code, six virtual viewpoint images with different positions and postures, and virtual cameras, and synthesizing them. Specifically, the modeling data includes the 3D models 831 - 836 of a person, the 3D models 843 and 844 of columns, and the 3D model 850 of a pedestal.

[0126] The coordinates of the 3D models of the foreground such as a person and the 3D models of the background such as a field, which are the basis for these, are represented by a single three - dimensional coordinate shown in FIG. 3(a), and their positional relationships accurately reproduce the positional relationships and postures of the players on the actual field.

[0127] Also, even if the 3D models are of the same person (object), by using the virtual viewpoint images obtained by imaging with a plurality of virtual cameras having different time codes, the following data can be output. That is, by calculating the generation range of the modeling data based on their drawing areas, it can be output as modeling data arranging a series of plays of a sports player in time series.

[0128] Note that the target object is not limited to one, and as shown in Embodiment 1, it is of course possible to specify the time codes of different virtual viewpoint images and the virtual cameras of the positions and postures of the virtual cameras for a plurality of objects.

[0129] As described above, according to this embodiment, it is possible to generate modeling data of an object according to a plurality of pieces of specified information specified by a user operation. That is, based on the drawing ranges of a plurality of virtual cameras according to the specified information, it is possible to generate modeling data of a plurality of objects.

[0130] For example, it is possible to output, as modeling data, a series of plays (continuous movements) of professional sports players synthesized into one 3D model figure. For example, it is possible to output modeling data in which scenes such as a spike scene of a volleyball player or a jump scene of a figure skater are arranged in time series.

[0131] [Other Embodiments] In the above-described embodiment, an example of generating modeling data in a virtual viewpoint image generated using a 3D model generated based on a plurality of captured images has been described. However, the present invention is not limited to this, and for example, the present embodiment is also applicable to a moving image using a 3D model generated using computer graphics (CG) software or the like.

[0132] The present disclosure can also be realized by supplying a program that realizes one or more functions of the above-described embodiment to a system or device via a network or a storage medium, and having one or more processors in a computer of the system or device read and execute the program. Further, it can also be realized by a circuit (for example, ASIC) that realizes one or more functions.

Description of Reference Numerals

[0133] 201 Virtual Camera Control Unit 204 Modeling Data Generation Unit

Claims

1. Specific means for identifying the object that is the target for generating modeling data in a moving image generated using shape data representing the three-dimensional shape of the object, Acquisition means for acquiring the shape data of the object identified by the specific means, Generation means for generating modeling data including a plurality of the shape data respectively corresponding to different times, acquired by the acquisition means, An information processing apparatus characterized by comprising the above.

2. The information processing apparatus according to claim 1, wherein when the shape data acquired by the acquisition means is not in the same data format as the modeling data of the object, the generation means generates by converting the data format of the shape data acquired by the acquisition means.

3. The moving image is a virtual viewpoint image that uses shape data generated based on a plurality of captured images acquired by a plurality of imaging devices. The information processing apparatus according to claim 1 or 2, characterized by the above.

4. The specific means identifies the time code of the virtual viewpoint image, the position of the virtual viewpoint related to the virtual viewpoint image, and the viewing direction from the virtual viewpoint. The information processing apparatus according to claim 3, characterized by the above.

5. The specific means identifies the object that is the target of the modeling data based on the position of the virtual viewpoint related to the virtual viewpoint image and the viewing direction from the virtual viewpoint. The information processing apparatus according to claim 3 or 4, characterized by the above.

6. The specific means identifies the object based on the drawing range related to the display means on which the virtual viewpoint image is displayed. The information processing apparatus according to claim 5, characterized by the above.

7. The specific means identifies the object according to the position of the background structure included in the drawing range related to the display means. The information processing apparatus according to claim 6, characterized by the above.

8. The generation means generates the modeling data including a support part that supports the object according to the position and orientation of the object. The information processing apparatus according to any one of claims 1 to 7, characterized by the above.

9. The generation means generates the modeling data using the ground included in the moving image as the pedestal of the object. The information processing apparatus according to any one of claims 1 to 8, characterized by the above.

10. The acquisition means acquires the shape data of each of a plurality of the objects that are different from each other at the same time and are specified by the specifying means. The generation means generates the modeling data based on the shape data of each of a plurality of the objects that are different from each other at the same time and are acquired by the acquisition means. The information processing apparatus according to any one of claims 1 to 9, characterized in that.

11. It has setting means for setting a priority for each of the plurality of objects. The generation means generates the modeling data generated based on the plurality of acquired shape data according to the priority of the object set by the setting means. The information processing apparatus according to claim 10, characterized in that.

12. The generation means generates the modeling data in which the color is made lighter for the object with a lower priority compared to the object with a higher priority. The information processing apparatus according to claim 11, characterized in that.

13. The generation means generates the modeling data in which the object with a higher priority is arranged in front of the object with a lower priority when the plurality of objects overlap on the moving image. The information processing apparatus according to claim 11 or 12, characterized in that.

14. The shape data is further generated by providing a sub - pedestal for the object and providing a depression corresponding to the sub - pedestal on the pedestal. The information processing apparatus according to claim 9, characterized in that.

15. The shape of the sub - pedestal is generated based on the object specified by the specifying means. The information processing apparatus according to claim 14, characterized in that.

16. The specifying means specifies the object in the image at a predetermined time in the moving image. The acquisition means acquires the shape data of the object corresponding to the predetermined time. The information processing apparatus according to claim 1 or 2, characterized in that.

17. The modeling data is data for modeling a model figure or a relief. The information processing apparatus according to any one of claims 1 to 16, characterized in that.

18. The plurality of shape data included in the modeling data corresponds to the same scene. The information processing apparatus according to any one of claims 1 to 17, characterized in that.

19. An information processing apparatus according to any one of claims 1 to 18, a modeling apparatus that models a three-dimensional object of the object based on the modeling data generated by the generation means of the information processing apparatus, A system characterized by comprising:

20. A specifying step of specifying the object that is the target for generating modeling data in a moving image generated using shape data representing the three-dimensional shape of the object; An acquisition step of acquiring the shape data of the object specified in the specifying step; A generation step of generating modeling data including a plurality of the shape data respectively corresponding to different times acquired in the acquisition step; An information processing method characterized by comprising:

21. A program for causing a computer to function as the information processing apparatus according to any one of claims 1 to 18.

Citation Information

Patent Citations

  • Video distribution method, video reception method, server, terminal device and video distribution system

    JP2016010145A

  • Information processing apparatus and information processing method

    JP2016035623A

  • System, device and method of 3D printing

    JP2016163996A

  • Image retrieval system, image retrieval device, image retrieval method and program

    JP2019159593A

  • Doll molding system, information processing method, and program

    JP2020062322A