Information processing device, information processing method, and program

JP7920369B2Active Publication Date: 2026-09-14CANON KK
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025088199
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2026-09-14
Estimated Expiration
2041-02-26

AI Technical Summary

Benefits of technology

【0008】 本開示によれば、動画像の任意のシーンにおけるオブジェクトの立体物を容易に造形することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007920369000001
    Figure 0007920369000001
  • Figure 0007920369000002
    Figure 0007920369000002
  • Figure 0007920369000003
    Figure 0007920369000003
Patent Text Reader

Abstract

To allow for creating a three-dimensional representation of an object in any scene of a video image.SOLUTION: An information processing device provided herein is configured to determine an object for which modeling data is to be generated and a level of priority of the object, acquire a 3D model of the object created based on a plurality of captured images, and generate modeling data including the 3D model on the basis of the 3D model and the level of priority.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a technique for generating modeling data of an object from a moving image.

Background Art

[0002] In recent years, it has become possible to form a figure of an object based on a three-dimensional model (hereinafter referred to as a 3D model), which is data representing the three-dimensional shape of the object, using a modeling apparatus such as a 3D printer. Objects to be modeled include not only characters appearing in games and animations but also real-life persons. By inputting a 3D model obtained by imaging or scanning a real-life person or the like into a 3D printer, a figure can be modeled at a size of one-tenth to one-twentieth of that of the person or the like.

[0003] Patent Document 1 discloses a method for modeling a doll of a desired object by allowing a user to select a desired scene or an object included in the scene from a list of highlight video scenes created by imaging a sports game.

Prior Art Literature

Patent Literature

[0004]

Patent Literature 1

Summary of Invention

Problem to be Solved by Invention

[0005] However, in Patent Document 1, an object to be modeled can only be selected from the highlight scenes in the list, and it is difficult to model a doll of an object included in a scene not listed.

[0006] An object of the present disclosure is to enable easy modeling of a three-dimensional object of an object in any scene of a moving image.

Means for Solving Problem

[0007] An information processing device relating to one aspect of this disclosure is: Acquisition means for acquiring viewpoint information indicating the position and orientation of a virtual camera set in a virtual space associated with the target area of ​​imaging, and a plurality of first 3D models corresponding to a plurality of objects generated based on a plurality of captured images, and identification means for identifying a second 3D model among the plurality of first 3D models that is included in the imaging range of the virtual camera, based on the viewpoint information, and the second 3D model A decision-making mechanism for determining priority for, The second Based on the 3D model and the priority, Second The present invention is characterized by having a generation means for generating modeling data, including a 3D model. [Effects of the Invention]

[0008] According to this disclosure, it is possible to easily create three-dimensional objects of any scene in a video. [Brief explanation of the drawing]

[0009] [Figure 1] Diagram showing an example of the configuration of an information processing system. [Figure 2] Diagram showing an example configuration of an image generation device. [Figure 3] Diagram showing a virtual camera and its operation screen. [Figure 4] Flowchart showing the process of generating data for modeling. [Figure 5] A diagram showing the range and examples of data generation for modeling. [Figure 6] A diagram illustrating examples of information managed by a database. [Figure 7] Flowchart showing the process of generating data for modeling. [Figure 8] This figure shows examples of specifying and generating the rendering ranges for multiple virtual cameras. [Modes for carrying out the invention]

[0010] Embodiments of the present invention will be described below with reference to the drawings. Note that the following embodiments are not intended to limit the scope of the claims of this disclosure, and not all combinations of features described in these embodiments are necessarily essential to the solutions of this disclosure. The same reference numerals are used for identical components, and their descriptions are omitted.

[0011] [Embodiment 1] This embodiment describes a method for generating modeling data in a system that renders a virtual viewpoint image using three-dimensional shape data (hereinafter referred to as a 3D model) representing the three-dimensional shape of an object, which is obtained from captured images from multiple viewpoints. A virtual viewpoint image is an image generated by an end user and / or a designated operator specifying the position and line of sight of the virtual viewpoint, and is also called a free viewpoint image or arbitrary viewpoint image. The virtual viewpoint image may be a video or a still image, but this embodiment will be explained using the video as an example. In the following explanation, the virtual viewpoint will be replaced with a virtual camera in its original sense. In the following explanation, the position of the virtual viewpoint will correspond to the position of the virtual camera, and the line of sight from the virtual viewpoint will correspond to the orientation of the virtual camera.

[0012] (System Configuration) Figure 1 shows an example configuration of an information processing system (virtual viewpoint image generation system) that generates data for creating a three-dimensional object of an object in a virtual viewpoint image. Figure 1(a) shows an example configuration of the information processing system 100, and Figure 1(b) shows an example of the installation of the sensor system that the information processing system has. The information processing system 100 has n sensor systems 101a-101n, an image recording device 102, a database 103, an image generation device 104, and a tablet 105. Each of the sensor systems 101a-101n has at least one imaging device, which is a camera. In the following, unless otherwise specified, the n sensor systems from sensor system 101a to sensor system 101n will not be distinguished and will be referred to as multiple sensor systems 101.

[0013] An example of installation of the multi-sensor system 101 and a virtual camera will be described with reference to FIG. 1(b). As shown in FIG. 1(b), the multi-sensor system 101 is installed so as to surround a region 120 that is an imaging target region, and the cameras of the multi-sensor system 101 image the region 120 from different directions respectively. The virtual camera 110 images the region 120 from a direction different from that of the cameras of the multi-sensor system 101. Details of the virtual camera 110 will be described later.

[0014] When the imaging target is a professional sports game such as rugby or soccer, the region 120 is a field (ground) of a stadium, and n (for example, 100) multi-sensor systems 101 are installed so as to surround the field. Further, the imaging target region 120 may include not only people on the field but also balls and other objects. The imaging target is not limited to the field of a stadium, and may be a music concert held in an arena or the like, or a CM shooting held in a studio, as long as the multi-sensor system 101 can be installed. The number of the sensor systems 101 to be installed is not limited. Further, the multi-sensor system 101 does not need to be installed over the entire circumference of the region 120, and may be installed only on a part of the periphery of the region 120 depending on restrictions on the installation location or the like. Further, the plurality of cameras included in the multi-sensor system 101 may include imaging devices having different functions such as a telephoto camera and a wide-angle camera.

[0015] The plurality of cameras included in the multi-sensor system 101 perform imaging synchronously and acquire a plurality of images. Each of the plurality of images may be a captured image, or may be an image obtained by performing image processing such as processing for extracting a predetermined region on the captured image.

[0016] Note that each of the sensor systems 101a-101n may include a microphone (not shown) in addition to a camera. The microphones of each of the plurality of sensor systems 101 collect audio in synchronization. Based on the collected audio, it is possible to generate an acoustic signal that is reproduced together with the display of an image in the image generation apparatus 104. Hereinafter, for the sake of simplifying the description, the description of audio will be omitted, but it is assumed that basically both images and audio are processed.

[0017] The image recording apparatus 102 acquires a plurality of images from the plurality of sensor systems 101, and stores the obtained plurality of images and the time codes used for imaging together in the database 103. The time code is time information represented by an absolute value for uniquely identifying the imaging time, and is time information that can be specified in a format such as day:hour:minute:second.frame number, for example.

[0018] The database 103 manages event information, 3D model information, and the like. The event information includes data indicating the storage location of 3D model information for each object, which is associated with all time codes of an event to be imaged. The objects may include people and things that the user wants to form, and may also include people and things that are not targets of forming. The 3D model information includes information related to the 3D model of the object.

[0019] The image generation apparatus 104 receives, as inputs, an image corresponding to a time code from the database 103 and information about the virtual camera 110 set by a user operation from the tablet 105. The virtual camera 110 is set in a virtual space associated with the area 120, and allows the area 120 to be viewed from a viewpoint different from that of any of the cameras of the plurality of sensor systems 101. Details of the virtual camera 110, its operation method and operations will be described later with reference to the drawings.

[0020] The image generation device 104 generates a 3D model for generating virtual viewpoint images based on images for each time code obtained from the database 103, and generates a virtual viewpoint image using the generated 3D model for generating virtual viewpoint images and information about the viewpoint of the virtual camera. The viewpoint information of the virtual camera includes information indicating the position and orientation of the virtual viewpoint. Specifically, the viewpoint information includes parameters representing the three-dimensional position of the virtual viewpoint and parameters representing the orientation of the virtual viewpoint in the pan, tilt, and roll directions. The virtual viewpoint image is an image representing the view from the virtual camera 110 and is also called a free-viewpoint image. The virtual viewpoint image generated by the image generation device 104 is displayed on a touch panel such as a tablet 105.

[0021] In this embodiment, the image generation device 104 generates modeling data (object shape data) for use in the modeling device 106 from 3D models corresponding to objects present in the drawing range of at least one virtual camera. That is, the image generation device 104 sets conditions for identifying the object to be modeled for the virtual viewpoint image, acquires shape data representing the three-dimensional shape of the object based on the set conditions, and generates modeling data based on the acquired shape data. Details of the modeling data generation process will be described later with reference to the figures. The format of the modeling data generated by the image generation device 104 can be any format that the modeling device 106 can process, and may be a general-purpose polygon mesh format for handling common 3D models, or a proprietary format of the modeling device 106. That is, the modeling device 106 can process shape data for generating virtual viewpoint images, and the image generation device 104 identifies shape data for generating virtual viewpoint images as modeling data. On the other hand, the molding device 106 is unable to process the shape data for generating the virtual viewpoint image, so the image generation device 104 generates the molding data using the shape data for generating the virtual viewpoint image.

[0022] The tablet 105 is a portable device having a touch panel that has both a display unit for displaying images and an input unit for receiving user operations. The tablet 105 may also be a portable device with other functions, such as a smartphone. The tablet 105 accepts user operations to set information related to the virtual camera. The tablet 105 also displays virtual viewpoint images generated by the image generation device 104 or virtual viewpoint images stored in the database 103. The tablet 105 then accepts user operations to set conditions for identifying the object to be modeled based on the virtual viewpoint image. This user operation sets an arbitrary spatial and temporal range for generating modeling data. Details of this operation method will be described later with reference to a diagram. Note that the operation of the virtual camera is not limited to user operations on the touch panel of the tablet 105, but may also be user operations on an operating device such as a three-axis controller.

[0023] As shown in Figure 1, the image generation device 104 and the tablet 105 may be configured as separate devices or as an integrated unit. In the integrated configuration, the image generation device 104 has a touch panel or the like to receive input from the virtual camera and displays the virtual viewpoint image generated by the image generation device 104 on its touch panel. The virtual viewpoint image may also be displayed on the LCD screen of a device other than the tablet 105 or the image generation device 104.

[0024] The molding device 106 is, for example, a 3D printer, and takes the molding data generated by the image generation device 104 as input to create a three-dimensional object of the corresponding object, such as a doll (3D model figure) or a relief. The molding method of the molding device 106 is not limited to stereolithography, inkjet, powder deposition, etc., as long as it can create a three-dimensional object. Furthermore, the molding device 106 is not limited to a device that outputs three-dimensional objects such as dolls, but may also be a device that prints on plates or paper.

[0025] The image generation device 104 and the molding device 106 may be configured as separate devices, as shown in Figure 1, or they may be configured as an integrated unit.

[0026] The configuration of the information processing system 100 is not limited to the configuration shown in Figure 1(a), in which the tablet 105 and the molding device 106 are connected one-to-one with the image generation device 104. For example, it may be an information processing system in which multiple tablets 105 and multiple molding devices 106 are connected to the image generation device 104.

[0027] (Configuration of the image generation device) An example configuration of the image generation device 104 will be explained using a diagram. Figure 2 shows an example configuration of the image generation device 104, with Figure 2(a) showing an example of the functional configuration of the image generation device 104 and Figure 2(b) showing an example of the hardware configuration of the image generation device 104.

[0028] As shown in Figure 2(a), the image generation device 104 includes a virtual camera control unit 201, a 3D model generation unit 202, an image generation unit 203, and a modeling data generation unit 204. The image generation device 104 uses the aforementioned functional units to generate modeling data based on the drawing range of at least one virtual camera. Here, we will describe the overview of each function, and the details of the processing will be described later.

[0029] The virtual camera control unit 201 receives virtual camera operation information from the tablet 105 or the like. This virtual camera operation information includes at least the virtual camera's position and orientation, and its time code. Details of this virtual camera operation information will be described later with reference to a diagram. If the image generation device 104 has a touch panel or the like and is configured to receive virtual camera operation information, the virtual camera control unit 201 receives virtual camera operation information from the image generation device 104.

[0030] The 3D model generation unit 202 generates a 3D model representing the three-dimensional shape of an object within the region 120 based on multiple captured images. Specifically, the 3D model generation unit 202 obtains a foreground image, which is an image containing the foreground region corresponding to an object such as a person or a ball, and a background image, which is an image containing the background region other than the foreground region. The 3D model generation unit 202 then generates a 3D model of the foreground for each object based on the multiple foreground images.

[0031] These 3D models are generated using shape estimation methods such as Visual Hull and consist of point clouds. However, the data format of the 3D models representing the shape of each object is not limited to this. The background 3D model may be acquired in advance by an external device.

[0032] The 3D model generation unit 202 records the generated 3D model along with the time code in the database 103.

[0033] The 3D model generation unit 202 may be included in the image recording device 102 instead of the image generation device 104. In that case, the image generation device 104 only needs to read the 3D model generated by the image recording device 102 from the database 103 via the 3D model generation unit 202.

[0034] The image generation unit 203 retrieves a 3D model from the database 103 and generates a virtual viewpoint image based on the retrieved 3D model. Specifically, the image generation unit 203 retrieves appropriate pixel values ​​from the image for each point that makes up the 3D model and performs coloring. Then, the image generation unit 203 places the colored 3D model in a three-dimensional virtual space, projects it onto a virtual camera (virtual viewpoint), and renders it to generate a virtual viewpoint image.

[0035] However, the method for generating virtual viewpoint images is not limited to this, and various methods may be used, such as generating virtual viewpoint images by projective transformation of captured images without using a 3D model.

[0036] The modeling data generation unit 204 calculates and determines the drawing range or projection range of the virtual camera using the virtual camera's position and orientation, as well as the time code. Then, based on the determined virtual camera drawing range, the modeling data generation unit 204 determines the modeling data generation range and generates modeling data from the 3D models included in that range. Details of these processes will be described later with reference to the diagrams.

[0037] (Hardware configuration of the image generation device) Next, the hardware configuration of the image generation device 104 will be explained using Figure 2(b). As shown in Figure 2(b), the image generation device 104 includes a CPU 211, RAM 212, ROM 213, an operation input unit 214, a display unit 215, and a communication I / F (interface) unit 216.

[0038] The CPU (Central Processing Unit) 211 processes data using programs and data stored in the RAM (Random Access Memory) 212 and ROM (Read Only Memory) 213.

[0039] The CPU 211 controls the overall operation of the image generation device 104 and executes processing to realize each function shown in Figure 2(a). The image generation device 104 may have one or more dedicated hardware components separate from the CPU 211, and at least a portion of the processing performed by the CPU 211 may be executed by the dedicated hardware. Examples of dedicated hardware include ASICs (Application-Specific Integrated Circuits), FPGAs (Field-Programmable Gate Arrays), and DSPs (Digital Signal Processors).

[0040] ROM213 holds programs and data. RAM212 has a work area for temporarily storing programs and data read from ROM213. RAM212 also provides a work area used by CPU211 when executing various processes.

[0041] The operation input unit 214 is, for example, a touch panel, which accepts user input and acquires information entered through the accepted user operation. Examples of input information include information about the virtual camera and information about the timecode of the virtual viewpoint image to be generated. The operation input unit 214 may also be connected to an external controller to accept user input information regarding the operation. The external controller may be, for example, a three-axis controller such as a joystick or an operating device such as a mouse. However, the external controller is not limited to these.

[0042] The display unit 215 displays a virtual viewpoint image generated by a touch panel or screen. In the case of a touch panel, the operation input unit 214 and the display unit 215 are integrated into a single configuration.

[0043] The communication interface unit 216 transmits and receives information with the database 103, tablet 105, and molding device 106, for example, via a LAN. The communication interface unit 216 may also transmit information to an external screen via an image output port compatible with the following communication standards. Examples of image output ports compatible with these communication standards include HDMI® (High-Definition Multimedia Interface) and SDI (Serial Digital Interface). Furthermore, the communication interface unit 216 may transmit image data and molding data via Ethernet, for example.

[0044] (Virtual camera (virtual viewpoint) and its operation screen) Next, we will explain the virtual camera and its operation screen, using the example of capturing images of a rugby match at a rugby field. Figure 3 shows the virtual camera and its operation screen. Figure 3(a) shows the coordinate system, Figure 3(b) shows an example of a field to which the coordinate system in Figure 3(a) is applied, Figures 3(c) and 3(d) show examples of the drawing range of the virtual camera, and Figure 3(e) shows an example of the movement of the virtual camera. Figure 3(f) shows an example of the display of a virtual viewpoint image as seen from the virtual camera.

[0045] First, we will explain the coordinate system that represents the three-dimensional space of the object to be imaged, which serves as the basis for setting the virtual viewpoint. As shown in Figure 3(a), in this embodiment, a Cartesian coordinate system is used that represents the three-dimensional space with three axes: the X, Y, and Z axes. This Cartesian coordinate system is set for each object shown in Figure 3(b), namely the rugby field 391, the ball 392 on it, the players 393, etc. Furthermore, it may also be set for facilities (structures) within the rugby field such as spectator stands and billboards surrounding the field 391. Specifically, the origin (0,0,0) is set to the center of the field 391. The X axis is set in the direction of the long side of the field 391, the Y axis in the direction of the short side of the field 391, and the Z axis in the direction perpendicular to the field 391. Note that the direction of each axis is not limited to these. The position and orientation of the virtual camera 110 are specified using such a coordinate system.

[0046] Next, the rendering range of the virtual camera will be explained using a diagram. In the square pyramid 300 shown in Figure 3(c), vertex 301 represents the position of the virtual camera 110, and the line-of-sight vector 302, originating from vertex 301, represents the orientation of the virtual camera 110. Vector 302 is also called the optical axis vector of the virtual camera. The position of the virtual camera is expressed by the components (x, y, z) of each axis, and the orientation of the virtual camera 110 is expressed by a unit vector with scalar components of each axis. The vector 302 representing the orientation of the virtual camera 110 is assumed to pass through the center points of the front clipping plane 303 and the rear clipping plane 304. The viewing frustum of the virtual camera, which is the projection range (rendering range) of the 3D model, is the space 305 sandwiched between the front clipping plane 303 and the rear clipping plane 304.

[0047] Next, the components indicating the rendering range of the virtual camera will be explained using a diagram. Figure 3(d) is a view of the virtual viewpoint in Figure 3(c) from above (Z-axis). The rendering range is determined by the following values: the distance 311 from vertex 301 to the front clipping plane 303, the distance 312 from vertex 301 to the rear clipping plane 304, and the field of view 313 of the virtual camera 110. These values ​​may be predetermined values ​​(default values) set in advance, or they may be set values ​​that have been changed by user operation. In addition, the field of view 313 may be a value obtained by using the focal length of the virtual camera 110 as a variable. Note that the relationship between the field of view and focal length is a common technique, so its explanation will be omitted.

[0048] Next, we will explain how to change the position of the virtual camera 110 (move the virtual viewpoint) and how to change the orientation of the virtual camera 110 (rotate it). The virtual viewpoint can be moved and rotated in a space represented by three-dimensional coordinates. Figure 3(e) is a diagram illustrating the movement and rotation of the virtual camera. In Figure 3(e), the dashed arrow 306 represents the movement of the virtual camera (virtual viewpoint), and the dashed arrow 307 represents the rotation of the moved virtual camera (virtual viewpoint). The movement of the virtual camera is expressed by the components of each axis (x, y, z), and the rotation of the virtual camera is expressed by yaw (rotation around the Z axis), pitch (rotation around the X axis), and roll (rotation around the Y axis). In this way, the virtual camera can be freely moved and rotated in the three-dimensional space of the subject (field), so it can be generated as a virtual viewpoint image in which any area of ​​the subject is the drawing range.

[0049] Next, we will explain the operation screen for setting the position and orientation of the virtual camera (virtual viewpoint). Figure 3(f) is a diagram illustrating an example of the operation screen for the virtual camera (virtual viewpoint).

[0050] In this embodiment, the generation range of the modeling data is determined based on the drawing range of at least one virtual camera; therefore, the virtual camera operation screen 320 shown in Figure 3(f) can also be said to be an operation screen for determining the generation range of the modeling data. In Figure 3(f), the virtual camera operation screen 320 is displayed on the touch panel of the tablet 105. However, the display destination of the virtual camera operation screen 320 is not limited to this, and it may also be the touch panel of the image generation device 104, etc.

[0051] On the operation screen 320, the rendering range of the virtual camera (the rendering range related to capturing the virtual viewpoint image) is displayed as a virtual viewpoint image, aligning it with the screen frame of the operation screen 320. This display allows the user to visually set the conditions for identifying the object to be modeled.

[0052] The operation screen 320 has a virtual camera operation area 322 that accepts user operations to set the position and orientation of the virtual camera 110, and a time code operation area 323 that accepts user operations to set the time code. First, the virtual camera operation area 322 will be described. Since the operation screen 320 is displayed on a touch panel, the virtual camera operation area 322 accepts general touch operations 325 such as taps, swipes, and pinch-in / out as user operations. These touch operations 325 adjust the position and focal length (angle of view) of the virtual camera. The virtual camera operation area 322 also accepts touch operations 324 such as pressing and holding each axis of the Cartesian coordinate system as a user operation. These touch operations 324 rotate the virtual camera 110 around the X, Y, or Z axis, and the orientation of the virtual camera is adjusted. By assigning movement, rotation, and scaling of the virtual camera to each touch operation on the operation screen 320 in this way, the virtual camera 110 can be freely operated. These operation methods are publicly known, so their explanation will be omitted.

[0053] Furthermore, operations related to the position and orientation of the virtual camera are not limited to touch operations on a touch panel; they may also be performed using control devices such as joysticks.

[0054] Next, the timecode manipulation area 323 will be described. The timecode manipulation area 323 has a main slider 332 and knob 342, and a sub-slider 333 and knob 343, and has multiple elements for manipulating the timecode. The timecode manipulation area 323 also has a virtual camera add button 350 and an output button 351.

[0055] The main slider 332 is an input element that allows you to set a desired timecode from the total timecodes of the imaging data. When the position of the knob 342 is moved to the desired position by dragging or other means, the timecode corresponding to the position of the knob 342 is specified. In other words, any timecode can be specified by adjusting the position of the knob 342 on the main slider 332.

[0056] Sub-slider 333 is an input element that displays a magnified portion of the total timecode and allows for detailed timecode setting for the magnified portion. When the position of knob 343 is moved to the desired position by dragging, the timecode corresponding to the position of knob 343 is specified. The main slider 332 and sub-slider 333 appear to be the same length on the screen, but the range of timecode that can be selected differs. For example, the main slider 332 allows selection from the length of a match, which is 3 hours, while the sub-slider 333 allows selection from a portion of that, such as 30 seconds. In this way, the scale of the sliders differs, and the sub-slider allows for more detailed timecode specification, such as in seconds or frames.

[0057] The timecode specified using the knob 342 of the main slider 332 or the knob 343 of the sub-slider 333 may be displayed numerically in the format day:hour:minute:second.frame number. The sub-slider 333 may be displayed on the operation screen 320 at all times or only temporarily. For example, it may be displayed when a timecode display instruction is received or when an instruction for a specific operation such as pausing is received. The timecode interval that can be selected with the sub-slider 333 may be variable.

[0058] Based on the virtual camera's position and orientation set by the above operations, and the time code, the generated virtual viewpoint image is displayed in the virtual camera's operation area 322. In Figure 3(f), as an example, the subject is a rugby match, and a crucial pass scene leading to a score is displayed. As will be explained in more detail later, this crucial scene will be generated or output as shape data for a modeling object.

[0059] Figure 3(f) shows the case where the time code for the moment the ball is released is specified, but for example, by operating the sub-slider 333 in the time code operation area 323, the time code while the ball is in the air can also be easily specified.

[0060] In Figure 3(f), the position and orientation of the virtual camera are specified so that three players are within the drawing range, but this is not limited to this. For example, by performing touch operations 324 and 325 on the virtual camera operation area 322, the space (the drawing range of the virtual camera) 305 can be freely manipulated, and further manipulated so that the surrounding players are within the drawing range.

[0061] The output button 351 is operated when determining the drawing range of the virtual camera set by user operations on the virtual camera operation area 322 and the time code operation area 323, and outputting the shape data of the object to be modeled. Once the time code indicating the time range and the position and orientation indicating the spatial range are determined for the virtual camera, the drawing range of the virtual camera corresponding to these determinations is calculated. Then, based on the calculated drawing range of the virtual camera, the generation range of the modeling data is determined. Details of this process will be described later with reference to a diagram.

[0062] The virtual camera addition button 350 is used when multiple virtual cameras are used to generate the modeling data. Details of the modeling data generation process using multiple virtual cameras will be explained in Embodiment 2, so the explanation is omitted here.

[0063] The virtual camera operation screen is not limited to the operation screen 320 shown in Figure 3(f); it is sufficient if it allows operations to set the virtual camera's position, orientation, and timecode. For example, the virtual camera operation area 322 and the timecode operation area 323 do not need to be separate. For instance, if an operation such as a double tap is performed on the virtual camera operation area 322, an operation such as pausing may be processed as an operation to set the timecode.

[0064] Furthermore, although the case where the operation input unit 214 is a tablet 105 has been described, the operation input unit 214 is not limited to this and may be an operating device having a general display and a three-axis controller, etc.

[0065] (Generation process for modeling data) Next, the process of generating modeling data in the image generation device 104 will be explained using diagrams. Figure 4 is a flowchart showing the flow of the modeling data generation process. This series of processes is realized by the CPU 211 executing a predetermined program and operating each functional unit shown in Figure 2(a). Hereafter, steps will be denoted as "S". The same will be used in the following explanation. Note that the explanation assumes that the shape data representing the three-dimensional shape of the object used to generate the virtual viewpoint image is in a different format from the modeling data.

[0066] In S401, the modeling data generation unit 204 accepts the specification of the time code for the virtual viewpoint image via the virtual camera control unit 201. The method for specifying the time code may be, for example, using the slider shown in Figure 3(f), or by directly inputting numbers.

[0067] In S402, the modeling data generation unit 204 receives virtual camera operation information via the virtual camera control unit 201. The virtual camera operation information includes at least information regarding the position and orientation of the virtual camera. The method of operating the virtual camera may be, for example, a method using a tablet as shown in Figure 3(f), or a method using an operating device such as a joystick.

[0068] In S403, the modeling data generation unit 204, via the virtual camera control unit 201, determines whether or not an output instruction has been received, that is, whether or not the output button 351 has been operated. If the modeling data generation unit 204 determines that an output instruction has been received (YES in S403), the process proceeds to S404. If the modeling data generation unit 204 determines that an output instruction has not been received (NO in S403), the process returns to S401, and the processes in S401 and S402 are executed again. In other words, the process of accepting user input regarding the time code of the virtual viewpoint image and the position and orientation of the virtual camera continues until an output instruction is received.

[0069] In S404, the modeling data generation unit 204 determines the drawing range of the virtual camera using the time code of the virtual viewpoint image received in processing S401 and the viewpoint information of the virtual camera, including the position and orientation of the virtual camera, received in processing S402. The method for determining the drawing range of the virtual camera may be, for example, by user operation of the front clip surface 303 and the rear clip surface 304 shown in Figures 3(c) and 3(d), or it may be determined based on the calculation result of a calculation formula set in advance in the device.

[0070] In S405, the modeling data generation unit 204 determines the modeling data generation range corresponding to the virtual camera based on the drawing range of the virtual camera determined in S404. The method for determining the modeling data generation range will be explained below. Figure 5 shows the modeling data generation range and examples of generation. Figure 5(a) shows the modeling data generation range in three-dimensional space, and Figure 5(b) shows the yz plane of the modeling data generation range shown in Figure 5(a). Figure 5(c) shows an example of a 3D model within the modeling data generation range in three-dimensional space. Figure 5(d) shows the yz plane when spectator seats are present within the modeling data generation range. Figure 5(e) shows an example of a 3D model figure corresponding to the 3D model shown in Figure 5(c). Figure 5(f) shows another example of a 3D model figure corresponding to the 3D model shown in Figure 5(c).

[0071] In Figure 5(a), a virtual camera (position 301, space (frustum of view) 305, etc.) is displayed in three-dimensional space, and the plane of the object to be imaged at Z=0 (for example, the field of a stadium) 500 is displayed in the same space. As shown in Figure 5(a), the generation range 510 for the modeling data is determined based on the plane (base) 501 that is included in the space (frustum of view) 305, which is the drawing range of the virtual camera, within the plane 500 at Z=0.

[0072] The modeling data generation range 510 is a part of the space (frustum of view) 305, which is the rendering range of the virtual camera. The modeling data generation range 510 is the space enclosed by the plane 501, the front plane (front surface) 513 located near the front clipping surface 303, and the rear plane (rear surface) 514 located near the rear clipping surface 304. The sides of the space that constitutes the modeling data generation range 510 are set based on the virtual camera's position 301 and its field of view, and the plane 500. The top of the space that constitutes the modeling data generation range 510 is also set based on the virtual camera's position 301 and its field of view.

[0073] Planes 513 and 514 will be explained using Figure 5(b). Figure 5(b) is a simplified side view of the virtual camera and the plane at Z=0 in Figure 5(a). Plane 513 is a plane that passes through intersection 503 or 505, is located at a predetermined distance from the front clipping plane 303 in space 305, and is parallel to the front clipping plane 303. Plane 514 is a plane that passes through intersection 504 or 506, is located at a predetermined distance from the rear clipping plane 304 in space 305, and is parallel to the rear clipping plane 304. The predetermined distance is set in advance. Plane 501 is a rectangle with the intersection 503 to 506 as its vertices.

[0074] Note that plane 514 may be the same as the rear clip surface 304. Also, planes 513 and 514 do not have to be parallel to the front clip surface and the rear clip surface, respectively. For example, they may be planes that pass through the vertex of plane 501 and are perpendicular to plane 501.

[0075] In addition to the above, the generation range 510 for the modeling data may be determined by considering the spectator seats in the stadium. This example will be explained using Figure 5(d). In Figure 5(d), similar to Figure 5(b), the drawing range is shown when the virtual camera and the plane at Z=0 in Figure 5(a) are viewed from the side in a simplified manner. In Figure 5(d), the plane 514 may be determined based on the position of the spectator seats 507 in the stadium. Specifically, the plane 514 is defined as a plane that passes through the intersection of the spectator seats 507 and the plane 501 and is parallel to the rear clipping plane. This ensures that the area behind the spectator seats 507 is not included in the generation range 510 for the modeling data.

[0076] The conditions for determining the range of data to be generated for modeling are not limited to stadium seating; the user may manually specify objects, etc. The plane 500 that intersects with the frustum is not limited to Z=0; it can be any shape. For example, it may have irregularities like an actual field. If the field, which is the 3D model of the actual background, has irregularities, these may be corrected into a flat 3D model. The configuration may not have a plane or curved surface that intersects with the frustum. In that case, the rendering range of the virtual camera may be used directly as the range for generating the modeling data.

[0077] In S406, the modeling data generation unit 204 retrieves 3D models of objects included in the modeling data generation range determined in S405 from the database 103.

[0078] Examples of event information and 3D model information tables managed by database 103 are explained using diagrams. Figure 6 shows the tables of information managed by database 103, with Figure 6(a) showing the event information table and Figure 6(b) showing the 3D model information table. As shown in Figure 6(a), the event information table 610 managed by database 103 shows the storage location of 3D model information for each object for all time codes of the events being imaged. For example, the event information table 610 shows that the storage location of the 3D model information for the time code "16:14:24.041" and the object is object A is "DataA100".

[0079] Table 620 of the 3D model information managed by database 103 stores data for the following items, as shown in Figure 6(b): "Total Point Cloud Coordinates," "Texture," "Average Coordinates," "Centroid Coordinates," and "Maximum / Minimum Coordinates." "Total Point Cloud Coordinates" stores data on the coordinates of each point in the point cloud that makes up the 3D model. "Texture" stores data on the texture image applied to the 3D model. "Average Coordinates" stores data on the coordinates of the point obtained by averaging all the coordinates of the point cloud that makes up the 3D model. "Centroid Coordinates" stores data on the coordinates of the centroid point based on all the coordinates of the point cloud that makes up the 3D model. "Maximum / Minimum Coordinates" stores data on the coordinates of the maximum and minimum points among the coordinates of the point cloud that makes up the 3D model. Note that the items of data stored in Table 620 of the 3D model information are not limited to all of "Total Point Cloud Coordinates," "Texture," "Average Coordinates," "Centroid Coordinates," and "Maximum / Minimum Coordinates." For example, it may only include "Total Point Cloud Coordinates" and "Texture," or other items may be added to these items.

[0080] By using the information shown in Figure 6, when a certain time code is specified, it is possible to obtain 3D model information for the specified time code, such as the coordinates of all point clouds for each object, and the coordinates of the maximum and minimum values ​​for each axis of the three-dimensional coordinate system.

[0081] An example of obtaining a 3D model included in the modeling data generation range from the database 103 will be explained using Figure 5(c). The modeling data generation unit 204 refers to the 3D model information for each object, which is associated with the time code specified in S401, in the database 103, and determines whether or not each object is included in the modeling data generation range determined in S406.

[0082] The determination method could be, for example, a method that determines whether all point cloud coordinates included in the 3D model information for each object, such as a person, are included in the generation range, or a method that determines whether only the average value of all point cloud coordinates or the maximum and minimum values ​​of each axis of the three-dimensional coordinates are included in the generation range.

[0083] Figure 5(c) shows an example where, as a result of the above determination, the generation range 510 for the modeling data included three 3D models. As will be explained in more detail later, an example of these three 3D models becoming modeling data is shown as 3D models 531-533 in Figure 5(e).

[0084] Furthermore, the determination result of whether each object is included in the generation range for the modeling data may be displayed on the tablet's operation screen 320. As a result of the determination, for example, for an object that lies on the boundary of the generation range 510 for the modeling data, a warning may be displayed stating that the entire object will not be output as modeling data because it lies on the boundary.

[0085] The judgment process in S406 may be performed at any time while user input is being received regarding the timecode of the virtual viewpoint image in S401 and the position and orientation of the virtual camera in S402. If the judgment process in S406 is performed at any time in this manner, a warning display may be shown to indicate that the target object is on the boundary.

[0086] The method for obtaining the 3D model of an object is not limited to an automatic method based on the above-mentioned determination results; it can also be obtained manually. For example, the user can specify the object for which the 3D model is to be obtained by accepting user operations such as tapping on the object that the user wants to use as the target for the modeling data on the tablet's operation screen 310.

[0087] Furthermore, in the S406 process, 3D model information of the background field and spectator seats, which are included in the generation range of the modeling data, may be obtained from the database. In that case, the obtained 3D models of the field and spectator seats may be partially converted to form a base. For example, even if the 3D model of the field has no thickness, the base 540 may be made into a rectangular parallelepiped shape as shown in Figure 5(e) by converting it into a 3D model with a specified height added.

[0088] In S407, the modeling data generation unit 204 generates a 3D model that does not actually exist but is used to assist the 3D model obtained from the database. For example, the player in the foreground 3D figure model 532 in Figure 5(e) is jumping and is not touching the field but is in the air. Therefore, in order to display the 3D figure model, a support column (support part) is needed as an auxiliary to support the 3D figure model and fix it to a base. In S407, the modeling data generation unit 204 first determines whether an auxiliary part is needed using the 3D model information for each object obtained from the database 103. This determination includes determining whether the target 3D model is in the air and whether the target 3D model is self-supporting on the base. In determining whether the target 3D model is in the air, for example, the coordinates or minimum coordinates of the entire point cloud included in the 3D model information may be used. In determining whether the target 3D model is self-supporting on the base, for example, the coordinates of the entire point cloud included in the 3D model information, or the centroid coordinates and maximum / minimum coordinates may be used.

[0089] If the modeling data generation unit 204 determines that an auxiliary part is necessary, it generates a 3D model of the auxiliary part that fixes the target 3D figure model to the base, based on the 3D model information. If the target 3D model is in the air, the modeling data generation unit 204 may generate the 3D model of the auxiliary part based on the coordinates of the entire point cloud, or the minimum coordinate and its surrounding coordinates. If the target 3D model does not stand on its own on the base, the modeling data generation unit 204 may generate the 3D model of the auxiliary part based on the coordinates of the entire point cloud, or the center of gravity coordinate and maximum / minimum coordinate. The support column may be positioned vertically from the center of gravity coordinate included in the 3D model information for each object. In addition, a 3D model of the support column may be added even for objects that have contact with the field.

[0090] On the other hand, if the modeling data generation unit 204 determines that the auxiliary unit is not necessary, it proceeds to S408 without generating a 3D model of the auxiliary unit.

[0091] In S408, the modeling data generation unit 204 synthesizes the set of 3D models generated in the previous step up to S407, converts the synthesized set of 3D models into the output format, and outputs it as modeling data for a three-dimensional object. An example of the modeling data to be output is explained using Figure 5(e).

[0092] In Figure 5(e), the modeling data includes the foreground 3D models 531, 532, and 533, as well as the support column 534 and the base 540, which are included in the modeling data generation range 510. The coordinates of the foreground 3D models, such as people, and the background 3D models, such as the field, are represented by a single three-dimensional coordinate system as shown in Figure 3(a), and their positional relationship can accurately reproduce the positions and postures of players on an actual field.

[0093] When the modeling data shown in Figure 5(e) is output to the modeling device 106, the modeling device 106 can print a 3D model figure in that exact shape. Therefore, although Figure 5(e) is described as modeling data, it can also be considered as a 3D model figure output by a 3D printer.

[0094] Furthermore, in order to reduce the risk of damage during delivery or transportation, the modeling data in Figure 5(e) may be output separately in S408, separating the foreground 3D model from the background 3D model. For example, as shown in Figure 5(f), the base 540 corresponding to the background field and the 3D models 531, 532, and 533 corresponding to the foreground people may be output separately, and small sub-bases 541, 542, and 543 may be attached to the 3D models 531, 532, and 533. In this case, by providing recesses 551-553 in the base 540 that correspond to the size of the sub-bases 541-543, the foreground 3D models 531, 532, and 533 can be attached to the base 540. Since the sub-bases 541, 542, and 543 are 3D models that do not actually exist, they may be added in S407, similar to the support pillars, as auxiliary components for the foreground and other 3D models.

[0095] The shapes of the sub-bases 541-543 and recesses 551-553 may be polygons other than the quadrilaterals shown in Figure 5(f). Furthermore, the shapes of the sub-bases 541-543 and recesses 551-553 may be different, like puzzle pieces. In this case, after the 3D printing device outputs the foreground 3D model and background 3D model separately, the foreground 3D model can be properly mounted without making mistakes in mounting position or direction. The positions of the sub-bases 541-543 and recesses 551-553 on the base 540 for each object may be changeable by user operation.

[0096] The base 540 may be output with embedded modeling data containing information that does not actually exist. For example, the modeling data for the base 540 may be output with embedded information such as the time code of the virtual camera used in the modeling data, and information regarding the position and orientation of the virtual camera. Alternatively, the modeling data for the base 540 may be output with embedded three-dimensional coordinate values ​​so that the generation range of the modeling data can be identified.

[0097] As described above, according to this embodiment, it is possible to generate modeling data for objects according to specified information specified by user operation. That is, based on the rendering range of the virtual camera according to the specified information, it is possible to generate modeling data for objects included in the rendering range of the virtual camera.

[0098] For example, in field sports such as rugby and soccer played in a stadium, it is possible to capture decisive moments leading to goals, specifying a time code and spatial range, and generate modeling data while preserving the actual positions and postures of the players. Furthermore, it is possible to use this data to output to a 3D printer and generate a 3D model figure of that scene.

[0099] Furthermore, if the shape data representing the three-dimensional shape of the object used to generate the virtual viewpoint image is in the same format as the object's modeling data, the following processing will be performed in S408. That is, the modeling data generation unit 204 will output the data, which is the 3D model of the object generated in the previous step up to S407, with the 3D model of the auxiliary part added, as the object's modeling data. In this way, if the format of the shape data for generating the virtual viewpoint image is the same as the format of the modeling data, the shape data for generating the virtual viewpoint image can be output directly as modeling data.

[0100] [Embodiment 2] In this embodiment, a method for generating modeling data using multiple virtual cameras and their drawing ranges will be described.

[0101] In this embodiment, the configuration of the information processing system is the same as in Figure 1, and the configuration of the image generation device is the same as in Figure 2, so their explanations will be omitted, and the differences will be explained. In this embodiment, the operation method for using multiple virtual cameras in the tablet 105 or the image generation device 104, and the process of generating modeling data based on their multiple drawing ranges will be explained.

[0102] (Generation process for modeling data) The process for generating modeling data according to this embodiment will be explained with reference to the figures. Figure 7 is a flowchart showing the flow of the modeling data generation process according to this embodiment. This series of processes is realized by the CPU 211 executing a predetermined program to operate each functional unit shown in Figure 2(a). Figure 8 is a diagram illustrating an example of the operation screen for the virtual viewpoint (virtual camera). Figure 8(a) shows the case when virtual camera 1 is selected, Figure 8(b) shows the case when virtual camera 2 is selected, Figure 8(c) shows the case when the priority setting screen is displayed in Figure 8(b), and Figure 8(d) shows an example of a 3D model. The operation screens shown in Figures 8(a) to 8(c) are screens with expanded functionality compared to the operation screen shown in Figure 3(f), and only the differences will be explained here.

[0103] In S701, the modeling data generation unit 204 specifies the virtual camera to accept operations. In this embodiment, since multiple virtual cameras are handled, in S702 and S703, which follow the processing of S701, a specification is made to identify which virtual camera will accept operations. Initially, there is only one virtual camera, so the identifier of that single camera is specified. If there are multiple virtual cameras, the identifier of the one selected by the user is specified. The method for selecting a virtual camera on the operation screen will be described later.

[0104] In S702, the modeling data generation unit 204 receives the time code specification for the virtual viewpoint image obtained by imaging with the virtual camera specified in S701 via the virtual camera control unit 201. The method for specifying the time code is the same as the method for specifying the time code described in S401, so that explanation is omitted.

[0105] In S703, the modeling data generation unit 204 receives virtual camera operation information specified in S701 via the virtual camera control unit 201. The method for operating the virtual camera is the same as the method for operating the virtual camera described in S402, so that explanation is omitted.

[0106] In S704, the modeling data generation unit 204 receives the following instructions from the user and determines which instruction the received instruction corresponds to. The user instructions received here are one of the following: selection of a virtual camera, addition of a virtual camera, or output of modeling data. "Selection of a virtual camera" is an instruction to select a different virtual camera displayed on the tablet 105, which is different from the virtual camera specified in S701. "Addition of a virtual camera" is an instruction to add a different virtual camera that is not displayed on the tablet 105, which is different from the virtual camera specified in S701. The instruction to add a virtual camera is performed, for example, when the add button 350 shown in Figure 8(a) is pressed. "Output of modeling data" is an instruction to output modeling data.

[0107] When the modeling data generation unit 204 determines that the operation instruction is "add a virtual camera," it proceeds to process S705.

[0108] In S705, the modeling data generation unit 204 adds a virtual camera. The screen displayed on the tablet 105 switches from the operation screen shown in Figure 8(a) to the operation screen shown in Figure 8(b), and a new tab 802 is added to the operation screen 320, and a virtual camera is also added. As a result, the added tab screen can accept the time code of the virtual viewpoint image obtained from imaging by the added virtual camera, as well as the position and orientation of the added virtual camera.

[0109] Figure 8 shows a virtual viewpoint image of a volleyball player spiking, as an example of a sports player's play scene. Figure 8(a) shows the moment when the volleyball player steps forward and crouches before jumping, and Figure 8(b) shows the moment when the volleyball player jumps and raises their arms before hitting the ball, with the time code advanced from the scene in Figure 8(a).

[0110] Note that Figures 8(a) to 8(c) show the case where the number of tabs and virtual cameras is 2, but this is not the only case. The number of tabs and virtual cameras may be 3 or more. Also, as will be described later, in the output example in Figure 8(d), six virtual cameras are used, and the volleyball player's spiking scene is captured in chronological order by manipulating the timecode of the virtual viewpoint image obtained from the imaging of each virtual camera, as well as the position and orientation of each virtual camera.

[0111] Furthermore, the timecode of the virtual viewpoint image obtained by imaging with the added virtual camera, as well as the position and orientation of the added virtual camera, may be initialized using the virtual camera values ​​that were displayed on the screen when the additional instruction was received.

[0112] If the system determines in S704 that the operation instruction is "select a virtual camera," the modeling data generation unit 204 returns to S701 and executes the series of processes from S701 to S703 for the selected virtual camera. The instruction to select a virtual camera is made, for example, by a user selection operation on tabs 801-802 shown in Figure 8(a). If an image including tabs 801-802 is displayed on the touch panel, the user operation can be a general touch operation.

[0113] In the spike scene shown in Figure 8, by precisely manipulating the timecode and other parameters of the virtual viewpoint images obtained from each virtual camera, it is possible to easily specify moments such as when the volleyball player reaches the highest point of their jump or the moment when the volleyball player hits the ball.

[0114] If the system determines in S704 that the operation instruction is "output of modeling data", the modeling data generation unit 204 proceeds to S706. The instruction to output modeling data is given, for example, by user operation on the output button 351 shown in Figure 8(a). When the output button 351 is pressed by the user, the time codes of the virtual viewpoint images obtained from imaging by all virtual cameras specified up to the previous step, the positions of all virtual cameras, and Once that position is confirmed, the process moves on to S706.

[0115] In steps S706 through S710, the modeling data generation unit 204 performs processing for all virtual cameras specified in S701.

[0116] In S707, the modeling data generation unit 204 determines the drawing range of the virtual camera using the time code of the virtual viewpoint image received in processing S702 and the viewpoint information of the virtual camera, including the position and orientation of the virtual camera, received in processing S703. The method for determining the drawing range of the virtual camera is the same as the method for determining the drawing range of the virtual camera described in S404, and therefore the explanation is omitted.

[0117] In S708, the modeling data generation unit 204 determines the modeling data generation range corresponding to the virtual camera based on the drawing range of the virtual camera determined in S707. The method for determining the modeling data generation range is the same as the method for determining the modeling data generation range explained in S405, and therefore the explanation is omitted.

[0118] In S709, the modeling data generation unit 204 retrieves 3D models of objects included in the modeling data generation range of the virtual camera determined in S708 from the database 103. The method for retrieving the 3D models is the same as the method for retrieving 3D models described in S406, and therefore the explanation is omitted.

[0119] In S711, the modeling data generation unit 204 synthesizes 3D models included in the modeling data generation range determined for each virtual camera acquired up to the previous step, based on the priority of each object.

[0120] Object-specific priority is a parameter set during operations on the virtual cameras of S702 and S703, and is referenced when generating modeling data. Object-specific priority is explained using Figure 8(c). As shown in Figure 8(c), on the operation screen 320, for example, a long tap on an object will display the priority setting screen 820. The user selects a priority from the selection items (e.g., high, medium, low) displayed on the priority setting screen 820, and the selected priority is set. Note that the selection items for priority are not limited to three. Initial values ​​for object-specific priority may be set in advance, and if a priority is not set on the setting screen 820, the initial values ​​may be set as the priority for each object.

[0121] In priority-based synthesis, for example, objects with medium or low priority are processed to have lower color density and lighter colors compared to objects with high priority, resulting in modeling data that makes high-priority objects stand out. Figure 8(d) illustrates how multiple objects synthesized in this way can be seen. As will be explained in detail later, Figure 8(d) is an example in which 3D models 831-836 of six objects were synthesized using a virtual camera that captured six virtual viewpoint images with different time codes, and output as modeling data. In Figure 8(d), 3D models 833 and 834 are set to "high" priority and synthesized with normal colors, while 3D models 831, 832, 835, and 836 are set to "low" priority and synthesized with lighter colors. By using this priority setting, it is possible to generate modeling data that emphasizes objects with time codes of interest.

[0122] Furthermore, if multiple objects overlap, the modeling data may be generated by compositing them so that objects with a "high" priority are placed in front of objects with a "low" priority.

[0123] In S712, the modeling data generation unit 204 does not actually exist, but it generates a 3D model to assist the 3D model obtained from the database. The method for generating the 3D model for the auxiliary unit is the same as the method for generating the 3D model for the auxiliary unit described in S407, so that explanation is omitted.

[0124] In step S713, the modeling data generation unit 204 synthesizes the set of 3D models generated in the previous step, up to S712, to generate modeling data. An example of the synthesized modeling data is explained using Figure 8(d).

[0125] In Figure 8(d), the modeling data is generated by acquiring 3D models included in the generation range of each modeling data from database 103 using time codes and six virtual viewpoint images and virtual cameras with different positions and orientations, and then combining them. Specifically, the modeling data includes 3D models of people 831-836, 3D models of support columns 843 and 844, and 3D model of a pedestal 850.

[0126] The coordinates of the 3D models of the foreground (people, etc.) and the 3D models of the background (field, etc.) that served as the basis for these models are represented by a single three-dimensional coordinate system, as shown in Figure 3(a). Their positional relationship accurately reproduces the positions and postures of players on an actual field.

[0127] Furthermore, even if the 3D model is the same person (object), by using virtual viewpoint images obtained from multiple virtual cameras with different time codes, the following data can be output. In other words, by calculating the generation range of the modeling data based on their rendering areas, a series of plays by a sports player can be output as modeling data arranged in chronological order.

[0128] Furthermore, the target object is not limited to one; as shown in Embodiment 1, it is naturally possible to specify different virtual viewpoint image time codes and virtual camera positions and orientations for multiple objects.

[0129] As described above, according to this embodiment, it is possible to generate modeling data for objects according to multiple specified pieces of information specified by user operation. In other words, it is possible to generate modeling data for multiple objects based on the rendering ranges of multiple virtual cameras according to the specified information.

[0130] For example, it's possible to output modeling data that synthesizes a series of plays (continuous movements) of a professional athlete into a single 3D model figure. For instance, it's possible to output modeling data that arranges volleyball player spikes or figure skater jumps in chronological order.

[0131] [Other embodiments] The embodiments described above describe an example of generating modeling data in a virtual viewpoint image generated using a 3D model generated based on multiple captured images. However, the embodiments are not limited to this, and can also be applied to moving images using 3D models generated using, for example, computer graphics (CG) software.

[0132] This disclosure can also be implemented by supplying a program that implements one or more of the functions of the embodiments described above to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be implemented by a circuit (e.g., an ASIC) that implements one or more functions. [Explanation of symbols]

[0133] 201 Virtual Camera Control Unit 204 Modeling data generation unit

Claims

1. An acquisition means for acquiring viewpoint information indicating the position and orientation of a virtual camera set in a virtual space associated with the target area of ​​imaging, and a plurality of first 3D models corresponding to a plurality of objects generated based on a plurality of captured images, Based on the viewpoint information, a means for identifying a second 3D model among the plurality of first 3D models that is included in the imaging range of the virtual camera, A determination means for determining the priority of the second 3D model, A generation means for generating molding data including the second 3D model based on the second 3D model and the priority, An information processing device characterized by having the following features.

2. The aforementioned molding data includes the second 3D model corrected according to the priority. The information processing apparatus according to feature 1.

3. The aforementioned molding data includes the second 3D model, whose color information has been corrected according to the priority. The information processing apparatus according to feature 2.

4. The aforementioned priority is determined based on user operation. The information processing apparatus according to feature 1.

5. The aforementioned viewpoint information is acquired based on user operations. The information processing apparatus according to feature 1.

6. The aforementioned molding data includes a support portion that supports the shape of the second 3D model. The information processing apparatus according to any one of claims 1 to 5.

7. The aforementioned modeling data is data for creating a model figure or relief. The information processing apparatus according to any one of claims 1 to 6.

8. An acquisition step of acquiring viewpoint information indicating the position and orientation of a virtual camera set in a virtual space associated with the target area of ​​imaging, and multiple first 3D models corresponding to multiple objects generated based on multiple captured images, Based on the viewpoint information, a selection step is made to identify a second 3D model among the plurality of first 3D models that is included in the imaging range of the virtual camera, A decision step for determining the priority of the second 3D model, A generation step of generating molding data including the second 3D model based on the second 3D model and the priority, An information processing method characterized by having the following features.

9. A program for causing a computer to function as an information processing device according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Doll molding system, information processing method, and program

    JP2020062322A

  • Methods and apparati for implementing programmable pipeline for three-dimensional printing including multi-material applications

    US20140324204A1

  • 3D Printing Using 3D Video Data

    US20180046167A1