Information processing device, information processing method, and program

The information processing apparatus addresses the lack of scene knowledge in virtual viewpoint images by generating a virtual viewpoint image from a captured image, enhancing the understanding and manipulation of the scene.

JP2025100763AActive Publication Date: 2025-07-03CANON KK
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025068136
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-03
Estimated Expiration
2041-02-26

AI Technical Summary

Technical Problem

Existing technologies lack the ability to deepen knowledge about the scene corresponding to a generated figure in virtual viewpoint images, particularly in generating new virtual viewpoint images from a scene of a virtual viewpoint image.

Method used

An information processing apparatus that acquires a captured image of a product generated based on a specific time in a virtual viewpoint image and generates a virtual viewpoint image of that scene using the captured image.

Benefits of technology

Enables the generation of a virtual viewpoint image of a predetermined scene, allowing for deeper understanding and manipulation of the scene corresponding to the generated figure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025100763000001_ABST
    Figure 2025100763000001_ABST
Patent Text Reader

Abstract

To allow for generating a virtual viewpoint image of a given scene corresponding to a modeled object.SOLUTION: An information processing device disclosed herein is configured to: acquire captured image of a generated object generated according to a scene corresponding to a specific time in a scene of a virtual viewpoint image; and generate a virtual viewpoint image of the scene corresponding to the specific time using the captured image of the generated object.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, an information processing method, and a program.

Background Art

[0002] There is a technology for calculating the position and orientation of a mobile terminal based on a captured image obtained by capturing a two-dimensional marker with a mobile terminal such as a smartphone, and generating a virtual viewpoint image according to the calculated position and orientation. This technology is used in various fields, for example, in technologies related to augmented reality (AR). Patent Document 1 discloses a technology for correcting the inclination of a virtual viewpoint image caused by camera shake or the like during imaging of a two-dimensional marker.

[0003] On the other hand, in recent years, using a 3D printer or the like, a figure of a person or the like has also been generated based on a 3D model (shape data representing a three-dimensional shape) of an object obtained by imaging or scanning an actual person or the like.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] Although it is also possible to generate a figure of a person or the like based on a scene of a virtual viewpoint image, there is a demand for deepening knowledge about the generated figure, such as which scene of the original virtual viewpoint image the scene corresponding to the generated figure is. In addition, it is also possible to generate a new virtual viewpoint image corresponding to another virtual viewpoint from a scene of a virtual viewpoint image. In this case as well, there is a demand for deepening knowledge about the generated virtual viewpoint image.

Means for Solving the Problems

[0006] An information processing apparatus according to an aspect of the present disclosure includes: an acquisition unit that acquires a captured image by capturing a product generated based on a scene corresponding to a specific time in a virtual viewpoint image; and a generation unit that generates a virtual viewpoint image of the scene corresponding to the specific time using the captured image of the product.

Advantages of the Invention

[0007] According to the present disclosure, a virtual viewpoint image of a predetermined scene corresponding to a modeled object can be generated.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Modes for Carrying Out the Invention

[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the following embodiments do not limit the present disclosure according to the claims, and not all combinations of the features described in the present embodiments are essential for the solution means of the present disclosure. The same reference numerals are assigned to the same components, and the description thereof will be omitted.

[0010] [Embodiment 1] In this embodiment, a mode of generating a virtual viewpoint image corresponding to the same 3D model as the D model used for generating the product or acquiring information related thereto from a captured image of the product generated based on the same 3D model as the 3D model used for generating the virtual viewpoint image will be described. Note that the virtual viewpoint image is an image generated by an end user and / or a selected operator operating the position and orientation (direction) of a camera (virtual camera) corresponding to the virtual viewpoint, and is also called a free viewpoint image, an arbitrary viewpoint image, or the like. The virtual viewpoint image may be a moving image or a still image, but in this embodiment, the case of a moving image will be described as an example. The 3D model is shape data representing the three-dimensional shape of an object.

[0011] (Configuration of the system) FIG. 1 is a diagram showing a configuration example of an information processing system (virtual viewpoint image generation system) that generates, as a second output, a virtual viewpoint image corresponding to the 3D model of the first output (product) or information related thereto based on a captured image of the first output. FIG. 1(a) shows a configuration example of the information processing system 100, and FIG. 1(b) shows an installation example of a sensor system included in the information processing system. The information processing system (virtual viewpoint image generation system) 100 includes a database 103, an image generation device 104, a mobile terminal 105, and a modeling device 106.

[0012] The database 103 manages event information, 3D model information, and the like. The event information includes data indicating a storage destination of 3D model information for each object associated with the total time code of the event to be imaged. The object may include a person or thing that the user wants to model, or may include a person or thing that is not a modeling target. The 3D model information includes information related to the 3D model of the object.

[0013] These pieces of information may be, for example, information acquired by the sensor system of the information processing system 100, or may be information acquired by a sensor system of another system different from the information processing system 100. The sensor system has, for example, sensor systems 101a - 101n each having at least one imaging device, i.e., a camera, as shown in Fig. 1(b). Hereinafter, unless otherwise specified, the n sensor systems from sensor system 101a to sensor system 101n are not distinguished, and are collectively referred to as a plurality of sensor systems 101. The plurality of sensor systems 101 are installed so as to surround the region 120 which is the imaging target region, and the cameras of the plurality of sensor systems 101 image the region 120 from different directions. The virtual camera 110 images the region 120 from a direction different from those of the cameras of the plurality of sensor systems 101. Details of the virtual camera 110 will be described later.

[0014] When the imaging target is a professional sports game such as rugby or soccer, the region 120 is the stadium field, and the n (for example, 100) plurality of sensor systems 101 are installed so as to surround the field. Also, in the imaging target region 120, there may be not only people on the field but also balls and other objects. Note that the imaging target is not limited to the stadium field, and may be a music live held in an arena or the like, or a CM shooting held in a studio, as long as the plurality of sensor systems 101 can be installed. Note that the number of sensor systems 101 to be installed is not limited. Also, the plurality of sensor systems 101 do not have to be installed all around the region 120, and may be installed only in a part of the periphery of the region 120 depending on restrictions on the installation location and the like. Also, the plurality of cameras of the plurality of sensor systems 101 may include imaging devices with different functions such as a telephoto camera and a wide - angle camera.

[0015] The plurality of cameras of the plurality of sensor systems 101 perform imaging synchronously to acquire a plurality of images. Note that each of the plurality of images may be an imaging image, or may be an image obtained by performing image processing such as processing to extract a predetermined region from the imaging image.

[0016] In addition, each of the sensor systems 101a-101n may have a microphone (not shown) in addition to the camera. The microphones of the plurality of sensor systems 101 each synchronously collect sound. Based on the collected sound, an acoustic signal can be generated that is reproduced together with the display of the image in the image generation device 104. Hereinafter, for the sake of simplicity of explanation, the description of the sound will be omitted, but basically the image and the sound are processed together.

[0017] The plurality of images acquired by the plurality of sensor systems 101 and the time code used for imaging are combined and stored in the database 103. The time code is time information represented by an absolute value for uniquely identifying the imaging time, and is, for example, time information that can be specified in a format such as day:hour:minute:second.frame number.

[0018] An example of a table of event information and 3D model information managed by the database 103 will be described with reference to the drawings. FIG. 2 is a diagram showing an example of a table of information managed by the database 103. FIG. 2(a) shows a table of event information, and FIG. 2(b) shows a table of 3D model information. As shown in FIG. 2(a), the table 210 of event information managed by the database 103 indicates the storage destination of 3D model information for each object for all time codes of the event to be imaged. In the table 210 of event information, for example, it is shown that the storage destination of the 3D model information with the time code "16:14:24.041" and the object being object A is "DataA100". Note that the table 210 of event information is not limited to all time codes, and may indicate the storage destination of 3D model information for some time codes. The time code is an absolute value for uniquely identifying the imaging time, and can be specified, for example, in a format such as "day:hour:minute:second.frame number".

[0019] As shown in Fig. 2(b), the table 220 of 3D model information managed by the database 103 stores data for each item of "3D model", "texture", "average coordinates", "center of gravity coordinates", and "maximum / minimum coordinates". In "3D model", for example, data related to the 3D model itself such as point clouds and meshes is stored. In "texture", data related to the texture image applied to the 3D model is stored. In "average coordinates", data related to the coordinates of the point obtained by averaging all the coordinates of the point cloud constituting the 3D model is stored. In "center of gravity coordinates", data related to the coordinates of the center of gravity point based on all the coordinates of the point cloud constituting the 3D model is stored. In "maximum / minimum coordinates", data related to the coordinates of the maximum / minimum points among the coordinates of the point cloud constituting the 3D model is stored. Note that the items of data stored in the table 220 of 3D model information are not limited to all of "3D model", "texture", "average coordinates", "center of gravity coordinates", and "maximum / minimum coordinates". For example, only "3D model" and "texture" may be used, or other items may be added to these items. Also, the database 103 manages information (not shown) related to the 3D model. The information related to the 3D model includes, for example, information about the date and place of the rugby match and the confrontation card, and information about rugby players if the imaging target is a rugby match.

[0020] By using the information shown in Fig. 2, when a certain time code is specified, it is possible to obtain the 3D model and texture image for each object as the 3D model information at the specified time code. For example, in Fig. 2(a), DataN100 can be obtained as the 3D model information of object N at the time code "16:14:24.041". The 3D model information for other time codes and objects can be obtained in the same way by specifying them. Also, a certain time code can be specified, and the 3D model information of all objects at that time code can be obtained collectively. For example, the time code "16:14:25.021" can be specified, and the 3D model information DataA141, DataB141, ···, DataN141 of all objects can be obtained.

[0021] As described above, an image obtained by imaging with a plurality of sensor systems is called a multi-viewpoint image, and a 3D model representing the three-dimensional shape of an object can be generated from such a multi-viewpoint image. Specifically, from the multi-viewpoint image, a foreground image in which a foreground region corresponding to an object such as a person or a ball is extracted and a background image in which a background region other than the foreground region is extracted are obtained, and a 3D model of the foreground can be generated for each object based on the plurality of foreground images. These 3D models are generated by a shape estimation method such as Visual Hull, for example, and are composed of a point cloud. However, the data format of the 3D model representing the shape of each object is not limited to this. Then, the 3D model generated in this way is recorded in the database 103 for each time code and for each object. Note that the method for generating the 3D model is not limited to this, and it is sufficient that the 3D model is recorded in the database 103.

[0022] Regarding the 3D model of the background object such as a field or a spectator seat, it may be recorded in the database 103 as 3D model information at a time code such as 00:00:00.000.

[0023] Returning to the description of FIG. 1(a). The image generation device 104 acquires a 3D model from the database 103 and generates a virtual viewpoint image based on the acquired 3D model. Specifically, the image generation device 104 acquires an appropriate pixel value from the multi-viewpoint image for each point constituting the acquired 3D model and performs a coloring process. Then, the colored 3D model is arranged in a three-dimensional virtual space, projected onto a virtual camera, and rendered to generate a virtual viewpoint image. In addition, the image generation device 104 generates information regarding the virtual viewpoint image.

[0024] The virtual camera 110 is set, for example, in a virtual space associated with the target region 120 as shown in FIG. 1(b), and can view the region 120 from a viewpoint different from that of any camera of the plurality of sensor systems 101. The viewpoint of the virtual camera 110 is determined by its position and orientation. Details of the position and orientation of the virtual camera 110 will be described later.

[0025] The image generation device 104 may determine the position and orientation of the virtual camera based on the information sent from the mobile terminal 105. Details of this process will be described later with reference to the drawings.

[0026] The virtual viewpoint image is an image representing what is seen from the virtual camera 110, and is also called a free viewpoint video. The virtual viewpoint image generated by the image generation device 104 and information related thereto may be displayed on the image generation device 104, or may be returned as a response to the mobile terminal 105, which is the transmission source of the imaging data, and displayed on the mobile terminal 105. Naturally, the virtual viewpoint image may be displayed on both the image generation device 104 and the mobile terminal 105.

[0027] The modeling device 106 is, for example, a 3D printer or the like, and based on the virtual viewpoint image generated by the image generation device 104, forms a first output 107 such as a human figure (3D model figure). The first output 107 is an output generated based on the 3D model recorded in the database 103, and in this embodiment, is an object imaged by the mobile terminal 105. Examples of the first output include 3D model figures formed using a 3D model by a modeling device 106 such as a 3D printer. Here, the 3D model figure, which is an example of the first output 107, will be described with reference to the drawings.

[0028] FIG. 3 is a diagram for explaining a process of imaging a 3D model figure and generating related information of the 3D model figure based on the captured image. FIG. 3(a) shows an example of a 3D model figure, FIG. 3(b) shows an example of a captured image of the 3D model figure in FIG. 3(a), and FIG. 3(c) shows an example of a virtual viewpoint image corresponding to the 3D model figure in FIG. 3(a). In FIG. 3(a), a 3D model of a rugby game is recorded in the database 103, and an example of a model formed using the 3D model at a certain time code is shown. Specifically, a scene where an offload pass is made as a decisive scene leading to a score during a rugby game is selected, and an example of a model formed using the 3D model associated with that time code is shown.

[0029] As shown in FIG. 3(a), the 3D model figure has a pedestal 301, a first figure body 302, a second figure body 303, and a third figure body 304. The first figure body 302, the second figure body 303, and the third figure body 304 respectively correspond to each player who is an object, and are fixed on the pedestal 301. In FIG. 3(a), the player of the third figure body 304 is holding a ball and is in the 3D model figure at the moment of making an off-road pass.

[0030] In addition, two-dimensional codes 311 and 312 for holding information such as the imaging direction and posture at the time of imaging with the mobile terminal 105 are provided on the front and side surfaces of the pedestal 301. Two two-dimensional codes 311 and 312 are provided on the pedestal 301 of the first output object shown in FIG. 3(a), but the number and position of the two-dimensional codes provided on the first output object are not limited to this. In addition, the information included in the two-dimensional code may include not only the imaging direction and posture, but also image information related to the first output object, such as game information in which the 3D model was imaged. Note that the form of imparting information to the first output object is not limited to the two-dimensional code, and other forms such as watermark information may be used.

[0031] The first output object 107 is not limited to a three-dimensional object such as a 3D model figure, and may be any object generated using the 3D model in the database 103. For example, it may be a shaped object printed on a plate or the like, or a virtual viewpoint image displayed on a display.

[0032] After the image generation device 104 acquires the 3D model in the database 103, it may output the 3D model to the shaping device 106 to shape the first output object 107 such as a 3D model figure.

[0033] Return to the description of FIG. 1. The mobile terminal 105 captures a first output such as a 3D model figure and transmits the captured data to the image generation device 104. The captured data (captured image) may be image data or data including attribute information in the image data. Here, an example of a captured image will be described with reference to FIG. 3(b). FIG. 3(b) shows a state where an image obtained by capturing the 3D model figure shown in FIG. 3(a) as the first output is displayed on the display unit of the mobile terminal. The mobile terminal 105 transmits the captured data of the first output to the image generation device, and as a response thereto, receives and displays the virtual viewpoint image generated by the image generation device 104 and image information related thereto as the second output. Details of these processes will be described later with reference to figures.

[0034] Note that in this embodiment, the case where the virtual viewpoint image is a moving image will be mainly described, but it may also be a still image.

[0035] (Configuration of the Image Generation Device) A configuration example of the image generation device 104 will be described with reference to figures. FIG. 4 is a diagram showing a configuration example of the image generation device 104. FIG. 4(a) shows a functional configuration example of the image generation device 104, and FIG. 4(b) shows a hardware configuration example of the image generation device 104.

[0036] As shown in FIG. 4(a), the image generation device 104 includes an imaging data processing unit 401, a virtual camera control unit 402, a 3D model acquisition unit 403, an image generation unit 404, and an output data control unit 405. Using the above-described functional units, the image generation device 104 generates a virtual viewpoint image generated from the 3D model and image information related thereto as the second output using the captured data obtained by capturing the first output generated from the 3D model. Here, an overview of each function will be described, and details of each process will be described later.

[0037] The imaging data processing unit 401 receives the imaging data obtained by the mobile terminal 105 capturing the first output. Then, the imaging data processing unit 401 acquires the generation information of the first output from the received imaging data. The generation information of the first output is information for identifying the data for generating the first output. For example, it is the identifier of the database in which the 3D model that is the source of the first output is recorded, the time code of the 3D model, etc. The identifier of the database is information for uniquely identifying the corresponding database. Therefore, the 3D model that is the source of the first output can be uniquely identified by this information.

[0038] Also, the imaging data processing unit 401 acquires information such as the position and orientation of the imaging device that captured the first output and the focal length (hereinafter referred to as imaging information) from the imaging data. The position and orientation of the imaging device may be acquired from the two-dimensional codes 311, 312, etc. included in the imaging data, or may be acquired by other methods. Examples of other acquisition methods are shown in Embodiment 2.

[0039] Note that the generation information of the first output and the imaging information that can be acquired from the imaging data are not limited to these. For example, if the first output is related to a sports game, related information such as the game result and the opposing team information may be included.

[0040] The virtual camera control unit 402 controls the virtual camera according to the position and orientation of the imaging device and the focal length acquired via the imaging data processing unit 401. Details of the virtual camera and its position and orientation will be described later with reference to the drawings.

[0041] The 3D model acquisition unit 403 specifies the identifier and time code of the database acquired via the imaging data processing unit 401, and acquires the corresponding 3D model from the database 103. Note that the time code may be specified by a user operation on the mobile terminal 105 or the like.

[0042] The image generation unit 404 generates a virtual viewpoint image based on the 3D model acquired by the 3D model acquisition unit 403. Specifically, for each point constituting the 3D model, the image generation unit 404 acquires an appropriate pixel value from the image and performs a coloring process. Then, the image generation unit 404 arranges the colored 3D model in a three-dimensional virtual space, projects it onto the virtual camera (virtual viewpoint) controlled by the virtual camera control unit 402, and renders it to generate a virtual viewpoint image.

[0043] The virtual viewpoint image generated here is generated using the generation information (information for identifying the 3D model) of the first output object acquired by the imaging data processing unit 401 from the imaging data of the first output object and the imaging information (the position and orientation of the imaging device that captured the first output object, etc.).

[0044] Note that the method for generating the virtual viewpoint image is not limited to this, and various methods such as a method of generating a virtual viewpoint image by projective transformation of a captured image without using a 3D model may be used.

[0045] The output data control unit 405 outputs the virtual viewpoint image generated by the image generation unit 404 as a second output object to an external device, for example, the mobile terminal 105. Further, the related information regarding the first output object acquired by the imaging data processing unit 401 may be output as a second output object to the external device.

[0046] Note that the output data control unit 405 may generate modeling data from the 3D model, output the generated modeling data to the modeling device 106, and obtain the first output object with the modeling device 106.

[0047] (Hardware Configuration of the Image Generation Device) Next, the hardware configuration of the image generation device 104 will be described with reference to FIG. 4(b). As shown in FIG. 4(b), the image generation device 104 includes a CPU 411, a RAM 412, a ROM 413, an operation input unit 414, a display unit 415, and a communication I / F (interface) unit 416.

[0048] The CPU (Central Processing Unit) 411 processes using programs and data stored in the RAM (Random Access Memory) 412 and the ROM (Read Only Memory) 413.

[0049] The CPU 411 performs overall operation control of the image generation device 104 and executes processing for realizing each function shown in Fig. 4(a). Note that the image generation device 104 may have one or more dedicated hardware different from the CPU 411, and at least a part of the processing by the CPU 411 may be executed by the dedicated hardware. Examples of the dedicated hardware include an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), and a DSP (Digital Signal Processor).

[0050] The ROM 413 holds programs and data. The RAM 412 has a work area for temporarily storing programs and data read from the ROM 413. Also, the RAM 412 provides a work area used when the CPU 411 executes each process.

[0051] The operation input unit 414 is, for example, a touch panel, receives input operations by the user, and acquires information input by the received user operations. Examples of the input information include information regarding a virtual camera and information regarding the time code of the virtual viewpoint image to be generated. Note that the operation input unit 414 may be connected to an external controller and receive input information from the user regarding operations. The external controller is, for example, a three-axis controller such as a joystick or an operation device such as a mouse. Note that the external controller is not limited to these.

[0052] The display unit 415 is a touch panel or a screen and displays the generated virtual viewpoint image. In the case of a touch panel, the operation input unit 414 and the display unit 415 are integrated.

[0053] The communication I / F unit 416 transmits and receives information to and from, for example, the database 103, the mobile terminal 105, the modeling device 106, etc. via a LAN or the like. Further, the communication I / F unit 416 may transmit information to an external screen via an image output port corresponding to the following communication standards. Examples of the image output port corresponding to the communication standards include HDMI (Registered Trademark) (High-Definition Multimedia Interface), SDI (Serial Digital Interface), and the like. Also, the communication I / F unit 416 acquires 3D model information and the like from the database 103 via, for example, Ethernet. The communication I / F unit 416 receives imaging data from the mobile terminal 105 and transmits a second output such as a virtual viewpoint image and image information related thereto via, for example, Ethernet or short-range communication.

[0054] (Virtual viewpoint (virtual camera) and its operation screen) The virtual camera 110 and its position and orientation will be described with reference to the drawings. FIG. 5 is a diagram for explaining the virtual camera 110 and its position and orientation. FIG. 5(a) shows a coordinate system, FIG. 5(b) shows an example of a field to which the coordinate system of FIG. 5(a) is applied, FIGS. 5(c) and 5(d) show examples of the drawing area of the virtual camera, and FIG. 5(e) shows an example of the movement of the virtual camera.

[0055] First, a coordinate system representing the three-dimensional space of the imaging target, which serves as a reference when setting virtual specifications, will be described. As shown in FIG. 5(a), in this embodiment, an orthogonal coordinate system representing the three-dimensional space with three axes of the X-axis, Y-axis, and Z-axis is used. This orthogonal coordinate system is set for each object shown in FIG. 5(b), that is, the rugby field 591, the ball 592 existing thereon, the player 593, and the like. Further, it may be set for facilities in the rugby field such as the spectator seats and billboards around the field 591. Specifically, the origin (0, 0, 0) is set at the center of the field 591. Then, the X-axis is set in the long side direction of the field 591, the Y-axis is set in the short side direction of the field 591, and the Z-axis is set in the vertical direction with respect to the field. Note that the directions of the respective axes are not limited to these. Using such a coordinate system, the position and the attitude of the virtual camera 110 are set.

[0056] Subsequently, the drawing range of the virtual camera will be described with reference to the drawings. In the quadrangular pyramid 500 shown in FIG. 5(c), the vertex 501 represents the position of the virtual camera 110, and the vector 502 in the line-of-sight direction with the vertex 501 as the base point represents the attitude of the virtual camera 110. Note that the vector 502 is also called the optical axis vector of the virtual camera. The position of the virtual camera is expressed by the components (x, y, z) of each axis, and the attitude of the virtual camera 110 is expressed by a unit vector with the components of each axis as scalars. The vector 502 representing the attitude of the virtual camera 110 is assumed to pass through the center points of the front clip plane 503 and the rear clip plane 504. The frustum of the virtual camera, which is the projection range (drawing range) of the 3D model, is the space 505 sandwiched between the front clip plane 503 and the rear clip plane 504.

[0057] Next, the component indicating the drawing range of the virtual camera will be described with reference to the drawings. Fig. 5(d) is a view of the virtual viewpoint in Fig. 5(c) seen from above (the Z-axis). The drawing range is determined by the following values. Each value of the distance 511 from the vertex 501 to the front clip plane 503, the distance 512 from the vertex 501 to the rear clip plane 504, and the field angle 513 of the virtual camera 110 may be a preset predetermined value (specified value), or a set value in which the predetermined value is changed by a user operation. Further, the field angle 513 may be a value obtained based on a variable focal length of the virtual camera 110 separately. Note that the relationship between the field angle and the focal length is a general technique and its description is omitted.

[0058] Next, the change in the position of the virtual camera 110 (movement of the virtual viewpoint) and the change in the orientation of the virtual camera 110 (rotation) will be described. The virtual viewpoint can be moved and rotated within the space represented in three-dimensional coordinates. Fig. 5(e) is a diagram for explaining the movement and rotation of the virtual camera. In Fig. 5(e), the dashed arrow 506 represents the movement of the virtual camera (virtual viewpoint), and the dashed arrow 507 represents the rotation of the moved virtual camera (virtual viewpoint). The movement of the virtual camera is represented by the components (x, y, z) of each axis, and the rotation of the virtual camera is represented by Yaw, which is a rotation around the Z-axis, Pitch, which is a rotation around the X-axis, and Roll, which is a rotation around the Y-axis. Since the virtual camera can freely move and rotate in the three-dimensional space of the subject (field) in this way, a virtual viewpoint image in which an arbitrary area of the subject becomes the drawing range can be generated.

[0059] In the present embodiment, as described above, the position, orientation, and focal length of the imaging device are acquired as imaging information from the imaging data of the first output 107 (3D model figure, etc.) imaged by the mobile terminal 105 or the like. Then, control such as moving and rotating the virtual camera is performed according to the acquired values of the position, orientation, and focal length of the imaging device.

[0060] (Generation process of the second output) Next, the generation process of the second output in the image generation device 104 will be described with reference to the drawings. FIG. 6 is a flowchart showing the flow of the generation process of the second output. Based on the imaging data obtained by imaging the first output generated from the shape data (3D model) showing the three-dimensional shape of the object, the virtual viewpoint image generated from the 3D model, or image information related thereto, etc. will be described as the second output with reference to the drawings. This series of processes is realized by the CPU 411 executing a predetermined program to operate each functional unit shown in FIG. 4(a). Hereinafter, steps will be denoted as "S". The same applies to the following description.

[0061] In S601, the image generation unit 404 reads the imaging data obtained by imaging the first output 107. Specifically, the image generation unit 404 acquires the imaging data obtained by imaging the first output 107 with a mobile terminal 105 or the like via the imaging data processing unit 401. The first output 107 is, for example, a 3D model figure output by the modeling device 106 based on the 3D model in the database 103. Examples of the first output 107 include the 3D model figure shown in FIG. 3(a).

[0062] In S602, the image generation unit 404 acquires the generation information (database identifier and time code) of the first output and the imaging information (position and orientation of the imaging device and focal length) from the imaging data acquired in the process of S601. Examples of the imaging data include the image shown in FIG. 3(b). Examples of the method for acquiring the imaging information include the method of acquiring from the two-dimensional code described in the configuration of the above-described image generation device. When the acquired imaging data is the image of FIG. 3(b), the identifier of the database containing the 3D model that is the source of the first output and the time code, and the information regarding the position and orientation of the imaging device are acquired from the two-dimensional code 311 included in this image.

[0063] In S603, the image generation unit 404 acquires the corresponding 3D model from the database 103 by using the identifier and time code of the database included in the generation information of the first output obtained in the process of S602.

[0064] In S604, the image generation unit 404 performs controls such as moving and rotating the virtual camera according to the position, orientation, and focal length of the imaging device included in the imaging information acquired in the process of S602. Then, the image generation unit 404 generates a virtual viewpoint image as the second output by using the controlled virtual camera and the 3D model acquired in S603. Examples of the method for generating the virtual viewpoint image include the method for generating the virtual viewpoint image described in the configuration of the above-described image generation device.

[0065] Here, the virtual viewpoint image, which is an example of the second output, will be described with reference to FIG. 3(c). FIG. 3(c) shows a 3D model associated with the time code obtained from the imaging data and a virtual viewpoint image generated using a virtual camera at the position and pose obtained from the imaging data. In other words, the 3D model figure in FIG. 3(a) (the first output) and the virtual viewpoint image in FIG. 3(c) (the second output) are such that the 3D model associated with the same time code is viewed from the same position and direction. The operation screen 320 has a virtual camera operation area 322 that accepts a user operation for setting the position and pose of the virtual camera 110, and a time code operation area 323 that accepts a user operation for setting the time code. First, the virtual camera operation area 322 will be described. Since the operation screen 320 is displayed on the touch panel, the virtual camera operation area 322 accepts general touch operations 325 such as tap, swipe, pinch-in / out, etc. as user operations. By this touch operation 325, the position and focal length (angle of view) of the virtual camera are adjusted. Also, in the virtual camera operation area 322, touch operations 324 such as continuously pressing each axis of the orthogonal coordinate system are accepted as user operations. By this touch operation 324, the virtual camera 110 rotates around the X-axis, Y-axis, or Z-axis, and the pose of the virtual camera is adjusted. By assigning movement, rotation, and scaling of the virtual camera to each touch operation on the operation screen 320 in this way, the virtual camera 110 can be freely operated. These operation methods are well-known and their description will be omitted.

[0066] Next, the time code operation area 323 will be described. The time code operation area 323 has a main slider 332 and a knob 342, and has a plurality of elements for operating the time code. The time code operation area 323 has an output button 350.

[0067] The main slider 332 is an input element that can perform an operation of setting to a desired time code among all the time codes of the imaging data. When the position of the knob 342 is moved to a desired position by a drag operation or the like, the time code corresponding to the position of the knob 342 is specified. That is, by adjusting the position of the knob 342 of the main slider 332, an arbitrary time code is specified.

[0068] In the process of S604, instead of using the time code included in the generation information of the first output as it is, time codes in a predetermined range from a time code slightly before to a time code slightly after that may be used. In this case, a 3D model associated with the time codes in this predetermined range is acquired. Therefore, a definitive scene modeled as a 3D model figure, which is the first output, can be seen as a series of plays from a time code slightly before that as virtual viewpoint images, which are the second output. The virtual viewpoint images reproduced in this way are shown in the order of FIGS. 3(d)-3(f).

[0069] In FIG. 3(e), it is the same virtual viewpoint image 362 as the virtual viewpoint image in FIG. 3(c) described above, and shows the virtual viewpoint image 362 generated using the 3D model with the same time code (T1) as the 3D model figure which is the first output in FIG. 3(a). In FIG. 3(d), it shows the virtual viewpoint image 361 generated using the 3D model with a time code (T1 - Δt1) slightly before the virtual viewpoint image 362 shown in FIG. 3(e). In FIG. 3(f), it shows the virtual viewpoint image 363 generated using the 3D model with a time code (T1 + Δt1) slightly after the virtual viewpoint image 362 shown in FIG. 3(e). When imaging the 3D model figure which is the first output with a mobile terminal 105 or the like, it can be played back on the mobile terminal 105 or the like in the order of FIGS. 3(d), 3(e), and 3(f) as the virtual viewpoint image of the second output. Of course, there are frames not shown between the virtual viewpoint image 361 shown in FIG. 3(d) and the virtual viewpoint image 363 shown in FIG. 3(f), and the virtual viewpoint images from the time code (T1 - Δt1) to the time code (T1 + Δt1) can be played back at 60 fps or the like. Regarding the width (2Δt1) from the slightly earlier time code (T1 - Δt1) to the slightly later time code (T1 + Δt1) described above, it may be included in the generation information of the first output in advance, or may be specified by a user operation. Regarding the position and orientation of the virtual camera, the position and orientation obtained from the captured image (such as FIG. 3(b)) of the first output may be used for the virtual viewpoint image (FIG. 3(e)) of the time code, and different positions and orientations may be used for the virtual viewpoint images of other time codes. Adjustment of the time code, position, and orientation may be performed, for example, on the mobile terminal 105 by adjusting the position of the time code knob 342 by a user operation or by a tap operation on the output button 350.

[0070] Note that the generated virtual viewpoint image of the second output may be transmitted to the mobile terminal 105 and displayed on the screen of the mobile terminal 105, and in addition to the virtual viewpoint image, information such as a game result may also be displayed as image information regarding the virtual viewpoint image.

[0071] As described above, according to the present embodiment, from the captured image of the output generated based on the same 3D model as the 3D model used in the generation of the virtual viewpoint image, it is possible to output a virtual viewpoint image corresponding to the same 3D model as the 3D model used in the generation of the output or image information related thereto. That is, an imaging image is obtained by imaging a 3D model figure generated based on a scene corresponding to a specific time in the virtual viewpoint image, and using the imaging image of this 3D model figure, a virtual viewpoint image of the scene corresponding to the specific time can be generated.

[0072] For example, when using a 3D model for a rugby game, when imaging a 3D model figure with a mobile terminal, it is possible to display a virtual viewpoint image using the same 3D model at least at the same time code.

[0073] Also, at that time, it is possible to display the virtual viewpoint image viewed from the same direction as the direction in which the 3D model figure was imaged.

[0074] Also, in the above, the case where a corresponding 3D model is specified based on a two-dimensional code obtained from imaging data and a virtual viewpoint image corresponding to the specified 3D model is generated has been described. However, the method for specifying the corresponding 3D model is not limited to the method using a two-dimensional code. For example, feature amounts of objects (for example, the color of a player's uniform and the back number) are extracted in each scene to create a plurality of patterns and managed in advance in a database, and based on the result of image processing such as pattern matching, the corresponding 3D model may be specified. By specifying the corresponding 3D model in this way, it is possible to generate a virtual viewpoint image corresponding to the specified 3D model, or image information related to the virtual viewpoint image.

[0075] [Embodiment 2] In the present embodiment, a mode will be described in which detailed position and orientation are obtained from imaging data obtained by imaging a first output generated based on a 3D model used in the generation of a virtual viewpoint image with a mobile terminal or the like, and a second output is generated using the same 3D model as the position and orientation.

[0076] In this embodiment, since the configuration of the information processing system is the same as that in FIG. 1 and the configuration of the image generation device is the same as that in FIG. 4, the description thereof will be omitted, and the differences will be described.

[0077] The configuration in which the imaging data of the first output object captured by the mobile terminal 105 in Embodiment 1 is processed by the image generation device 104 to generate a second output object based on the information included in the imaging data is the same. Also in this embodiment, the flowchart of the generation process of the second output object can be executed in the same manner as in FIG. 6, but the method of processing the imaging data in S604 is different from that in Embodiment 1.

[0078] In this embodiment, the second output object will be generated using the viewing angle of the imaging data of the first output object captured by the mobile terminal 105. The generation of the second output object reflecting the viewing angle of the imaging data of such a first output object will be described with reference to FIG. 7.

[0079] FIG. 7 is a diagram for explaining a process of capturing a 3D model figure and generating related information of the 3D model figure based on the captured image. FIG. 7(a) shows an example of a 3D model figure which is the same as the first output object example in FIG. 3(a). FIG. 7(b) shows an example of the captured image of the 3D model figure in FIG. 7(a), and FIG. 7(c) shows an example of a virtual viewpoint image in which an object corresponding to the 3D model figure in FIG. 7(a) exists.

[0080] As shown in FIG. 7(a), the 3D model figure has a pedestal 301, a first figure body 302, a second figure body 303, and a third figure body 304. The first figure body 302, the second figure body 303, and the third figure body 304 respectively correspond to each player who is an object, and are fixed on the pedestal 301. On the upper surface of the pedestal 301, in the vicinity of the front, markers 701, 702, and 703 for recognizing coordinates are provided. For convenience, these markers 701 - 703 are called coordinate markers (also called calibration targets, etc.). Note that the coordinate markers 701 - 703 may be visible or invisible such as with watermark information. The number of coordinate markers that can be provided on the first output is not limited to three, and may be less than three or more than three. The location where the coordinate markers 701 - 703 are provided is not limited to the pedestal 301. The coordinate markers 701 - 703 may be provided so as not to be conspicuous on, for example, the uniforms or back numbers of the figure bodies 302 - 304. Also, the coordinate markers 701 - 703 may be embedded so as not to be conspicuous on the field and lines on the pedestal 301. Note that the shape of each of the coordinate markers 701 - 703 is not limited, and any shape that can uniquely identify each of the coordinate markers 701 - 703 is acceptable.

[0081] As shown in FIG. 7(b), the captured image data includes the coordinate markers 701 - 703. Coordinate information can be obtained from these multiple coordinate markers 701 - 703, and the position and coordinates of the mobile terminal 105 at the time of imaging can be accurately calculated using the obtained coordinate information. Note that this calculation method is called camera calibration, etc., and since various methods are known, the description thereof is omitted. Since the number of coordinate markers required for calculation (for example, six points) differs depending on the method, when the required number of coordinate markers is not included in the captured image of the first output by the mobile terminal 105, a warning image may be displayed on the screen of the mobile terminal 105 to notify the imaging of the required number of coordinate markers.

[0082] The coordinate markers 701-703 included in the first output are defined based on the coordinate system shown in FIG. 5(a). Since this is the same coordinate system as the 3D data included in the database 103, the position and orientation of the imaging camera obtained by camera calibration can be acquired in the same coordinate system as the 3D model. Therefore, in the present embodiment, using the position and coordinates obtained from the imaging data and the focal length, the virtual viewpoint image generated can have the same field of view as the imaging mobile terminal 105. The virtual viewpoint image, which is an example of the second output generated in this way, is shown in FIG. 7(c). As can be seen from FIGS. 7(b) and 7(c), the field of view for imaging the 3D model figure (first output) in FIG. 7(b) is the same as the field of view of the virtual viewpoint image (second output) in FIG. 7(c).

[0083] Note that the coordinate markers 701-703 are attached to the first output and are not included in the 3D model recorded in the database 103, so they are not included in the virtual viewpoint image shown in FIG. 7(c), which is the second output.

[0084] Note that instead of using the time code (T2) included in the generation information of the first output as it is, the 3D model may be acquired using the time code from a little before (T2-Δt2) to a little after (T2+Δt2). In that case, virtual viewpoint images for all of those time codes may be displayed with the field of view of the imaging data, or the field of view of the imaging data may be used only with the time code included in the generation information of the first output. In that case, virtual viewpoint images may be generated with different fields of view at the time code a little before (T2-Δt2), and virtual viewpoint images may be generated such that the field of view gradually becomes that of the imaging data towards the time code (T2).

[0085] Subsequently, regarding another example in which the first output in FIG. 7(a) is the imaging target of the mobile terminal 105, it will be described using FIG. 7(d) showing an example of the imaging data and FIG. 7(e) showing an example of the second output generated based on the imaging data in FIG. 7(d).

[0086] As shown in Fig. 7(d), it is imaging data obtained by zooming in on the third figure body 304 included in the 3D model figure which is the first output in Fig. 7(a) using the zoom function of the mobile terminal 105. In the imaging data, if a plurality of coordinate markers (not shown) are included in the third figure body 04, the position and orientation of the imaging camera with the same angle of view as Fig. 7(d) can be calculated using the above-described camera calibration.

[0087] Then, when a virtual camera is controlled according to the calculated position, orientation, and focal length to generate a virtual viewpoint image, the second output shown in Fig. 7(e) can be obtained. As shown in Fig. 7(d) and Fig. 7(e), the angle of view of the imaging data of the first output (3D model figure) in Fig. 7(d) and the second output (virtual viewpoint image) in Fig. 7(e) is the same.

[0088] Note that when processing the imaging data obtained by zooming in on a figure body (athlete) using the zoom function of the mobile terminal 105, a method that does not use the above-described plurality of coordinate markers may also be used.

[0089] For example, in advance, the position and orientation corresponding to the recommended zoom angle of view of the corresponding figure body (athlete) may be used as the zoom angle information of the figure body (athlete), and may be given to the corresponding figure body (athlete) of the first output using a two-dimensional code or watermark information. When the imaging data received by the image generation device 104 from the mobile terminal 105 includes the zoom angle information of the corresponding figure body (athlete), a virtual viewpoint image of the second output in a state where the corresponding figure body (athlete) is zoomed may be generated by controlling a virtual camera based on this. The examples of the imaging data and the generated second output in this case are the same as those in Fig. 7(d) and Fig. 7(e) described above, respectively.

[0090] As described above, according to this embodiment, it is possible to generate and display a second output such as a virtual viewpoint image having the same angle of view as the angle of view of the captured image obtained by capturing the output with, for example, a mobile terminal. For example, when the first output is a 3D model figure, a virtual viewpoint image can be generated as the second output with the same 3D model that generated it and at the angle of view at which it was captured.

[0091] For example, when the first output is a 3D model figure generated using a 3D model targeting a rugby game, when the user captures an image at a preferred angle of view using a mobile terminal, a virtual viewpoint image (second output) at the same angle of view can be played back as a video.

[0092] For example, when the first output is a 3D model figure, if the user zooms in on and captures one of the players included in it, a virtual viewpoint image in the state of having zoomed in on that player can be generated as the second output.

[0093] [Other Embodiments] The present disclosure can also be realized by supplying a program that realizes one or more functions of the above-described embodiment to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (for example, ASIC) that realizes one or more functions.

Explanation of Reference Numerals

[0094] 103 Database 104 Image Generation Device 604 Image Generation Unit

Claims

【Claim 1】 An acquisition unit that acquires a captured image by capturing a product generated based on a scene corresponding to a specific time in a virtual viewpoint image; A generation unit that generates a virtual viewpoint image of a scene corresponding to the specific time using the captured image of the product; An information processing apparatus comprising the above.

Citation Information

Patent Citations

  • I-figure and image processing system using i-figure

    JP2013149106A

  • Medical observation support system and 3-dimensional model of organ

    JP2016168078A

  • Inspection equipment for pier upper work lower surface, inspection system of inspection equipment for pier upper work lower surface and inspection method for pier upper work lower surface

    JP2018151964A

  • Image processing apparatus, image processing method, and program

    JP2017134775A