Information processing device, information processing method, and program
The information processing device captures and processes images of generated products to create virtual viewpoint images, addressing the need to identify and manipulate scenes within generated figures, thereby enhancing knowledge and control over these images.
Patent Information
- Application Number
- JP2021030907
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-02-26
- Publication Date
- 2025-05-08
- Estimated Expiration
- 2041-02-26
AI Technical Summary
There is a need to deepen knowledge about generated figures, such as identifying the scene in the original virtual viewpoint image that corresponds to a generated scene, and to generate new virtual viewpoint images from existing ones while maintaining this knowledge.
An information processing device that captures a product generated based on a scene corresponding to a specific time in a virtual viewpoint image and generates a virtual viewpoint image of the scene using the captured image, allowing for the acquisition and utilization of scene-specific information.
Enables the generation of a virtual viewpoint image of a given scene corresponding to a sculpture, facilitating deeper understanding and manipulation of generated figures and scenes.
Smart Images

Figure 0007672842000001 
Figure 0007672842000002 
Figure 0007672842000003
Abstract
Description
[Technical field]
[0001] The present disclosure relates to an information processing device, an information processing method, and a program. [Background technology]
[0002] There is a technology that calculates the position and attitude of a mobile terminal such as a smartphone based on an image of a two-dimensional marker captured by the mobile terminal, and generates a virtual viewpoint image according to the calculated position and attitude. This technology is used in various fields, for example, in technology related to augmented reality (AR). Patent Document 1 discloses a technology that corrects the tilt of a virtual viewpoint image caused by camera shake when capturing an image of a two-dimensional marker.
[0003] Meanwhile, in recent years, 3D printers and other devices have been used to generate figures of people and other objects based on 3D models (shape data representing three-dimensional shapes) of objects obtained by imaging or scanning real people and other objects. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] JP 2017-134775 A Summary of the Invention [Problem to be solved by the invention]
[0005] Although it is possible to generate figures such as people based on a scene in a virtual viewpoint image, there is a demand for deeper knowledge of the generated figures, such as which scene in the original virtual viewpoint image the generated figure corresponds to, etc. Also, a new virtual viewpoint image corresponding to another virtual viewpoint may be generated from a scene in a virtual viewpoint image, and in this case as well, there is a demand for deeper knowledge of the generated virtual viewpoint image. [Means for solving the problem]
[0006] An information processing device according to one aspect of the present disclosure is characterized in having an acquisition means for capturing an image of a product generated based on a scene corresponding to a specific time in a virtual viewpoint image, and a generation means for generating a virtual viewpoint image of the scene corresponding to the specific time using the captured image of the product. Effect of the Invention
[0007] According to the present disclosure, it is possible to generate a virtual viewpoint image of a predetermined scene corresponding to a structure. [Brief description of the drawings]
[0008] [Figure 1] FIG. 1 shows an example of the configuration of an information processing system. [Diagram 2] A diagram showing an example of information managed by a database [Diagram 3] FIG. 1 shows a first output example and a second output example. [Figure 4] FIG. 1 is a diagram showing an example of the configuration of an image generating device. [Diagram 5] Diagram explaining the virtual camera [Figure 6] A flowchart showing the flow of a process for generating a second output product. [Figure 7] FIG. 1 shows a first output example and a second output example. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the following embodiments do not limit the present disclosure related to the claims, and not all of the combinations of features described in the present embodiments are essential to the solution of the present disclosure. Note that the same components are given the same reference numbers and their description will be omitted.
[0010] [Embodiment 1] In this embodiment, a virtual viewpoint image corresponding to the same 3D model as the D model used in generating the product is generated from a captured image of the product generated based on the same 3D model as the 3D model used in generating the virtual viewpoint image, or information about the virtual viewpoint image is acquired. Note that the virtual viewpoint image is an image generated by an end user and / or a designated operator manipulating the position and attitude (direction) of a camera (virtual camera) corresponding to a virtual viewpoint, and is also called a free viewpoint image or an arbitrary viewpoint image. The virtual viewpoint image may be a video or a still image, but in this embodiment, a video image is taken as an example. The 3D model is shape data that represents the three-dimensional shape of an object.
[0011] (System Configuration) Fig. 1 is a diagram showing a configuration example of an information processing system (virtual viewpoint image generation system) that generates a virtual viewpoint image corresponding to a 3D model of a first output object (product) or information related thereto as a second output object based on a captured image of the first output object (product). Fig. 1(a) shows a configuration example of an information processing system 100, and Fig. 1(b) shows an installation example of a sensor system possessed by the information processing system. The information processing system (virtual viewpoint image generation system) 100 has a database 103, an image generation device 104, a mobile terminal 105, and a modeling device 106.
[0012] The database 103 manages event information, 3D model information, and the like. The event information includes data indicating a storage location of 3D model information for each object associated with all time codes of the event to be imaged. The objects may include people and objects that the user wants to model, or may include people and objects that are not the target of modeling. The 3D model information includes information about the 3D model of the object.
[0013] These pieces of information may be, for example, information acquired by a sensor system included in the information processing system 100, or may be information acquired by a sensor system included in a system different from the information processing system 100. For example, as shown in FIG. 1(b), the sensor system includes sensor systems 101a-101n, each of which includes at least one camera serving as an imaging device. In the following, unless otherwise specified, the n sensor systems from the sensor system 101a to the sensor system 101n are not distinguished from each other and are referred to as a multiple sensor system 101. The multiple sensor system 101 is installed so as to surround an area 120, which is a target area for imaging, and the cameras of the multiple sensor system 101 image the area 120 from different directions. The virtual camera 110 images the area 120 from a direction different from that of the camera of the multiple sensor system 101. Details of the virtual camera 110 will be described later.
[0014] When the imaging target is a professional sports match such as rugby or soccer, the area 120 is a stadium field, and n (e.g., 100) multiple sensor systems 101 are installed to surround the field. The imaging target area 120 may include not only people on the field but also balls and other objects. The imaging target is not limited to the stadium field, and may be a live music performance held in an arena or the like, or a commercial shoot held in a studio, as long as multiple sensor systems 101 can be installed. There is no limit to the number of sensor systems 101 to be installed. The multiple sensor systems 101 do not have to be installed all around the area 120, and may be installed only in a part of the periphery of the area 120 depending on restrictions on the installation location, etc. The multiple cameras of the multiple sensor system 101 may include imaging devices with different functions, such as a telephoto camera and a wide-angle camera.
[0015] The multiple cameras of the multiple sensor system 101 capture images synchronously to obtain multiple images. Each of the multiple images may be a captured image, or may be an image obtained by performing image processing, such as a process of extracting a predetermined area, on the captured image.
[0016] Each of the sensor systems 101a-101n may have a microphone (not shown) in addition to a camera. The microphones of the multiple sensor systems 101 synchronously pick up sound. Based on the picked up sound, an acoustic signal can be generated that is played back together with the display of an image in the image generating device 104. Hereinafter, for the sake of simplicity, a description of sound will be omitted, but it is assumed that images and sound are basically processed together.
[0017] The multiple images acquired by the multiple sensor system 101 and the time codes used for capturing images are combined and stored in a database 103. The time code is time information expressed as an absolute value for uniquely identifying the time when an image was captured, and can be specified in a format such as day:hour:minute:second.frame number.
[0018] An example of a table of event information and 3D model information managed by the database 103 will be described with reference to the drawings. FIG. 2 shows an example of a table of information managed by the database 103, where FIG. 2(a) shows the table of event information and FIG. 2(b) shows the table of 3D model information. As shown in FIG. 2(a), the table 210 of event information managed by the database 103 shows the storage destination of 3D model information for each object for all time codes of the event to be imaged. For example, the table 210 of event information shows that the storage destination of 3D model information in which the time code is "16:14:24.041" and the object is object A is "DataA100". Note that the table 210 of event information is not limited to all time codes, and may show the storage destination of 3D model information for some time codes. The time code is an absolute value for uniquely identifying the time of image capture, and can be specified in a format such as "day:hour:minute:second.frame number".
[0019] As shown in FIG. 2B, the table 220 of the 3D model information managed by the database 103 stores data on each of the items of "3D model", "texture", "average coordinate", "center of gravity coordinate", and "maximum and minimum coordinate". In the "3D model", data on the 3D model itself such as a point cloud or a mesh is stored. In the "texture", data on a texture image to be applied to the 3D model is stored. In the "average coordinate", data on the coordinate of a point obtained by averaging all the coordinates of the point cloud constituting the 3D model is stored. In the "center of gravity coordinate", data on the coordinate of a point that is the center of gravity based on all the coordinates of the point cloud constituting the 3D model is stored. In the "maximum and minimum coordinate", data on the coordinate of the maximum and minimum points among the coordinates of the point cloud constituting the 3D model is stored. Note that the items of data stored in the table 220 of the 3D model information are not limited to all of "3D model", "texture", "average coordinate", "center of gravity coordinate", and "maximum and minimum coordinate". For example, only "3D model" and "texture" may be stored, or other items may be added to these items. Furthermore, the database 103 manages information (not shown) related to the 3D model. If the imaging subject is a rugby match, the information related to the 3D model includes information related to the date and venue of the rugby match, match schedule, information related to the rugby players, and the like.
[0020] By using the information shown in FIG. 2, when a certain time code is specified, a 3D model and texture image for each object can be acquired as 3D model information at the specified time code. For example, in FIG. 2(a), DataN100 can be acquired as 3D model information for object N at time code "16:14:24.041". 3D model information for other time codes and objects can also be acquired by specifying in the same manner. In addition, a certain time code may be specified and 3D model information for all objects at that time code may be acquired all at once. For example, time code "16:14:25.021" may be specified and 3D model information DataA141, DataB141, ..., DataN141 of all objects may be acquired.
[0021] As described above, images obtained by imaging with a multi-sensor system are called multi-view images, and a 3D model representing the three-dimensional shape of an object can be generated from such a multi-view image. Specifically, a foreground image in which a foreground region corresponding to an object such as a person or a ball is extracted, and a background image in which a background region other than the foreground region is extracted are acquired from the multi-view image, and a 3D model of the foreground can be generated for each object based on the multiple foreground images. These 3D models are generated by a shape estimation method such as Visual Hull, and are composed of a point cloud. However, the data format of the 3D model representing the shape of each object is not limited to this. The 3D models generated in this manner are recorded in the database 103 for each time code and for each object. Note that the method of generating the 3D model is not limited to this, and it is sufficient that the 3D model is recorded in the database 103.
[0022] It should be noted that 3D models of background objects such as the field and spectator seats may be recorded in the database 103 as 3D model information at time code 00:00:00.000 or the like.
[0023] Returning to the explanation of FIG. 1(a), the image generating device 104 acquires a 3D model from the database 103, and generates a virtual viewpoint image based on the acquired 3D model. Specifically, the image generating device 104 acquires appropriate pixel values from the multi-viewpoint image for each point constituting the acquired 3D model, and performs coloring processing. The colored 3D model is then placed in a three-dimensional virtual space, projected onto a virtual camera, and rendered to generate a virtual viewpoint image. The image generating device 104 also generates information related to the virtual viewpoint image.
[0024] 1(b), the virtual camera 110 is set in a virtual space associated with the region of interest 120 and can view the region 120 from a viewpoint different from that of any of the cameras in the multi-sensor system 101. The viewpoint of the virtual camera 110 is determined by its position and orientation. The position and orientation of the virtual camera 110 will be described in detail later.
[0025] The image generating device 104 may determine the position and orientation of the virtual camera based on information sent from the mobile terminal 105. The details of this process will be described later with reference to the drawings.
[0026] The virtual viewpoint image is an image that represents the view from the virtual camera 110, and is also called a free viewpoint video. The virtual viewpoint image generated by the image generation device 104 and information related thereto may be displayed on the image generation device 104, or may be returned as a response to the mobile terminal 105, which is another device that is the source of the captured image data, and displayed on the mobile terminal 105. Naturally, the virtual viewpoint image may be displayed on both the image generation device 104 and the mobile terminal 105.
[0027] The molding device 106 is, for example, a 3D printer, and molds a first output object 107, such as a doll (3D model figure), based on the virtual viewpoint image generated by the image generating device 104. The first output object 107 is an output object generated based on a 3D model recorded in the database 103, and in this embodiment, is an object imaged by the mobile terminal 105. An example of the first output object is a 3D model figure molded using a 3D model by the molding device 106, such as a 3D printer. Here, a 3D model figure, which is an example of the first output object 107, will be described with reference to the drawings.
[0028] FIG. 3 is a diagram for explaining a process of capturing an image of a 3D model figure and generating related information of the 3D model figure based on the captured image. FIG. 3(a) shows an example of a 3D model figure, FIG. 3(b) shows an example of a captured image of the 3D model figure of FIG. 3(a), and FIG. 3(c) shows an example of a virtual viewpoint image corresponding to the 3D model figure of FIG. 3(a). FIG. 3(a) shows an example of modeling using a 3D model at a certain time code, in which a 3D model of a rugby game is recorded in the database 103. Specifically, the example shows a scene in which an offload pass is made as a decisive scene that leads to a score during a rugby game, and the scene is modeled using a 3D model associated with the time code.
[0029] As shown in Fig. 3(a), the 3D model figure has a base 301, a first figure body 302, a second figure body 303, and a third figure body 304. The first figure body 302, the second figure body 303, and the third figure body 304 correspond to the respective players, which are objects, and are fixed on the base 301. In Fig. 3(a), the 3D model figure shows the moment when the player of the third figure body 304 has the ball and is about to make an offload pass.
[0030] In addition, two-dimensional codes 311, 312 that hold information such as the imaging direction and posture when the image was captured by the mobile terminal 105 are provided on the front and side of the base 301. Two two-dimensional codes 311, 312 are provided on the base 301 of the first output object shown in FIG. 3(a), but the number and positions of the two-dimensional codes provided on the first output object are not limited to this. Information included in the two-dimensional code may include not only the imaging direction and posture, but also image information related to the first output object, such as information on the match in which the 3D model was captured. Note that the form in which information is provided to the first output object is not limited to a two-dimensional code, and may be other forms such as watermark information.
[0031] The first output object 107 is not limited to a three-dimensional object such as a 3D model figure, and may be any object generated using a 3D model in the database 103. For example, it may be a shaped object printed on a plate or the like, or a virtual viewpoint image displayed on a display.
[0032] The image generating device 104 may obtain the 3D model in the database 103 and then output it to the modeling device 106 to model a first output object 107, such as a 3D model figure.
[0033] Returning to the explanation of FIG. 1, the mobile terminal 105 captures a first output object 107 such as a 3D model figure, and transmits the captured image data to the image generating device 104. The captured image data (captured image) may be image data, or may be image data including attribute information. An example of the captured image will now be described with reference to FIG. 3(b). FIG. 3(b) shows a state in which an image obtained by capturing an image of the 3D model figure shown in FIG. 3(a) as the first output object is displayed on the display unit of the mobile terminal. The mobile terminal 105 transmits the captured image data of the first output object to the image generating device, and in response thereto, receives and displays the virtual viewpoint image generated by the image generating device 104 and image information related thereto as the second output object. Details of these processes will be described later with reference to the figures.
[0034] In this embodiment, the virtual viewpoint image will be mainly described as a moving image, but it may be a still image.
[0035] (Configuration of the image generating device) An example of the configuration of the image generating device 104 will be described with reference to the drawings. Fig. 4 is a diagram showing an example of the configuration of the image generating device 104, in which Fig. 4(a) shows an example of the functional configuration of the image generating device 104, and Fig. 4(b) shows an example of the hardware configuration of the image generating device 104.
[0036] As shown in Fig. 4(a), the image generating device 104 has an imaging data processing unit 401, a virtual camera control unit 402, a 3D model acquisition unit 403, an image generating unit 404, and an output data control unit 405. The image generating device 104 uses the above-mentioned functional units to generate a virtual viewpoint image generated from the 3D model and image information related thereto as a second output object, using imaging data obtained by imaging a first output object generated from the 3D model. Here, an overview of each function will be described, and details of each process will be described later.
[0037] The imaging data processing unit 401 receives imaging data obtained by the mobile terminal 105 capturing an image of the first output product. The imaging data processing unit 401 then acquires generation information of the first output product from the received imaging data. The generation information of the first output product is information for identifying data for generating the first output product. For example, the generation information is an identifier of a database in which a 3D model that is the basis of the first output product is recorded, a time code of the 3D model, and the like. The database identifier is information for uniquely identifying a corresponding database. Therefore, the 3D model that is the basis of the first output product can be uniquely identified by these pieces of information.
[0038] Furthermore, the imaging data processing unit 401 acquires information such as the position and orientation of the imaging device that captured the first output object, and the focal length from the imaging data (hereinafter referred to as imaging information). The position and orientation of the imaging device may be acquired from the two-dimensional codes 311, 312, etc., included in the imaging data, or may be acquired by other methods. Examples of other acquisition methods are described in the second embodiment.
[0039] Note that the generation information of the first output product and the imaging information that can be acquired from the imaging data are not limited to these. For example, if the first output product is related to a sports game, the information may include information related to the game result and the opposing team information.
[0040] A virtual camera control unit 402 controls the virtual camera according to the position, orientation, and focal length of the imaging device acquired via the imaging data processing unit 401. Details of the virtual camera, its position, and its orientation will be described later with reference to the drawings.
[0041] The 3D model acquisition unit 403 specifies the database identifier and the time code acquired via the imaging data processing unit 401, and acquires a corresponding 3D model from the database 103. Note that the time code may be specified by a user operation on the mobile terminal 105, etc.
[0042] The image generating unit 404 generates a virtual viewpoint image based on the 3D model acquired by the 3D model acquiring unit 403. Specifically, the image generating unit 404 acquires appropriate pixel values from the image for each point constituting the 3D model and performs coloring processing. Then, the image generating unit 404 places the colored 3D model in a three-dimensional virtual space, projects the image onto a virtual camera (virtual viewpoint) controlled by the virtual camera control unit 402, and performs rendering to generate a virtual viewpoint image.
[0043] The virtual viewpoint image generated here is generated using generation information of the first output object (information that identifies the 3D model) acquired by the imaging data processing unit 401 from the imaging data of the first output object, and imaging information (the position and orientation of the imaging device that captured the first output object, etc.).
[0044] Note that the method of generating the virtual viewpoint image is not limited to this, and various methods may be used, such as a method of generating a virtual viewpoint image by projective transformation of a captured image without using a 3D model.
[0045] The output data control unit 405 outputs the virtual viewpoint image generated by the image generation unit 404 as a second output to an external device, for example, the mobile terminal 105. In addition, related information regarding the first output obtained by the imaging data processing unit 401 may be output to the external device as a second output.
[0046] Note that the output data control unit 405 may generate modeling data from the 3D model, output the generated modeling data to the modeling device 106, and obtain a first output object in the modeling device 106.
[0047] (Hardware configuration of image generating device) Next, the hardware configuration of the image generating device 104 will be described with reference to Fig. 4(b). As shown in Fig. 4(b), the image generating device 104 has a CPU 411, a RAM 412, a ROM 413, an operation input unit 414, a display unit 415, and a communication I / F (interface) unit 416.
[0048] A CPU (Central Processing Unit) 411 performs processing using programs and data stored in a RAM (Random Access Memory) 412 and a ROM (Read Only Memory) 413 .
[0049] The CPU 411 controls the overall operation of the image generating device 104 and executes processing for realizing each function shown in Fig. 4(a). Note that the image generating device 104 may have one or more pieces of dedicated hardware different from the CPU 411, and at least a part of the processing by the CPU 411 may be executed by the dedicated hardware. Examples of the dedicated hardware include an ASIC (application specific integrated circuit), an FPGA (field programmable gate array), and a DSP (digital signal processor).
[0050] The ROM 413 holds programs and data. The RAM 412 has a work area for temporarily storing programs and data read from the ROM 413. The RAM 412 also provides a work area used by the CPU 411 when it executes each process.
[0051] The operation input unit 414 is, for example, a touch panel, and receives an input operation by a user, and acquires information input by the received user operation. Examples of the input information include information about a virtual camera and information about a time code of a virtual viewpoint image to be generated. The operation input unit 414 may be connected to an external controller and receive input information from a user regarding operations. The external controller is, for example, an operating device such as a three-axis controller such as a joystick or a mouse. The external controller is not limited to these.
[0052] The display unit 415 is a touch panel or a screen, and displays the generated virtual viewpoint image. In the case of a touch panel, the operation input unit 414 and the display unit 415 are integrated into one unit.
[0053] The communication I / F unit 416 transmits and receives information to and from the database 103, the mobile terminal 105, the modeling device 106, and the like, via, for example, a LAN. The communication I / F unit 416 may also transmit information to an external screen via an image output port that corresponds to the following communication standard. Examples of image output ports that correspond to the communication standard include HDMI (registered trademark) (High-Definition Multimedia Interface) and SDI (Serial Digital Interface). The communication I / F unit 416 also acquires 3D model information and the like from the database 103, for example, via Ethernet. The communication I / F unit 416 receives imaging data from the mobile terminal 105, and transmits second output objects such as virtual viewpoint images and image information related thereto, for example, via Ethernet or short-distance communication.
[0054] (Virtual viewpoint (virtual camera) and its operation screen) The virtual camera 110 and its position and orientation will be described with reference to the drawings. Fig. 5 is a diagram for explaining the virtual camera 110 and its position and orientation. Fig. 5(a) shows a coordinate system, Fig. 5(b) shows an example of a field to which the coordinate system of Fig. 5(a) is applied, Fig. 5(c) and Fig. 5(d) show examples of the drawing area of the virtual camera, and Fig. 5(e) shows an example of the movement of the virtual camera.
[0055] First, a coordinate system representing a three-dimensional space of an imaging target, which is a reference for setting a virtual designation, will be described. As shown in FIG. 5(a), in this embodiment, an orthogonal coordinate system is used in which a three-dimensional space is represented by three axes, an X axis, a Y axis, and a Z axis. This orthogonal coordinate system is set for each object shown in FIG. 5(b), that is, a rugby field 591, a ball 592 existing thereon, a player 593, and the like. Furthermore, it may be set for facilities in the rugby field, such as spectator seats and signs around the field 591. Specifically, the origin (0, 0, 0) is set at the center of the field 591. Then, the X axis is set in the long side direction of the field 591, the Y axis is set in the short side direction of the field 591, and the Z axis is set in the vertical direction to the field. Note that the directions of the axes are not limited to these. Using such a coordinate system, the position and attitude of the virtual camera 110 are set.
[0056] Next, the rendering range of the virtual camera will be described with reference to the drawings. In a quadrangular pyramid 500 shown in FIG. 5(c), a vertex 501 represents the position of the virtual camera 110, and a line-of-sight vector 502 with the vertex 501 as the base point represents the orientation of the virtual camera 110. The vector 502 is also called the optical axis vector of the virtual camera. The position of the virtual camera is represented by the components (x, y, z) of each axis, and the orientation of the virtual camera 110 is represented by a unit vector with the components of each axis as scalars. The vector 502 representing the orientation of the virtual camera 110 passes through the center points of the front clip plane 503 and the rear clip plane 504. The view frustum of the virtual camera, which is the projection range (rendering range) of the 3D model, is a space 505 sandwiched between the front clip plane 503 and the rear clip plane 504.
[0057] Next, the components indicating the rendering range of the virtual camera will be explained using the diagram. FIG. 5(d) is a diagram of the virtual viewpoint of FIG. 5(c) viewed from above (Z axis). The rendering range is determined by the following values. The values of the distance 511 from the vertex 501 to the front clipping plane 503, the distance 512 from the vertex 501 to the rear clipping plane 504, and the angle of view 513 of the virtual camera 110 may be a predetermined value (default value) set in advance, or may be a setting value changed by a user operation. The angle of view 513 may be a value calculated based on the focal length of the virtual camera 110 as a variable. Note that the relationship between the angle of view and the focal length is a common technique, and the explanation thereof will be omitted.
[0058] Next, a change in the position of the virtual camera 110 (movement of the virtual viewpoint) and a change in the attitude of the virtual camera 110 (rotation) will be described. The virtual viewpoint can be moved and rotated in a space expressed by three-dimensional coordinates. FIG. 5(e) is a diagram for explaining the movement and rotation of the virtual camera. In FIG. 5(e), a dashed arrow 506 indicates the movement of the virtual camera (virtual viewpoint), and a dashed arrow 507 indicates the rotation of the moved virtual camera (virtual viewpoint). The movement of the virtual camera is expressed by the components of each axis (x, y, z), and the rotation of the virtual camera is expressed by Yaw, which is a rotation around the Z axis, Pitch, which is a rotation around the X axis, and Roll, which is a rotation around the Y axis. In this way, since the virtual camera can freely move and rotate in the three-dimensional space of the subject (field), any area of the subject can be generated as a virtual viewpoint image that becomes the drawing range.
[0059] In this embodiment, as described above, the position, orientation, and focal length of the imaging device are obtained as imaging information from imaging data of the first output object 107 (such as a 3D model figure) captured by the mobile terminal 105, etc. Then, control such as moving and rotating the virtual camera is performed according to the obtained values of the position, orientation, and focal length of the imaging device.
[0060] (Generation of the second output) Next, the generation process of the second output object in the image generating device 104 will be described with reference to the drawings. FIG. 6 is a flow chart showing the flow of the generation process of the second output object. The process of generating a virtual viewpoint image generated from a 3D model or image information related thereto as the second output object based on imaging data obtained by imaging a first output object generated from shape data (3D model) indicating the three-dimensional shape of an object will be described with reference to the drawings. Note that this series of processes is realized by the CPU 411 executing a predetermined program and operating each functional unit shown in FIG. 4(a). Hereinafter, steps are denoted as "S". Note that the same applies to the following explanations.
[0061] In S601, the image generation unit 404 reads imaging data obtained by imaging the first output object 107. Specifically, the image generation unit 404 acquires imaging data obtained by imaging the first output object 107 with the mobile terminal 105 or the like via the imaging data processing unit 401. The first output object 107 is, for example, a 3D model figure output by the molding device 106 based on a 3D model in the database 103. An example of the first output object 107 is a 3D model figure shown in FIG. 3(a).
[0062] In S602, the image generating unit 404 acquires generation information of the first output object (database identifier and time code) and imaging information (position of the imaging device, its orientation, and focal length) from the imaging data acquired in the process of S601. For example, the imaging data may be the image shown in FIG. 3(b). For example, the imaging information may be acquired from the two-dimensional code described in the configuration of the image generating device above. When the acquired imaging data is the image shown in FIG. 3(b), the database identifier and time code containing the 3D model that is the basis of the first output object, and information regarding the position and orientation of the imaging device are acquired from the two-dimensional code 311 included in the image.
[0063] In S603, the image generation unit 404 acquires a corresponding 3D model from the database 103 using the database identifier and the time code included in the generation information of the first output product acquired in the processing of S602.
[0064] In S604, the image generating unit 404 controls the virtual camera to move and rotate according to the position of the imaging device and its attitude and focal length included in the imaging information acquired in the process of S602. Then, the image generating unit 404 generates a virtual viewpoint image as a second output object using the virtual camera after the control and the 3D model acquired in S603. Examples of a method for generating the virtual viewpoint image include the method for generating the virtual viewpoint image described in the configuration of the image generating device described above.
[0065] Here, a virtual viewpoint image, which is an example of the second output object, will be described with reference to FIG. 3(c). FIG. 3(c) shows a 3D model associated with a time code acquired from imaging data, and a virtual viewpoint image generated using a virtual camera with a position and orientation acquired from the imaging data. In other words, the 3D model figure (first output object) in FIG. 3(a) and the virtual viewpoint image (second output object) in FIG. 3(c) are like a 3D model associated with the same time code viewed from the same position and direction. The operation screen 320 has a virtual camera operation area 322 that accepts a user operation for setting the position and orientation of the virtual camera 110, and a time code operation area 323 that accepts a user operation for setting a time code. First, the virtual camera operation area 322 will be described. Since the operation screen 320 is displayed on a touch panel, the virtual camera operation area 322 accepts general touch operations 325 such as tapping, swiping, and pinching in and out as user operations. The position and focal length (angle of view) of the virtual camera are adjusted by this touch operation 325. Furthermore, in virtual camera operation area 322, touch operation 324 such as continuously pressing each axis of a Cartesian coordinate system is accepted as a user operation. This touch operation 324 rotates virtual camera 110 around the X-axis, Y-axis or Z-axis, and the attitude of the virtual camera is adjusted. In this way, by assigning the movement, rotation or enlargement / reduction of the virtual camera to each touch operation on operation screen 320, virtual camera 110 can be freely operated. These operation methods are well known and will not be described here.
[0066] Next, a description will be given of the time code operation area 323. The time code operation area 323 has a main slider 332 and a knob 342, and has multiple elements for operating the time code. The time code operation area 323 has an output button 350.
[0067] The main slider 332 is an input element that can be used to set a desired time code from among all time codes of the imaging data. When the position of the knob 342 is moved to a desired position by a dragging operation or the like, the time code corresponding to the position of the knob 342 is specified. In other words, by adjusting the position of the knob 342 of the main slider 332, an arbitrary time code can be specified.
[0068] In the process of S604, the time code included in the generation information of the first output may not be used as is, but may be used within a predetermined range from a time code slightly before to a time code slightly after. In this case, a 3D model that corresponds to this predetermined range of time codes is obtained. Therefore, the decisive scene modeled as the 3D model figure, which is the first output, can be viewed as a series of plays from a time code slightly before the decisive scene as a virtual viewpoint image, which is the second output. The virtual viewpoint images played in this way are shown in the order of Figures 3(d) to 3(f).
[0069] FIG. 3(e) shows a virtual viewpoint image 362 which is the same as the virtual viewpoint image 362 of FIG. 3(c) described above, and which was generated using a 3D model with the same time code (T1) as the 3D model figure which is the first output object of FIG. 3(a). FIG. 3(d) shows a virtual viewpoint image 361 which was generated using a 3D model with a time code (T1-Δt1) slightly before the virtual viewpoint image 362 shown in FIG. 3(e). FIG. 3(f) shows a virtual viewpoint image 363 which was generated using a 3D model with a time code (T1+Δt1) slightly after the virtual viewpoint image 362 shown in FIG. 3(e). When the 3D model figure which is the first output object is captured by the mobile terminal 105 or the like, it can be reproduced as a virtual viewpoint image of the second output object in the order of FIG. 3(d), FIG. 3(e), and FIG. 3(f) on the mobile terminal 105 or the like. Of course, there are frames not shown between the virtual viewpoint image 361 shown in FIG. 3(d) and the virtual viewpoint image 363 shown in FIG. 3(f), and the virtual viewpoint image from the time code (T1-Δt1) to the time code (T1+Δt1) can be played back at 60 fps or the like. The width (2Δt1) from the slightly earlier time code (T1-Δt1) to the slightly later time code (T1+Δt1) may be included in the generation information of the first output object in advance, or may be specified by a user operation. Regarding the position and orientation of the virtual camera, the position and orientation acquired from the captured image (FIG. 3(b), etc.) capturing the first output object may be used for the virtual viewpoint image of that time code (FIG. 3(e)), and a different position and orientation may be used for the virtual viewpoint images of other time codes. The time code and the position and orientation may be adjusted, for example, by adjusting the position of the time code knob 342 by a user operation on the mobile terminal 105, or by tapping the output button 350.
[0070] The generated virtual viewpoint image of the second output object may be transmitted to the mobile terminal 105 and displayed on the screen of the mobile terminal 105. In addition to the virtual viewpoint image, information related to the virtual viewpoint image, such as the results of a match, may also be displayed.
[0071] As described above, according to this embodiment, a virtual viewpoint image corresponding to the same 3D model as the 3D model used in generating an output object, or image information related thereto, can be output from a captured image of an output object output based on the same 3D model as the 3D model used in generating the virtual viewpoint image. That is, a captured image is obtained by capturing an image of a 3D model figure generated based on a scene corresponding to a specific time in the virtual viewpoint image, and a virtual viewpoint image of the scene corresponding to the specific time can be generated using the captured image of this 3D model figure.
[0072] For example, when a 3D model of a rugby match is used, when the 3D model figure is imaged with a mobile terminal, a virtual viewpoint image using the same 3D model at least at the same time code can be displayed.
[0073] At that time, a virtual viewpoint image can be displayed that is viewed from the same direction as the direction in which the 3D model figure was captured.
[0074] In the above, a case has been described in which a corresponding 3D model is identified based on a two-dimensional code obtained from imaging data, and a virtual viewpoint image corresponding to the identified 3D model is generated, but the method of identifying the corresponding 3D model is not limited to the method using the two-dimensional code. For example, a plurality of patterns may be created by extracting feature amounts of objects in each scene (e.g., the color of a player's uniform or the number on his / her back), and the patterns may be managed in advance in a database, and the corresponding 3D model may be identified based on the results of image processing such as pattern matching. By identifying the corresponding 3D model in this way, a virtual viewpoint image corresponding to the identified 3D model, or image information related to the virtual viewpoint image, may be generated.
[0075] [Embodiment 2] In this embodiment, a detailed position and orientation are obtained from imaging data obtained by imaging a first output object generated based on a 3D model used to generate a virtual viewpoint image using a mobile terminal or the like, and a second output object is generated using a 3D model with the same position and orientation.
[0076] In this embodiment, the configuration of the information processing system is the same as that in FIG. 1, and the configuration of the image generating device is the same as that in FIG. 4, so their description will be omitted and only the differences will be described.
[0077] The configuration described in the first embodiment is the same as that in which the image data of a first output product captured by the mobile terminal 105 is processed in the image generating device 104 to generate a second output product based on information included in the image data. In the present embodiment, the flow chart of the generation process of the second output product can be executed in the same manner as in Fig. 6, but the method of processing the image data in S604 is different from that in the first embodiment.
[0078] In this embodiment, the second output object is generated using the angle of view of the imaging data obtained by imaging the first output object with the mobile terminal 105. The generation of the second output object reflecting the angle of view of the imaging data of the first output object will be described with reference to FIG.
[0079] Fig. 7 is a diagram for explaining a process of capturing an image of a 3D model figure and generating related information of the 3D model figure based on the captured image. Fig. 7(a) shows an example of a 3D model figure, which is an example of a first output object similar to Fig. 3(a). Fig. 7(b) shows an example of a captured image of the 3D model figure of Fig. 7(a), and Fig. 7(c) shows an example of a virtual viewpoint image in which an object corresponding to the 3D model figure of Fig. 7(a) exists.
[0080] As shown in FIG. 7(a), the 3D model figure has a pedestal 301, a first figure body 302, a second figure body 303, and a third figure body 304. The first figure body 302, the second figure body 303, and the third figure body 304 correspond to the respective players, which are objects, and are fixed on the pedestal 301. On the upper surface of the pedestal 301, markers 701, 702, and 703 for recognizing coordinates are provided near the front. For convenience, these markers 701-703 are called coordinate markers (also called calibration targets, etc.). The coordinate markers 701-703 may be visible or invisible due to watermark information, etc. The number of coordinate markers that can be provided to the first output object is not limited to three, and may be less than three or more than three. The location where the coordinate markers 701-703 are provided is also not limited to the pedestal 301. The coordinate markers 701-703 may be, for example, inconspicuously affixed to the uniforms or back numbers of the figures 302-304. The coordinate markers 701-703 may also be inconspicuously embedded in the field, lines, etc. on the base 301. The shape of each of the coordinate markers 701-703 is not limited as long as it allows each of the coordinate markers 701-703 to be uniquely identified.
[0081] As shown in FIG. 7(b), the imaging data includes coordinate markers 701-703. Coordinate information is acquired from the multiple coordinate markers 701-703, and the position and coordinates of the mobile terminal 105 at the time of imaging can be accurately calculated using the acquired coordinate information. This calculation method is called camera calibration or the like, and various techniques are known, so a description thereof will be omitted. Since the number of coordinate markers required for calculation differs depending on the technique (e.g., six points), if the required number of coordinate markers are not included when the first output object is imaged by the mobile terminal 105, a warning image may be displayed on the screen of the mobile terminal 105 to notify the user to image the required number of coordinate markers.
[0082] The coordinates of the coordinate markers 701-703 included in the first output object are defined based on the coordinate system shown in FIG. 5(a). This is the same coordinate system as the 3D data included in the database 103, so the position and orientation of the imaging camera obtained by camera calibration can be acquired in the same coordinate system as the 3D model. Therefore, in this embodiment, the virtual viewpoint image generated using the position and coordinates acquired from the imaging data and the focal length can have the same angle of view as the mobile terminal 105 that captured the image. The virtual viewpoint image, which is an example of the second output object generated in this manner, is shown in FIG. 7(c). As can be seen from FIG. 7(b) and FIG. 7(c), the angle of view at which the 3D model figure (first output object) in FIG. 7(b) is captured is the same as the angle of view of the virtual viewpoint image (second output object) in FIG. 7(c).
[0083] Note that the coordinate markers 701-703 are added to the first output object and are not included in the 3D model recorded in the database 103, and therefore are not included in the second output object, the virtual viewpoint image shown in FIG. 7(c).
[0084] In addition, instead of using the time code (T2) included in the generation information of the first output product as it is, a 3D model may be acquired using a time code (T2-Δt2) slightly before that to a time code (T2+Δt2) slightly after that. In that case, the virtual viewpoint image for all the time codes may be displayed with the angle of view of the imaging data, or the angle of view of the imaging data may be used only with the time code included in the generation information of the first output product. In that case, a virtual viewpoint image may be generated with a different angle of view at the slightly earlier time code (T2-Δt2), and a virtual viewpoint image may be generated that gradually becomes the angle of view of the imaging data toward the time code (T2).
[0085] Next, another example in which the first output object of FIG. 7(a) is the subject of imaging by the mobile terminal 105 will be described with reference to FIG. 7(d) showing an example of the imaging data, and FIG. 7(e) showing an example of a second output object generated based on the imaging data of FIG. 7(d).
[0086] As shown in Fig. 7(d), this is imaging data captured by zooming in on the third figure 304 included in the 3D model figure, which is the first output object in Fig. 7(a), using the zoom function of the mobile terminal 105. If the third figure 304 includes multiple coordinate markers (not shown) in the imaging data, the position and orientation of the imaging camera that will have the same angle of view as Fig. 7(d) can be calculated using the camera calibration described above.
[0087] Then, the virtual camera is controlled according to the calculated position, orientation, and focal length to generate a virtual viewpoint image, and the second output object shown in Fig. 7(e) can be obtained. As shown in Fig. 7(d) and Fig. 7(e), the image data of the first output object (3D model figure) in Fig. 7(d) and the second output object (virtual viewpoint image) in Fig. 7(e) have the same angle of view.
[0088] When processing image data captured by zooming in on a figure (player) using the zoom function of the mobile terminal 105, a method that does not use the multiple coordinate markers described above may be used.
[0089] For example, the position and posture of the figure (player) that provides the recommended zoom angle of view may be set as zoom angle of view information for the figure (player) and may be assigned to the figure (player) of the first output object in advance using a two-dimensional code or watermark information. If the zoom angle of view information for the figure (player) is included in the imaging data received by the image generating device 104 from the mobile terminal 105, a virtual camera may be controlled based on the information to generate a virtual viewpoint image of the second output object in which the figure (player) is zoomed in. In this case, examples of the imaging data and the generated second output object are similar to those shown in Figs. 7(d) and 7(e), respectively.
[0090] As described above, according to this embodiment, it is possible to generate and display a second output object such as a virtual viewpoint image having the same angle of view as the angle of view of a captured image obtained by capturing an output object, for example, with a mobile terminal. For example, if the first output object is a 3D model figure, it is possible to generate a virtual viewpoint image as the second output object with the same 3D model and the angle of view at which the first output object was captured.
[0091] For example, if the first output is a 3D model figure generated using a 3D model of a rugby match, when the user uses a mobile device to capture an image at a preferred angle of view, a virtual viewpoint image (second output) with the same angle of view can be played back as a video.
[0092] For example, if the first output is a 3D model figure and an image of one of the players included in the figure is zoomed in, a virtual viewpoint image zoomed in on that player can be generated as the second output.
[0093] [Other embodiments] The present disclosure can also be realized by a process in which a program for implementing one or more functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that implements one or more functions. [Explanation of symbols]
[0094] 103 Database 104 Image generating device 604 Image Generation Unit
Claims
1. a first acquisition means for acquiring, from an image capturing device, image data obtained by capturing an image of a three-dimensional structure of an object corresponding to a specific scene, and camera parameters indicating a position and an attitude of the image capturing device used to capture the three-dimensional structure; a second acquisition means for acquiring three-dimensional shape data of the object based on identification information for identifying the specific scene specified based on the acquired photographic image data; A generating means for generating a virtual viewpoint image based on the acquired three-dimensional shape data and the acquired camera parameters; 13. An information processing device comprising:
2. the identification information is attached to the three-dimensional object in a visible or invisible form, and is acquired by photographing the three-dimensional object; 2. The information processing apparatus according to claim 1,
3. the identification information includes information indicating a storage location of the three-dimensional shape data, the second acquisition means acquires the three-dimensional shape data by accessing the storage location using the information indicating the storage location.
3. The information processing apparatus according to claim 2.
4. The identification information further includes a time code corresponding to the specific scene, the second acquisition means acquires three-dimensional shape data of the object corresponding to the time code; 4. The information processing apparatus according to claim 3.
5. The second acquisition means acquires three-dimensional shape data of the object corresponding to a predetermined range of time codes including the time code corresponding to the specific scene based on the time code corresponding to the specific scene. the generating means generates a virtual viewpoint image corresponding to the predetermined range of time codes.
5. The information processing apparatus according to claim 4.
6. The identification information is specified based on a two-dimensional code or watermark information attached to the three-dimensional object.
6. The information processing apparatus according to claim 1,
7. The identification information is specified based on an angle of view of the photographing device and a region of the three-dimensional object in a photographed image.
6. The information processing apparatus according to claim 1,
8. The identification information includes coordinate information of an object in a captured image.
6. The information processing apparatus according to claim 1,
9. The information processing apparatus according to claim 1 , wherein the identification information includes information about an angle of view of a photographing device corresponding to the photographed image.
10. The virtual viewpoint image generated by the generating means is output to at least one of a source of the captured image data and an external device.
10. The information processing apparatus according to claim 1,
11. An acquisition means for acquiring photographed image data obtained by photographing a three-dimensional model of an object corresponding to a specific scene; a display control means for controlling display of a virtual viewpoint image generated based on scene information indicating the specific scene identified based on the acquired photographed image data, three-dimensional shape data of the object generated based on a plurality of photographed images, and camera parameters indicating a position and an attitude of an information processing device when the three-dimensional structure is photographed; 13. An information processing device comprising:
12. acquiring, from a photographing device, photographed image data obtained by photographing a three-dimensional object corresponding to a specific scene, and camera parameters indicating a position and an attitude of the photographing device used to photograph the three-dimensional object; acquiring three-dimensional shape data of the object based on identification information for identifying the specific scene specified based on the acquired photographed image data; generating a virtual viewpoint image based on the acquired three-dimensional shape data and the acquired camera parameters; 13. An information processing method comprising:
13. acquiring photographed image data obtained by photographing a three-dimensional object corresponding to a specific scene; a step of controlling a display device to display a virtual viewpoint image generated based on scene information indicating the specific scene identified based on the acquired photographed image data, three-dimensional shape data of the object generated based on a plurality of photographed images, and camera parameters indicating a position and an attitude of an information processing device when the three-dimensional structure was photographed; 13. An information processing method comprising:
14. A program for causing a computer to function as the information processing device according to any one of claims 1 to 11.
Citation Information
Patent Citations
Video compositing apparatus, video compositing program and video compositing system
JP2005250748A
I-figure and image processing system using i-figure
JP2013149106A
Device, method, and program for model railroad appreciation, dedicated display monitor, and scene image data for synthesis
JP2016131342A
Image processing apparatus, image processing method, and program
JP2017134775A
Scene Modification for Augmented Reality using Markers with Parameters
US20160247320A1