Information processing device, information processing method, and program
The information processing device addresses the challenge of sudden object appearance in VR and MR by generating interpolated images that include hidden objects, improving the immersive experience through selective rendering and interpolation techniques.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2024-11-20
- Publication Date
- 2026-06-01
AI Technical Summary
Existing VR and MR technologies using HMDs face challenges in maintaining a high frame rate, leading to sudden appearance of hidden objects due to insufficient interpolation of images when objects become visible.
An information processing device that selects hidden objects in a virtual space and generates interpolated images by combining images of visible and hidden objects, reducing processing load through selective rendering and interpolation.
Smoothly displays the movement of hidden objects as they appear, enhancing the immersive experience by accurately representing objects in virtual spaces.
Smart Images

Figure 2026089256000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, and a program for generating an image.
Background Art
[0002] Technologies related to VR (Virtual Reality) and MR (Mixed Reality), which can realize a highly immersive experience using an HMD (Head Mounted Display) worn on the head, are becoming widespread. The HMD constructs a virtual space that reflects in real time the position and orientation information of the user acquired from the sensors mounted thereon. As a result, an image of a highly immersive virtual space is generated. The image of the virtual space is drawn using CG (Computer Graphics). Therefore, based on the positional relationship between the user and the CG object, the display of the object can be easily changed.
[0003] In Patent Document 1, when a first object is shielded by a second object, the game device erases or makes the second object transparent, thereby making the first object observable.
[0004] On the other hand, in VR and MR, a high frame rate is required to minimize the deviation between the image corresponding to the movement of the body and the displayed image. However, in a mobile terminal type HMD, there are limitations in specifications, and it is easy to cause insufficient frame rate. Based on such a background, a method for solving the insufficient frame rate in a mobile terminal type HMD has been developed.
[0005] In Patent Document 2, a technique for estimating (interpolating) the next frame based on a motion vector image in addition to an RGB image and a depth image representing a virtual space is described.
Prior Art Documents
Patent Documents
[0006] [Patent Document 1] Japanese Patent Publication No. 2021-028840 [Patent Document 2] U.S. Patent Application Publication No. 2023 / 0134355 Specification [Overview of the project] [Problems that the invention aims to solve]
[0007] Patent Document 2 generates an interpolated image, which is the image of the next frame, based on the RGB image, depth image, and motion vector image of an object visible from a virtual viewpoint. Therefore, the movement of objects that are not visible from the virtual viewpoint (i.e., hidden) cannot be estimated, and accurate interpolation of the image of the frame at the moment when a hidden object appears is not possible. As a result, there was a problem in that hidden objects appeared to suddenly pop out.
[0008] The present invention aims to generate interpolated images that more appropriately represent objects in a virtual space. [Means for solving the problem]
[0009] One aspect of the present invention is, A means for selecting one or more objects from among multiple objects placed in a virtual space, Based on a first image which is an image of the virtual space as seen from a reference viewpoint in the first frame, and a second image which is an image of one or more objects as seen from the reference viewpoint in the first frame, the virtual space in the second frame after the first frame A generation means for generating a third image, which is an image of the imaginary space, This is an information processing device characterized by having [a certain feature]. [Effects of the Invention]
[0010] According to the present invention, it is possible to generate interpolated images that more appropriately represent objects in a virtual space. [Brief explanation of the drawing]
[0011] [Figure 1] This is a hardware configuration diagram of the information processing device according to Embodiment 1. [Figure 2] This is a functional configuration diagram of the information processing device according to Embodiment 1. [Figure 3] This is a diagram illustrating the virtual space according to Embodiment 1. [Figure 4] This is a diagram illustrating the first group of images relating to Embodiment 1. [Figure 5] This diagram illustrates the second set of images relating to Embodiment 1. [Figure 6] This is a diagram illustrating the interpolated image according to Embodiment 1. [Figure 7] This is a flowchart of the processing of the information processing device according to Embodiment 1. [Figure 8] This is a flowchart of the process for determining a specific object according to Embodiment 1. [Figure 9] This is a functional configuration diagram of the information processing device according to Embodiment 2. [Figure 10] This is a diagram illustrating the positional information and orientation information according to Embodiment 2. [Figure 11] This is a flowchart of the processing of the information processing device according to Embodiment 2. [Figure 12] This is a flowchart of the determination process for generating a corrected image according to Embodiment 2. [Figure 13] This diagram illustrates the user switching screens according to Embodiment 2. [Modes for carrying out the invention]
[0012] Hereinafter, embodiments according to the present invention will be described with reference to the drawings. The following embodiments do not limit the present invention, and not all combinations of features described in these embodiments are essential for the solution means of the present invention. The configuration of the embodiments can be appropriately modified or changed according to the specifications of the device to which the present invention is applied and various conditions (usage conditions, usage environments, etc.). Also, a configuration may be formed by appropriately combining a part of each of the embodiments described below. In the following embodiments, the same reference numerals are given to the same configurations for explanation.
[0013] <Embodiment 1> FIG. 1 is a diagram showing an example of the hardware configuration of an information processing apparatus 101 in Embodiment 1. FIG. 2 is a functional block diagram showing the functional configuration of the information processing apparatus 101 in Embodiment 1. Also, hereinafter, with reference to FIGS. 3 to 6C, an example of generating an interpolation image of hidden objects in a scene where a plurality of hidden objects appear will be described. Here, for example, in a virtual space where objects of a cat and a dog are hidden behind an object of a wall, a scene where the objects of the cat and the dog appear from behind the object of the wall will be described.
[0014] Hereinafter, the configuration and operation of the information processing apparatus 101 will be described. FIG. 1 is a diagram showing an example of the hardware configuration of the information processing apparatus 101. The information processing apparatus 101 is, for example, a display device (head-mounted display; HMD) that can be worn on the user's head.
[0015] The information processing apparatus 101 generates a virtual space using CG technology. The information processing apparatus 101 arranges a virtual viewpoint (a reference viewpoint when observing the virtual space) and a plurality of objects, etc. in the generated virtual space. Thereafter, the information processing apparatus 101 generates a virtual image, etc. seen from the virtual viewpoint. The information processing apparatus 101 includes a CPU 102, a ROM 103, a RAM 104, a recording unit 105, a communication interface 106, a display unit 107, and an operation unit 108.
[0016] The CPU 102 is a system control unit that controls the entire information processing device 101. Furthermore, the CPU 102 realizes the information processing according to this embodiment by executing the information processing program. ru.
[0017] ROM103 is a read-only memory that stores programs and parameters that do not require modification. ROM103 stores the basic program and initial data, etc.
[0018] RAM104 is memory for temporarily storing input information and calculation results (results in information processing and image processing).
[0019] The recording unit 105 is a device capable of writing and reading various types of information. Specifically, the recording unit 105 is a hard disk or memory card built into or attached to the information processing device 101. Alternatively, the recording unit 105 is a storage medium (such as a memory card, removable disk, or IC card) that can be attached to or removed from the information processing device 101.
[0020] The information processing program according to this embodiment is stored in the recording unit 105. The information processing program is read from the recording unit 105, loaded into the RAM 104, and executed by the CPU 102. The information processing program may also be stored in the ROM 103. The recording unit 105 can also record the necessary data used by the information processing program executed by the CPU 102. The data recorded by the recording unit 105 includes, for example, objects placed in a virtual space and interpolated images generated by the information processing device 101.
[0021] The communication interface 106 is an interface unit that can send and receive data with the operating device.
[0022] The display unit 107 is an electronic display device (such as a liquid crystal display device) mounted on the information processing device 101.
[0023] The operation unit 108 includes an operation device that allows for pointing operations (such as hand tracking) and input of various commands. The operation unit 108 acquires operation instructions and commands from the user. The acquired operation instructions and commands are sent to the CPU 102.
[0024] Figure 2 shows an example of the functional configuration of the information processing device 101 according to Embodiment 1. The information processing device 101 has a determination unit 201, a rendering unit 202, and an interpolation unit 203.
[0025] The determination unit 201 determines (selects) objects that are entirely hidden when viewed from a virtual viewpoint (when viewed from a reference position in a reference direction) among the objects placed in the virtual space as "specific objects" for which individual images are generated. The determination unit 201 stores information such as an ID (identifier) that uniquely defines the determined specific object as internal data of the information processing device 101.
[0026] The determination unit 201 may determine one specific object, or it may determine multiple specific objects. Furthermore, the determination unit 201 does not have to determine a specific object from all objects placed in the virtual space. For example, the determination unit 201 may determine a specific object from among objects that have been set in advance by the user. The determination unit 201 may also determine a specific object based on conditions set in the information processing device 101. For example, a condition may be set that "objects placed as backgrounds are not determined as specific objects even if they are completely hidden."
[0027] The rendering unit 202 performs rendering based on the "position, orientation, and model information" of each object placed in the virtual space and the "position, orientation, and field of view, etc." of the virtual viewpoint. The orientation of the virtual viewpoint indicates the direction of observation of the virtual space from the virtual viewpoint. The field of view of the virtual viewpoint indicates the field of view of the virtual space from the virtual viewpoint. For objects visible from the virtual viewpoint, the rendering unit 202 generates a first RGB image in which each pixel has color information, a first depth image in which each pixel has depth information, and a first motion vector image in which each pixel has motion vector information. Hereinafter, the first RGB image, the first depth image, and the first motion vector image will be collectively referred to as the "first image group." The rendering unit 202 stores the first image group as internal data of the information processing device 101.
[0028] Furthermore, the rendering unit 202 generates a second RGB image in which each pixel has color information, a second depth image in which each pixel has depth information, and a second motion vector image in which each pixel has motion vector information for a specific object viewed from a virtual viewpoint. Hereinafter, the second RGB image, the second depth image, and the second motion vector image will be collectively referred to as the "second image group." The rendering unit 202 stores the second image group as internal data of the information processing device 101.
[0029] The rendering unit 202 may omit (skip) part of the rendering process or add processing to reduce the processing load when generating the second RGB image. For example, the rendering unit 202 may omit shading when calculating the color of a specific object, or reduce the number of polygons of a specific object compared to normal. The rendering unit 202 may also set different resolutions for the first image group and the second image group. For example, if the resolution of the first RGB image is set to 1920 pixels horizontally and 1080 pixels vertically, the resolution of the second RGB image may be set to 960 pixels horizontally and 540 pixels vertically.
[0030] The interpolation unit 203 generates an interpolated image based on the image generated by the rendering unit 202. Based on the first set of images, the interpolation unit 203 estimates the image of the next frame of the virtual space (each object) as seen from the virtual viewpoint and obtains this result as the first interpolated image. The interpolation unit 203 also estimates the image of the next frame of a specific object based on the second set of images and obtains this result as the second interpolated image.
[0031] Furthermore, the interpolation unit 203 combines the first interpolated image and the second interpolated image. Through this process, the interpolation unit 203 generates a third interpolated image that includes the objects visible from the virtual viewpoint and the specific object.
[0032] In this way, the rendering unit 202 and the interpolation unit 203 process together generate a third interpolated image that corresponds to the image of the next frame of the first image more accurately. Therefore, the rendering unit 202 and the interpolation unit 203 together can be considered as a "generation unit" for generating the interpolated image.
[0033] Figure 3 shows an example of a virtual space 301 in which multiple objects are placed in Embodiment 1. The virtual space 301 contains a virtual viewpoint 302, a wall object 303, a dog object 304, and a cat object 305.
[0034] The virtual viewpoint 302 is a virtual viewpoint (a reference point for observing virtual objects) placed in the virtual space 301. The position of the virtual viewpoint 302 can be changed based on the position information of the user wearing the HMD. The orientation (observation direction) of the virtual viewpoint 302 can be changed based on the orientation information of the user wearing the HMD.
[0035] Wall object 303 is an object that resembles a wall.
[0036] Dog object 304 is an object that resembles a dog. Dog object 304 is hidden behind wall object 303 and cannot be seen from virtual viewpoint 302. When dog object 304 moves to position 306 in the next frame, it will appear from behind wall object 303.
[0037] Cat object 305 is an object that resembles a cat. Cat object 305 is hidden behind wall object 303 and cannot be seen from virtual viewpoint 302. When cat object 305 moves to position 307 in the next frame, it will appear from behind wall object 303.
[0038] Therefore, in Embodiment 1, the determination unit 201 determines (selects) as a specific object an object that is obstructed by other objects and cannot be seen from the virtual viewpoint 302, such as the dog object 304 and the cat object 305.
[0039] Figures 4A to 4C show examples of the first set of images of the virtual space (objects) visible from the virtual viewpoint 302.
[0040] Figure 4A shows an example of a first RGB image 401 representing the colors of each object visible from the virtual viewpoint 302. The first RGB image 401 holds color information for each pixel. In the first RGB image 401, the color of the wall object 303 is represented in region 402, but the colors of the dog object 304 and the cat object 305 are not represented.
[0041] Figure 4B shows a first depth image 403 representing the depth of each object visible from the virtual viewpoint 302. The first depth image 403 holds depth information for each pixel. In the first depth image 403, the degree of depth is represented by the intensity of the colors. For example, the distance between the wall object 303 and the virtual viewpoint 302 is shorter than the distance between the background and the virtual viewpoint 302. Therefore, if areas farther from the virtual viewpoint 302 are drawn more intensely, the depth 404 will be drawn lighter than the background.
[0042] Figure 4C shows a first motion vector image 405 representing the motion vectors of each object visible from the virtual viewpoint 302. The first motion vector image 405 holds motion vector information for each pixel. The wall object 303 visible from the virtual viewpoint 302 does not move. Therefore, each pixel in the first motion vector image 405 holds an initial value (for example, (0,0,0)).
[0043] Figures 5A to 5C show examples of a second set of images of a specific object as seen from the virtual viewpoint 302. Specifically, Figures 5A to 5C are images representing the color, depth, and motion vectors of "dog object 304 and cat object 305". Note that in these images, the existence of objects other than the specific object is ignored during rendering.
[0044] Figure 5A shows a second RGB image 501 representing the color of a specific object as seen from a virtual viewpoint 302. The second RGB image 501 holds color information for each pixel. The second RGB image 501 depicts a region 502 representing the color of the dog object 304 and a region 503 representing the color of the cat object 305.
[0045] Figure 5B shows a second depth image 504 representing the depth from the virtual viewpoint 302 to a specific object. The second depth image 504 holds depth information for each pixel. In depth image 504 of 2, the depth 505 of the dog object 304 and the depth 506 of the cat object 305 are represented.
[0046] Figure 5C shows a second motion vector image 507 representing the motion vector of a specific object as seen from the virtual viewpoint 302. The second motion vector image 507 holds motion vector information for each pixel. The dog object 304 moves to position 306 in the next frame. Therefore, the motion vector 508 of the dog object 304 is drawn as a motion vector to the left. The cat object 305 moves to position 307 in the next frame. Therefore, the motion vector 509 of the cat object 305 is drawn as a motion vector to the right.
[0047] Figures 6A to 6C show examples of interpolated images generated by the information processing device 101 according to Embodiment 1.
[0048] Figure 6A shows the first interpolated image 601 of the object as seen from the virtual viewpoint 302. The first interpolated image 601 is the image estimated as the RGB image of the frame following the first RGB image 401. Since there is no movement in the wall object visible from the virtual viewpoint 302, the first interpolated image 601 is the same image as the RGB image of the previous frame (first RGB image 401).
[0049] Figure 6B shows a second interpolated image 602 of a specific object as seen from the virtual viewpoint 302. The second interpolated image 602 is an image estimated as the RGB image of the frame following the second RGB image 501. In the second interpolated image 602, compared to the second RGB image 501 shown in Figure 5A, the region 603 representing the color of the dog object 304 has shifted to the left, and the region 604 representing the color of the cat object 305 has shifted to the right.
[0050] Figure 6C shows a third interpolated image 605 that includes objects visible from the virtual viewpoint 302 and specific objects. The third interpolated image 605 is an RGB image created by combining the first interpolated image 601 and the second interpolated image 602.
[0051] Existing interpolation techniques generate a virtual image containing only objects visible from the virtual viewpoint shown in Figure 4, and then generate an interpolated image like the one shown in Figure 6A based on that virtual image. However, hidden objects are not rendered in the generated interpolated image. Therefore, in a series of videos including the interpolated image, the movement of hidden objects appearing cannot be displayed smoothly.
[0052] On the other hand, the information processing device 101 according to Embodiment 1 generates a third interpolated image by combining a first interpolated image of an object visible from a virtual viewpoint with a second interpolated image of a specific object. As a result, the third interpolated image includes a portion of the hidden object, as shown in Figure 6C. Therefore, in a series of videos including the interpolated images, the movement of the hidden object as it appears can be displayed smoothly.
[0053] Referring to the flowchart in Figure 7, a series of information processing operations performed by the information processing device 101 of Embodiment 1 will be described. The processing shown in the flowchart in Figure 7 is realized when the CPU 102 executes the information processing program according to Embodiment 1. The timing of this processing is not limited. For example, this processing may be performed at specified time intervals, or it may be performed when user operation is detected.
[0054] In step S701, the rendering unit 202 renders the virtual space as seen from the virtual viewpoint, based on the information (position, orientation, model information, etc.) of each object placed in the virtual space. A first set of images representing the interval is generated. The rendering unit 202 stores the first set of images as internal data of the information processing device 101.
[0055] In step S702, the interpolation unit 203 generates a first interpolated image (the image of the next frame of the first RGB image) based on the first set of images generated in step S701. The interpolation unit 203 stores the first interpolated image as internal data of the information processing device 101. Methods for generating the interpolated image include, for example, a method of dividing the image into multiple layers such as objects and background and then estimating, and a block matching method that estimates for each certain image region. Since the methods for generating the interpolated image are known techniques, a detailed explanation is omitted here.
[0056] In step S703, the determination unit 201 performs a process to determine (select) a specific object from among multiple objects placed in the virtual space. Details of the process in step S703 will be described later with reference to the flowchart in Figure 8.
[0057] In step S704, the rendering unit 202 determines whether one or more specific objects were determined in step S703. If it is determined that one or more specific objects were determined, the process proceeds to step S705. If it is determined that no specific objects were determined, the process of this flowchart ends. In this case, for example, the interpolation unit 203 treats the first interpolated image as the third interpolated image.
[0058] In step S705, the rendering unit 202 generates a second set of images representing only the specific object as seen from a virtual viewpoint, based on the information of the specific object (position, orientation, model information, etc.) determined in step S703. The rendering unit 202 stores the second set of images as internal data of the information processing device 101.
[0059] In step S706, the interpolation unit 203 generates a second interpolated image (the image of the next frame of the second RGB image) based on the second group of images generated in step S705. The interpolation unit 203 stores the second interpolated image as internal data of the information processing device 101.
[0060] In step S707, the interpolation unit 203 generates a third interpolated image by combining the first interpolated image and the second interpolated image. The interpolation unit 203 may refer to the first depth image (depth information) and the second depth image during the combining process. The interpolation unit 203 may also generate the third interpolated image by superimposing (overwriting) the first interpolated image onto the second interpolated image. Note that the interpolation unit 203 does not have to "perform the combining process for all pixels". For example, the interpolation unit 203 may identify pixels that retain color from the second RGB image and perform the combining process only on the identified pixels.
[0061] Furthermore, in this embodiment, an interpolated image may be generated after the first and second image groups have been generated. In other words, after the processing in step S701, the processing in steps S703 to S705 may be executed, and then the processing in steps S702, S706, and S707 may be executed.
[0062] The following describes the processing of the determination unit 201 in step S703, referring to the flowchart in Figure 8. In the flowchart in Figure 8, an object that is completely hidden (not even partially visible) in the first RGB image is determined to be a specific object.
[0063] In step S801, the determination unit 201 determines the objects that are placed in the virtual space and have not yet been subjected to processing in steps S802 to S803 (hereinafter, Select the "target object" (referred to as the "target object"). The determination unit 201 acquires information such as the target object's position, depth, and model information. The determination unit 201 may also acquire additional information necessary to determine a specific object (such as information on the target object's size).
[0064] In step S802, the determination unit 201 determines whether the entire target object is hidden by other objects in the first RGB image. Methods for determining object occlusion include comparing with a Z-buffer representing depth from a virtual viewpoint, or using a collision detection function set for each object. Since the occlusion determination method is a known technique, a detailed explanation is omitted here. If it is determined that the entire target object is hidden, the process proceeds to step S803. If it is determined that at least a part of the target object is not hidden, the process proceeds to step S804. Alternatively, instead of determining whether the entire target object is hidden, it may be determined whether a predetermined percentage (e.g., 90%) of the target object is hidden.
[0065] In step S803, the determination unit 201 determines (selects) the target object as a specific object. The determination unit 201 stores information such as an ID that can uniquely identify the target object as internal data of the information processing device 101.
[0066] In step S804, the decision unit 201 determines whether the processes in steps S802 to S803 have been executed for all objects placed in the virtual space. If it is determined that the processes in steps S802 to S803 have not been executed for at least one of the objects, the process returns to step S801. If it is determined that the processes in steps S802 to S803 have been executed for all objects, the process in this flowchart ends.
[0067] According to Embodiment 1, if an object is hidden by another object when viewed from a virtual viewpoint, the information processing device 101 generates a virtual image and an interpolated image of that object. This allows the information processing device 101 to generate an interpolated image in the next frame that shows the virtual space as viewed from the virtual viewpoint, more appropriately (so that the object is visible).
[0068] In Embodiment 1, the information processing device 101 generates a third interpolated image after generating a first interpolated image and a second interpolated image. However, the first and second interpolated images do not necessarily need to be generated. In other words, the information processing device 101 may omit the process of generating the first and second interpolated images and generate the third interpolated image based on the first and second image groups. Also, a specific object is likely to be located behind the object shown in the first image. For this reason, neither the first nor the second image group necessarily needs to include a depth image.
[0069] <Embodiment 2> In Embodiment 1, the information processing device 101 determines an object that is completely hidden as a specific object. Subsequently, the information processing device 101 generates an interpolated image of the specific object. This makes it possible to smoothly display the movement of the hidden object as it appears.
[0070] However, generating a virtual image and an interpolated image of a specific object while that object is hidden increases the processing load on the information processing device 101. Also, if the specific object is not moving, generating an interpolated image of that object is unnecessary. Furthermore, if the movement of the specific object is slow, and the specific object When an object is drawn small on a virtual image, the effect of generating an interpolated image of a specific object is small.
[0071] Therefore, in Embodiment 2, the information processing device 101 determines whether or not to generate an interpolated image of a specific object based on the position of the specific object and the position and orientation of the virtual viewpoint. This reduces the frequency of generating the virtual image and interpolated image of the specific object, thereby reducing the processing load.
[0072] Figure 9 is a functional block diagram showing an example of the functional configuration of the information processing device 101 in Embodiment 2. In Embodiment 2, in addition to the determination unit 201, rendering unit 202, and interpolation unit 203, the information processing device 101 also includes a second position acquisition unit 904, a first position acquisition unit 905, an attitude acquisition unit 906, and an interpolation determination unit 907.
[0073] The second position acquisition unit 904 acquires the position information of the "specific object determined by the determination unit 201". The position information of the specific object is 3D coordinate information in the virtual space. Based on the position information of the specific object, the second position acquisition unit 904 stores the position information of the specific object in the current frame and the position information of the previous frame.
[0074] Referring to Figure 10, an example of position information acquired by the second position acquisition unit 904 is shown. Assume that in the virtual space 1001 of the previous frame, a specific object, cat object 1002, is located at position (x11, y11, z11), and a virtual viewpoint 1003 is located at position (x21, y21, z21). Assume that in the virtual space 1004 of the current frame, a specific object, cat object 1005, is located at position (x12, y12, z12), and a virtual viewpoint 1006 is located at position (x22, y22, z22). If the position information of cat object 1005 is acquired in the current frame, the second position acquisition unit 904 stores the ID of the specific object, the position information of the current frame, and the position information of the previous frame, as shown in information 1007.
[0075] The first position acquisition unit 905 acquires position information of a virtual viewpoint placed in the virtual space. The first position acquisition unit 905 acquires the position information of the virtual viewpoint as three-dimensional coordinates. The first position acquisition unit 905 then stores the position information of the current frame and the position information of the previous frame. If the position information of the virtual viewpoint 1006 is acquired in the current frame, the first position acquisition unit 905 stores the ID of the virtual viewpoint, the position information of the current frame, and the position information of the previous frame, as shown in information 1008.
[0076] The attitude acquisition unit 906 acquires attitude information of a virtual viewpoint placed in the virtual space. The attitude information acquired here is the tilt of the three axes (roll, pitch, and yaw) relative to a specific direction in the virtual space. Each of the three axes can be expressed as 0 to 360 degrees, etc. The attitude acquisition unit 906 stores the attitude information of the virtual viewpoint for the current frame and the attitude information of the previous frame.
[0077] For example, suppose the pose information of virtual viewpoint 1003 in the previous frame is (r21, p21, y21), and the pose information of virtual viewpoint 1006 in the current frame is (r22, p22, y22). If the pose information of virtual viewpoint 1006 is acquired in the current frame, the pose acquisition unit 906 stores the ID of the virtual viewpoint, the pose information of the current frame, and the pose information of the previous frame, as shown in information 1009.
[0078] The interpolation determination unit 907 determines whether or not to generate an interpolated image of a specific object based on the position information and orientation information acquired by the "second position acquisition unit 904, first position acquisition unit 905, and orientation acquisition unit 906". The interpolation determination unit 907 stores the determination result as a string ("generate", "do not generate") or a flag (True, False). Only if it is determined that an interpolated image of a specific object should be generated, the rendering unit 202 generates an interpolated image of the specific object. The interpolation unit 203 generates an interpolated image of a specific object after generating a virtual image of the object. Therefore, if it is determined that an interpolated image of a specific object should not be generated, the interpolation unit 203 generates the image of the virtual space for the next frame based only on the first set of images, and not on the second set of images.
[0079] The series of information processing operations of the information processing device 101 according to Embodiment 2 will be explained with reference to the flowchart in Figure 11. The flowchart in Figure 11 is the flowchart shown in Figure 7 with steps S1105 and S1106 added.
[0080] In step S1105, the interpolation determination unit 907 sets a flag (determination flag) indicating whether or not to generate an interpolated image of a specific object. Here, if the interpolation determination unit 907 determines that an interpolated image of the specific object should be generated, it sets the determination flag to True. If the interpolation determination unit 907 determines that an interpolated image of the specific object should not be generated, it sets the determination flag to False. Details of the process in step S1105 will be described later with reference to the flowchart in Figure 12.
[0081] In step S1106, the rendering unit 202 determines whether or not to generate a virtual image of the specific object, according to the determination flag set in step S1105. If the determination flag is True, the rendering unit 202 determines to generate a virtual image of the specific object. If the determination flag is False, the rendering unit 202 determines not to generate a virtual image of the specific object. If it is determined to generate a virtual image of the specific object, the process proceeds to step S705. If it is determined not to generate a virtual image of the specific object, the processing of this flowchart ends.
[0082] The process in step S1105 will be explained by referring to the flowchart in Figure 12. In the flowchart in Figure 12, it is determined that an interpolated image of a specific object will be generated in the following four cases: The first case is when the change in the position of the specific object between the previous frame and the current frame is greater than or equal to the threshold ThA. The second case is when the change in the position of the virtual viewpoint between the previous frame and the current frame is greater than or equal to the threshold ThB. The third case is when the change in the pose (observation direction) of the virtual viewpoint between the previous frame and the pose of the current frame is less than or equal to the threshold ThC. The fourth case is when the distance between the specific object and the virtual viewpoint is within the threshold ThD.
[0083] In step S1201, the second position acquisition unit 904 acquires the position information of a specific object determined by the determination unit 201. The second position acquisition unit 904 retains the position information it already held as the position information of the previous frame, and retains the newly acquired position information as the position information of the current frame.
[0084] In step S1202, the interpolation determination unit 907 calculates the change in the position of a specific object based on the position information of the current frame and the position information of the previous frame obtained in step S1201. The change in the position of the specific object is calculated as the distance between two points in three-dimensional space. The change in the position of the specific object may be calculated using, for example, a formula for the distance between two points as shown in the following equation.
number
[0085] In step S1203, the interpolation determination unit 907 determines whether the amount of change in the position of the specific object calculated in step S1202 is greater than or equal to the threshold ThA. The threshold ThA may be a fixed value (e.g., 0.1m) or it may be user-configurable. This is also acceptable. If it is determined that the amount of change in the position of the specific object is less than the threshold ThA, proceed to step S1204. If it is determined that the amount of change in the position of the specific object is greater than or equal to the threshold ThA, the movement of the specific object is considered to be fast, so proceed to step S1213.
[0086] In step S1204, the first position acquisition unit 905 acquires position information of the virtual viewpoint. The first position acquisition unit 905 retains the position information it already possesses as the position information of the previous frame, and retains the newly acquired position information as the position information of the current frame.
[0087] In step S1205, the interpolation determination unit 907 calculates the change in the position of the virtual viewpoint based on the position information of the virtual viewpoint in the current frame and the position information of the previous frame obtained in step S1204. The change in the position of the virtual viewpoint is calculated as the distance between two points in three-dimensional space, similar to step S1202.
[0088] In step S1206, the interpolation determination unit 907 determines whether the amount of change in the position of the virtual viewpoint calculated in step S1205 is greater than or equal to the threshold ThB. The threshold ThB may be a fixed value, similar to the threshold ThA, or it may be set by the user. If it is determined that the amount of change in the position of the virtual viewpoint is less than the threshold ThB, the process proceeds to step S1207. If it is determined that the amount of change in the position of the virtual viewpoint is greater than or equal to the threshold ThB, the process proceeds to step S1213, as it is considered that the movement of the specific object is fast.
[0089] In step S1207, the posture acquisition unit 906 acquires posture information of the virtual viewpoint. The posture acquisition unit 906 retains the posture information it already possesses as the posture information of the previous frame, and retains the newly acquired posture information as the posture information of the current frame.
[0090] In step S1208, the interpolation determination unit 907 calculates the change in the pose of the virtual viewpoint based on the pose information of the current frame and the pose information of the previous frame of the virtual viewpoint acquired in step S1207. The change in the pose of the virtual viewpoint is, for example, the sum of the absolute values of the differences for each axis. The change in the pose of the virtual viewpoint can be calculated, for example, using the formula shown below.
number
[0091] In step S1209, the interpolation determination unit 907 determines whether the change in the position of the virtual viewpoint calculated in step S1208 is less than or equal to the threshold ThC. The threshold ThC may be a fixed value (e.g., 90 degrees) or it may be set by the user, similar to the threshold ThA. If it is determined that the change in the position of the virtual viewpoint is greater than the threshold ThC, the process proceeds to step S1210. If it is determined that the change in the position of the virtual viewpoint is less than or equal to the threshold ThC, the process proceeds to step S1213.
[0092] In step S1210, the interpolation determination unit 907 calculates the distance between the specific object and the virtual viewpoint in the current frame based on the position information of the specific object obtained in step S1201 and the position information of the virtual viewpoint obtained in step S1204. The distance between the specific object and the virtual viewpoint may be calculated as the distance between two points in three-dimensional space, similar to step S1202.
[0093] In step S1211, the interpolation determination unit 907 determines whether the distance between the specific object and the virtual viewpoint calculated in step S1210 is less than or equal to the threshold ThD. The threshold ThD may be a fixed value (e.g., 5.0m) or it may be user-configurable, similar to the threshold ThA. It may also be possible. If it is determined that the distance between the specific object and the virtual viewpoint is greater than the threshold ThD, proceed to step S1212. If it is determined that the distance between the specific object and the virtual viewpoint is less than or equal to the threshold ThD, the specific object is expected to be drawn larger on the virtual image, so proceed to step S1213.
[0094] In step S1212, the interpolation determination unit 907 determines that an interpolated image of the specific object should not be generated and sets the determination flag to False. The interpolation determination unit 907 stores the determination flag as internal data of the information processing device 101.
[0095] In step S1213, the interpolation determination unit 907 determines that an interpolated image of the specific object should be generated and sets the determination flag to True. The interpolation determination unit 907 stores the determination flag as internal data of the information processing device 101.
[0096] Note that the order of processing in the flowchart in Figure 12 may be changed. That is, the order of processing in "Steps S1201 to S1203", "Steps S1204 to S1206", "Steps S1207 to S1209", and "Steps S1210 to S1211" may be changed.
[0097] In the flowchart of Figure 12, the judgment flag is set to False if none of the above four cases apply, such as when the change in the position of a specific object between the previous frame and the current frame is greater than or equal to the threshold ThA. However, the judgment flag may also be set to False if none of the four cases apply. In addition, the interpolation judgment unit 907 may evaluate the probability of the specific object appearing in the next frame, and set the judgment flag to True if the probability is greater than a certain threshold, and to False otherwise. For example, the interpolation judgment unit 907 evaluates (estimates) the probability of the specific object appearing in the next frame based on the change in the position of the specific object, the change in the position of the virtual viewpoint, and the change in the pose of the virtual viewpoint. In this case, for example, the interpolation judgment unit 907 increases the probability of the specific object appearing in the next frame the greater the change in the position of the specific object. For example, the interpolation judgment unit 907 increases the probability of the specific object appearing in the next frame the greater the change in the position of the virtual viewpoint. Furthermore, any other method may be used if it is possible to evaluate the likelihood that a specific object will appear in the next frame (i.e., that the specific object will become visible from the virtual viewpoint).
[0098] Furthermore, whether or not to generate interpolated images of specific objects may be switchable by user operation (instruction), or the system may switch it depending on the processing load. Below, we will describe the screen when the user can switch whether or not to generate interpolated images of specific objects.
[0099] Here, we will explain assuming a virtual space 301 as shown in Figure 3. Figure 13A shows an example of screen transitions when the user has selected not to generate an interpolated image of a specific object. Figure 13B shows an example of screen transitions when the user has selected to generate an interpolated image of a specific object. Each time the user presses a designated button on the controller, the displayed screen switches between the screen in Figure 13A and the screen in Figure 13B. Note that, regardless of the user's selection, if the judgment flag is set to False, the screen in Figure 13A may be displayed, and if the judgment flag is set to True, the screen in Figure 13B may be displayed.
[0100] If the user chooses not to generate interpolated images for a specific object, then, as shown in Figure 13A, frame images 1408-1410 will be displayed in this order as time progresses. Frame image 1409 is the first interpolated image. Therefore, the specific objects, dog object 304 and cat object 305, are not drawn in frame image 1409. In addition, information (for example, the text "Smoothness Priority: OFF") is superimposed on each frame image to indicate that an interpolated image of the specific object is not being generated (frame image 1409 is not based on the second image).
[0101] If the user has selected to generate interpolated images of specific objects, then, as shown in Figure 13B, frame images 1411 to 1413 will be displayed in this order over time. Frame image 1412 is the third interpolated image. Therefore, frame image 1412 will depict parts of the specific objects, the dog object 304 and the cat object 305. In addition, each frame image will have information superimposed on it indicating that the system is in a state where interpolated images of specific objects are being generated (for example, the text "Smoothness Priority: ON").
[0102] According to Embodiment 2, a third interpolated image is not generated based on the second group of images unless there are specific circumstances. Therefore, when the effect of displaying a specific object in the interpolated image is low, the processing load of the information processing device 101 can be reduced.
[0103] Furthermore, in the above, "If A is greater than or equal to B, proceed to step S1; if A is less than (lower than) B, proceed to step S2" may be rephrased as "If A is greater than (higher than) B, proceed to step S1; if A is less than or equal to B, proceed to step S2." Conversely, "If A is greater than (higher than) B, proceed to step S1; if A is less than or equal to B, proceed to step S2" may be rephrased as "If A is greater than or equal to B, proceed to step S1; if A is less than (lower than) B, proceed to step S2." Therefore, as long as no contradiction arises, "greater than or equal to A" may be rephrased as "greater than (higher; longer; more) than A," and "less than or equal to A" may be rephrased as "less than (lower; shorter; fewer) than A." And "greater than (higher; longer; more) than A" may be rephrased as "greater than or equal to A," and "less than (lower; shorter; fewer) than A" may be rephrased as "less than or equal to A."
[0104] The various controls described above may or may not be performed by a single piece of hardware (e.g., a processor or circuit). Multiple pieces of hardware (e.g., multiple processors, multiple circuits, or a combination of one or more processors and one or more circuits) may share the processing to control the entire device.
[0105] Furthermore, the above-mentioned processors are processors in a broad sense, including general-purpose processors and specialized processors. General-purpose processors include, for example, CPUs (Central Processing Units), MPUs (Micro Processing Units), and DSPs (Digital Signal Processors). Specialized processors include, for example, GPUs (Graphics Processing Units), ASICs (Application Specific Integrated Circuits), and PLDs (Programmable Logic Devices). Programmable logic devices include, for example, FPGAs (Field Programmable Gate Arrays) and CPLDs (Complex Programmable Logic Devices).
[0106] Furthermore, although embodiments of the present invention have been described in detail, the present invention is not limited to these specific embodiments, and various forms that do not depart from the spirit of the invention are also included in the present invention. Moreover, each of the embodiments described above is merely one embodiment of the present invention, and it is possible to combine each embodiment as appropriate.
[0107] <Other Embodiments> The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit that implements one or more functions.
[0108] The above-disclosed embodiments include the following configurations, methods, and programs. (Composition 1) A means for selecting one or more objects from among multiple objects placed in a virtual space, A generation means that generates a third image, which is an image of the virtual space in a second frame following the first frame, based on a first image, which is an image of the virtual space as seen from a reference viewpoint in the first frame, and a second image, which is an image of one or more objects as seen from the reference viewpoint in the first frame. An information processing device characterized by having the following features. (Configuration 2) The selection means selects one or more objects from the plurality of objects that are not partially visible in the first image. The information processing device according to configuration 1, characterized by the above. (Composition 3) The first image is an RGB image of the virtual space, The second image is an RGB image of one or more of the objects, The generation means generates the third image based on the depth information and motion vector information of the objects captured in the first image in the first frame, and the depth information and motion vectors of one or more objects in the first frame. An information processing device according to configuration 1 or 2, characterized by the above. (Composition 4) The generating means is Based on the first image and the depth information and motion vector information of the objects captured in the first image in the first frame, a first interpolated image, which is an image of the virtual space in the second frame, is generated. Based on the second image and the depth information and motion vectors of the one or more objects in the first frame, a second interpolated image is generated, which is an image of the one or more objects in the second frame. Based on the first interpolated image and the second interpolated image, the third image is generated. The information processing apparatus according to configuration 3, characterized by the features described herein. (Composition 5) The generation means generates the third image based on the first image, without relying on the second image, unless otherwise specified. An information processing device according to any one of configurations 1 to 4, characterized by the above. (Composition 6) The aforementioned specific case is based on at least one of the following: the position of one or more objects, the viewing direction from the reference viewpoint, the position of the reference viewpoint, user operation, and the processing load of the information processing device. The information processing apparatus according to configuration 5, characterized by the features described herein. (Composition 7) In the aforementioned specific case, the amount of change in the position of one or more objects is greater than a first threshold. The information processing device according to configuration 6, characterized by the features described therein. (Composition 8) The aforementioned specific case is when the amount of change in the position of the reference viewpoint is greater than the second threshold. The information processing device according to configuration 6, characterized by the features described therein. (Composition 9) In the aforementioned specific case, the amount of change in the observation direction from the aforementioned reference viewpoint is less than the third threshold. The information processing device according to configuration 6, characterized by the features described therein. (Composition 10) In the aforementioned specific case, the distance between the reference position and the one or more objects is less than a fourth threshold. The information processing device according to configuration 6, characterized by the features described therein. (Composition 11) The generation means superimposes information on the third image whether or not the third image is based on the second image onto the third image. An information processing apparatus according to any one of configurations 5 to 10, characterized by the above. (method) A selection step in which one or more objects are selected from multiple objects placed in a virtual space, A generation step of generating a third image, which is an image of the virtual space in a second frame following the first frame, based on a first image, which is an image of the virtual space as seen from a reference viewpoint in the first frame, and a second image, which is an image of one or more objects as seen from the reference viewpoint in the first frame. An information processing method characterized by having the following features. (program) A program for causing a computer to function as one of the means of an information processing device described in any of configurations 1 to 11. [Explanation of Symbols]
[0109] 101: Information processing device, 201: Decision unit, 202: Rendering section, 203: Interpolation section
Claims
1. A selection means for selecting one or more objects from among multiple objects placed in a virtual space, A generation means that generates a third image, which is an image of the virtual space in a second frame following the first frame, based on a first image, which is an image of the virtual space as seen from a reference viewpoint in the first frame, and a second image, which is an image of one or more objects as seen from the reference viewpoint in the first frame. An information processing device characterized by having the following features.
2. The selection means selects one or more objects from the plurality of objects that are not partially visible in the first image. The information processing apparatus according to feature 1.
3. The first image is an RGB image of the virtual space, The second image is an RGB image of one or more of the objects, The generation means generates the third image based on the depth information and motion vector information of the objects captured in the first image in the first frame, and the depth information and motion vectors of one or more objects in the first frame. The information processing apparatus according to feature 1.
4. The generating means is Based on the first image and the depth information and motion vector information of the objects captured in the first image in the first frame, a first interpolated image, which is an image of the virtual space in the second frame, is generated. Based on the second image and the depth information and motion vectors of the one or more objects in the first frame, a second interpolated image is generated, which is an image of the one or more objects in the second frame. Based on the first interpolated image and the second interpolated image, the third image is generated. The information processing apparatus according to claim 3.
5. The generation means generates the third image based on the first image, without relying on the second image, unless otherwise specified. The information processing apparatus according to feature 1.
6. The aforementioned specific case is based on at least one of the following: the position of one or more objects, the viewing direction from the reference viewpoint, the position of the reference viewpoint, user operation, and the processing load of the information processing device. The information processing apparatus according to feature 5.
7. In the aforementioned specific case, the amount of change in the position of one or more objects is greater than a first threshold. The information processing apparatus according to feature 6.
8. The aforementioned specific case is when the amount of change in the position of the reference viewpoint is greater than the second threshold. The information processing apparatus according to feature 6.
9. In the aforementioned specific case, the amount of change in the observation direction from the aforementioned reference viewpoint is less than the third threshold. The information processing apparatus according to feature 6.
10. In the aforementioned specific case, the distance between the reference position and the one or more objects is less than a fourth threshold. The information processing apparatus according to feature 6.
11. The generation means superimposes information on the third image whether or not the third image is based on the second image onto the third image. The information processing apparatus according to feature 5.
12. A selection step in which one or more objects are selected from multiple objects placed in a virtual space, A generation step of generating a third image, which is an image of the virtual space in a second frame following the first frame, based on a first image, which is an image of the virtual space as seen from a reference viewpoint in the first frame, and a second image, which is an image of one or more objects as seen from the reference viewpoint in the first frame. An information processing method characterized by having the following features.
13. A program for causing a computer to function as one of the means of an information processing apparatus according to any one of claims 1 to 11.