Information processing apparatus, information processing method
The information processing apparatus addresses display delays in AR technology by correcting virtual object images based on real-time recognition of real object movements, thereby improving user immersion and comfort.
Patent Information
- Application Number
- JP2021543075
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-08-30
- Filing Date
- 2020-08-28
- Publication Date
- 2025-06-11
- Estimated Expiration
- 2040-08-28
AI Technical Summary
In AR technology, display delays of 3D objects occur not only due to changes in the position and orientation of the display device, but also due to the movement of real objects, particularly moving targets, leading to user discomfort and reduced immersion.
An information processing apparatus that performs first and second recognition processing on the position and orientation of a real object, and uses this information to perform first and second rendering processing on associated virtual objects. The apparatus corrects the image of the virtual object obtained from the first rendering processing based on the result of the second recognition processing before the second rendering processing is completed, allowing for immediate output of the corrected image.
This approach significantly reduces display delays of virtual objects by ensuring that their position and orientation align with the real object's changes in real-time, enhancing user immersion and comfort in the AR environment.
Smart Images

Figure 0007690883000001 
Figure 0007690883000002 
Figure 0007690883000003
Abstract
Description
Technical Field
[0001] This technology relates to an information processing apparatus that performs control for displaying a virtual object in association with a real object recognized by object recognition processing, and a method thereof. the one It relates to the technical field of this method.
Background Art
[0002] VR (Virtual Reality) technology that enables a user to perceive an artificially constructed virtual space has been put into practical use. In recent years, the spread of AR (Augmented Reality) technology, which is an advancement of VR technology, has been progressing. AR technology presents an augmented reality space (AR space) constructed by partially modifying the real space to the user. For example, in AR technology, a virtual object (virtual object) is superimposed on an image from an imaging device directed at the real space, providing a user experience as if the virtual object exists in the real space shown in the image. Alternatively, as AR technology, there is also one that projects an image of a virtual object onto the real space by a projector device to provide a user experience as if the virtual object exists in the real space.
[0003] When displaying a 3D object as a virtual object in association with a real object, such as superimposing it on the real object or displaying it in a predetermined positional relationship with the real object, the position and orientation of the real object are recognized, and the 3D object is drawn onto a two-dimensional image so that the 3D object is displayed in a position and orientation corresponding to the recognized position and orientation.
[0004] However, it may take a long time to draw a 3D object. If the user's head moves, for example, a change occurs in the position of the viewpoint before the drawn object is displayed to the user, a relative shift occurs between the position of the viewpoint and the position where the drawn object is displayed. Such a shift is recognized by the user as a delay in the object's follow-up with respect to the displacement of the real object. That is, it is recognized as a display delay of the object.
[0005] Regarding the display delay caused by such head movement, it is effective to correct the image of the drawn object based on the detection information of the position and orientation of the head (detection information of the viewpoint position and the line-of-sight direction) (see, for example, Patent Document 1 below). Specifically, from the information on the position and orientation of the head repeatedly detected at a predetermined cycle, for example, the amount of change in the position and orientation of the head from the start point to the end point of drawing is obtained, and based on this amount of change, image correction is performed to change the position and orientation of the object after drawing.
[0006] Thereby, it is possible to prevent the display delay of the object from occurring due to the time required for drawing.
Prior Art Documents
Patent Documents
[0007]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0008] However, in AR technology, the display delay of 3D objects does not occur only due to changes in the position and orientation of the display device. For example, when the target real object is a moving object, it can also occur due to the movement of the real object.
[0009] This technology has been made in view of the above circumstances, and aims to reduce the user's sense of discomfort and enhance the sense of immersion in the AR space by suppressing the display delay of virtual objects.
Means for Solving the Problems
[0010] The information processing apparatus according to the present technology performs, based on a captured image including a real object, first recognition processing regarding the position and orientation of the real object at a first time point, and second recognition processing regarding the position and orientation of the real object at a second time point after the first time point. An image recognition processing unit, a rendering control unit that controls a rendering processing unit to perform first rendering processing on an associated virtual object associated with the real object based on the first recognition processing and second rendering processing on an associated virtual object associated with the real object based on the second recognition processing, and a correction control unit that corrects a virtual object image, which is an image of the associated virtual object obtained by completion of the first rendering processing, based on the result of the second recognition processing before the second rendering processing is completed.
[0011] By performing image correction of the associated virtual object based on the recognition results of the position and orientation of the real object as described above, when the position and orientation of the real object change, it is possible to change the position and orientation of the associated virtual object in accordance with the change. And according to the above configuration, as long as the latest recognition result (the recognition result of the second recognition processing) of the associated virtual object image is obtained, it is possible to immediately output the image as an image corrected from the image obtained by the rendering processing (first rendering processing) based on the past recognition result without waiting for the completion of the rendering processing (second rendering processing) based on the latest recognition result.
[0012] In the information processing apparatus according to the present technology described above, it is conceivable that the correction control unit performs the correction of changing the position of the associated virtual object in the vertical and horizontal directions of the virtual object image based on the information on the position of the real object recognized by the image recognition processing unit in the vertical and horizontal directions of the real object.
[0013] Thereby, when the real object moves in the vertical and horizontal directions, it is possible to perform image correction that changes the position of the associated virtual object in the vertical and horizontal directions according to the movement.
[0014] In the information processing apparatus according to the present technology described above, it is conceivable that the correction control unit performs the correction of changing the size of the related virtual object for the virtual object image based on the information on the position of the real object recognized by the image recognition processing unit in the depth direction.
[0015] Thereby, when the real object moves in the depth direction, for example, when the real object approaches the user's viewpoint, the image of the related virtual object can be greatly changed, or conversely, when the real object moves away from the viewpoint, the image of the related virtual object can be made smaller. Thus, it is possible to change the size of the related virtual object according to the position of the real object in the depth direction.
[0016] In the information processing apparatus according to the present technology described above, it is conceivable that the correction control unit performs the correction of changing the position or orientation of the related virtual object according to a change in the viewpoint position or the line-of-sight direction of the user.
[0017] Thereby, it is possible to suppress the display delay caused by a change in the viewpoint position or the line-of-sight direction when the user moves the head or the like.
[0018] In the information processing apparatus according to the present technology described above, when the correction control unit selects one or more related virtual objects to be the target of the correction from among a plurality of related virtual objects each associated with a different real object, it is conceivable that the related virtual object of the real object with a large movement is preferentially selected.
[0019] Thereby, it is possible to prevent the image correction from being performed unnecessarily on the virtual object associated with the real object with a small movement or no movement.
[0020] In the information processing apparatus according to the present technology described above, it is conceivable that the processing cycle of the correction is set to be shorter than the processing cycle of the image recognition processing unit.
[0021] This makes it possible to shorten the delay time from when the recognition result of the real object is obtained until the image correction of the virtual object starts.
[0022] In the information processing apparatus according to the present technology described above, it is conceivable that the drawing control unit controls the drawing processing unit to draw the related virtual object and an unrelated virtual object, which is a virtual object independent of the image recognition processing of the real object, on different drawing planes in a plurality of drawing planes.
[0023] This makes it possible to perform appropriate image correction according to whether the virtual object is a related virtual object or not, such as performing image correction according to the user's viewpoint position and line-of-sight direction for the unrelated virtual object, and performing image correction according to the position and posture of the associated real object and the viewpoint position and line-of-sight direction for the related virtual object.
[0024] In the information processing apparatus according to the present technology described above, when the number of the related virtual objects is equal to or more than the number of the plurality of drawing planes, assuming that the number of the plurality of drawing planes is n (n is a natural number), the drawing control unit selects n - 1 of the related virtual objects, draws the selected related virtual objects exclusively on at least one drawing plane, and controls the drawing processing unit to draw the unselected related virtual objects and the unrelated virtual objects on the remaining one drawing plane.
[0025] As a result, when the virtual object includes an irrelevant virtual object for which image correction based on the object recognition result is unnecessary as a virtual object, and the number of relevant virtual objects is n or more with respect to the number n of drawing planes, image correction based on the recognition result of the relevant real object is performed for n - 1 relevant virtual objects, and for the remaining relevant virtual objects, image correction based on the viewpoint position and line-of-sight direction of the user is performed together with the irrelevant virtual objects. That is, when it is impossible to perform image correction based on the recognition result of the real object for all relevant virtual objects due to the relationship between the number of drawing planes and the number of relevant virtual objects, image correction based on the recognition result of the real object is preferentially performed for n - 1 relevant virtual objects.
[0026] In the information processing apparatus according to the present technology described above, it is conceivable that the drawing control unit performs the selection using a selection criterion in which the higher the movement amount of the real object, the higher the possibility of selection.
[0027] As a result, it is possible to preferentially select a relevant virtual object with a large movement amount and a high possibility of being perceived as a display delay as an object of image correction based on the recognition result of the real object.
[0028] In the information processing apparatus according to the present technology described above, it is conceivable that the drawing control unit performs the selection using a selection criterion in which the smaller the area of the real object, the higher the possibility of selection.
[0029] When a relevant virtual object is superimposed and displayed on a real object, even if the movement amount of the real object is large but the area of the real object is large, the ratio of the area of the position error generation part of the relevant virtual object to the area of the real object may be small, and in such a case, it is difficult to perceive a display delay. On the other hand, even if the movement amount of the real object is small but the area of the real object is small, the ratio may be large, and in such a case, it is easy to perceive a display delay.
[0030] In the information processing apparatus according to the present technology described above, it is conceivable that the drawing control unit makes the selection using a selection criterion in which the possibility of selection increases as the distance between the user's gaze point and the real object decreases.
[0031] Thereby, it becomes possible to select, as an object of image correction based on the recognition result of the real object, a related virtual object that is displayed near the user's gaze point and is highly likely to be perceived as a display delay.
[0032] In the information processing apparatus according to the present technology described above, it is conceivable that the drawing control unit controls the drawing processing unit so as to lower the update frequency of a drawing plane that draws an unrelated virtual object independent of the image recognition processing of the real object among a plurality of drawing planes, compared to the update frequency of a drawing plane that draws the related virtual object.
[0033] Thereby, it is possible to prevent drawing for all drawing planes from being performed at a high update frequency.
[0034] In the information processing apparatus according to the present technology described above, it is conceivable that the drawing control unit controls the drawing processing unit so as to lower the drawing update frequency of the related virtual object when the related virtual object is a related virtual object that performs an animation, compared to when the related virtual object is a related virtual object that does not perform an animation.
[0035] Thereby, in a case where a plurality of drawing planes have to be used, if the related virtual object to be drawn does not perform an animation, the drawing of the related virtual object is performed at a low update frequency, and if it performs an animation, the drawing of the related virtual object is performed at a high update frequency.
[0036] In the information processing apparatus according to the present technology described above, it is conceivable that the drawing control unit controls the drawing processing unit to use at least one drawing plane that is smaller in size than other drawing planes when drawing processing is performed for a plurality of drawing planes.
[0037] Accordingly, when it is necessary to use a plurality of drawing planes, it is possible to reduce the processing load of the drawing process.
[0038] In the information processing apparatus according to the present technology described above, when a part of the user's body overlaps a virtual object as viewed from the user's viewpoint position, the correction control unit performs the correction on a shielding virtual object that is a virtual object that shields the overlapping part of the virtual object. It is conceivable to adopt a configuration.
[0039] Accordingly, it is possible to suppress the display delay of the shielding virtual object.
[0040] In the information processing apparatus according to the present technology described above, it is conceivable that the shielding virtual object is a virtual object imitating the user's hand.
[0041] Accordingly, it is possible to suppress the display delay of the shielding virtual object imitating the user's hand.
[0042] In the information processing apparatus according to the present technology described above, the drawing control unit controls the drawing processing unit so as to exclusively use at least one drawing plane among the plurality of drawing planes that can be used by the drawing processing unit for the shielding virtual object. It is conceivable to adopt a configuration.
[0043] Accordingly, it is possible to preferentially perform image correction based on the object recognition result of the shielding virtual object.
[0044] In the information processing apparatus according to the present technology described above, before the completion of the first drawing process, based on the result of the first recognition process, a light source viewpoint image, which is an image of the related virtual object viewed from the position of the virtual light source that illuminates the related virtual object, is generated, and control is performed so that the generated light source viewpoint image is corrected based on the result of the second recognition process before the completion of the second drawing process. It is conceivable to further provide a virtual shadow image generation unit that generates a virtual shadow image, which is an image of a virtual shadow for the virtual related object, based on the corrected light source viewpoint image.
[0045] As a result, regarding the light source viewpoint image used for generating the virtual shadow image, even when the target real object moves, it becomes possible to immediately correct and use the image generated based on the past recognition result (the result of the first recognition process) based on the latest recognition result (the result of the second recognition process). When improving the realism by displaying the shadow (virtual shadow) of the related virtual object, it becomes possible to suppress the display delay of the shadow.
[0046] In the information processing apparatus according to the present technology described above, the virtual shadow image generation unit, before the completion of the first drawing process, based on the result of the first recognition process, calculates the distance from each point in the three-dimensional space projected onto each pixel of the drawing image by the drawing process unit to the virtual light source as the drawing-side light source distance for each, and as the light source viewpoint image, generates a depth image as a shadow map by the shadow map method. As the correction of the light source viewpoint image, a process of changing the position or size of the image area of the related virtual object in the shadow map based on the result of the second recognition process is performed, and the virtual shadow image is generated based on the corrected shadow map and the drawing-side light source distance. It is conceivable to adopt such a configuration.
[0047] That is, in the generation of a virtual shadow image by the shadow map method, correction is performed to change the position or size of the image area of the real object in the shadow map generated based on the result of the first recognition process based on the latest object recognition result (the result of the second recognition process).
[0048] Further, the control method according to the present technology performs first recognition processing on the position and orientation of the real object at a first point in time based on a captured image including the real object, and controls a rendering processing unit to perform first rendering processing on a related virtual object associated with the real object based on the first recognition processing. At a second point in time after the first point in time, second recognition processing on the position and orientation of the real object is performed based on a captured image including the real object, and the rendering processing unit is controlled to perform second rendering processing on a related virtual object associated with the real object based on the second recognition processing. Before the second rendering processing is completed, correction of a first image of the related virtual object obtained upon completion of the first rendering processing is performed based on the result of the second recognition processing.
[0049] Even with such a control method, the same operation as that of the information processing apparatus according to the present technology described above can be obtained.
Brief Description of the Drawings
[0050]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Embodiments for Carrying Out the Invention
[0051] Hereinafter, with reference to the accompanying drawings, embodiments of the present technology will be described in the following order. <1. Configuration of the AR System as an Embodiment> (1-1. System Configuration) (1-2. Example of the Internal Configuration of the Information Processing Device) <2. Delay Associated with the Rendering of Virtual Objects> <3. Delay Suppression Method as an Embodiment> <4. Processing Procedure> <5. Another Example of Reducing the Rendering Processing Load> <6. Regarding the Virtual Object for Occlusion> <7. Regarding the Shadow Display> <8. Modification Example> <9. Program and Storage Medium> <10. Summary of the Embodiment> <11. The Present Technology>
[0052] <1. Configuration of the AR System as an Embodiment> (1-1. System Configuration) FIG. 1 is a diagram showing a configuration example of an AR (Augmented Reality) system 50 configured to include an information processing device 1 as an embodiment. As shown in the figure, the AR system 50 as an embodiment is configured to include at least the information processing device 1.
[0053] Here, in FIG. 1, as examples of real objects Ro arranged in the real space, real objects Ro1, Ro2, and Ro3 are shown. In the AR system 50 of this example, for a predetermined real object Ro among these real objects Ro, a virtual object Vo is arranged to be superimposed by AR technology and displayed to the user. In the present disclosure, a virtual object Vo superimposed on the real space based on the image recognition result of the real object Ro may be referred to as a "related virtual object". In the illustrated example, the virtual object Vo2 is superimposed on the real object Ro2, and the virtual object Vo3 is superimposed on the real object Ro3, respectively. Specifically, the virtual object Vo is displayed so that the position of the virtual object Vo substantially coincides with the position of the real object Ro when viewed from the user's perspective. Note that the present technology is not limited to the technology of superimposing the virtual object Vo on the real object Ro. The virtual object Vo may be superimposed at a position associated with the position of the real object Vo. For example, it may be superimposed on the real space so as to fix the relative distance in a state separated from the real object Ro.
[0054] The position in the real space is defined by the values of three axes: the x-axis corresponding to the left-right direction axis, the y-axis corresponding to the up-down direction axis, and the z-axis corresponding to the depth direction axis.
[0055] The information processing device 1 acquires information for recognizing the real object Ro, recognizes the position of the real object Ro in the real space, and based on the recognition result, superimposes and displays the virtual object Vo on the real object Ro to the user.
[0056] FIG. 2 is a diagram showing an example of the external configuration of the information processing apparatus 1. The information processing apparatus 1 in this example is configured as a so-called head-mounted device that a user wears on at least a part of the head and uses. For example, in the example shown in FIG. 2, the information processing apparatus 1 is configured as a so-called eyewear type (glasses type) device, and at least one of the lenses 100a and 100b is configured as a transmissive display 10. Further, the information processing apparatus 1 includes a first imaging unit 11a and a second imaging unit 11b as an imaging unit 11, and also includes an operation unit 12 and a holding unit 101 corresponding to the frame of the eyewear. The holding unit 101 holds the display 10, the first imaging unit 11a and the second imaging unit 11b, and the operation unit 12 in a predetermined positional relationship with respect to the user's head when the information processing apparatus 1 is worn on the user's head. Although not shown in FIG. 2, the information processing apparatus 1 may include a sound collecting unit for collecting the user's voice or the like.
[0057] In the example shown in FIG. 2, the lens 100a corresponds to the lens on the right-eye side, and the lens 100b corresponds to the lens on the left-eye side. The holding unit 101 holds the display 10 so that the display 10 is positioned in front of the user's eyes when the information processing apparatus 1 is worn on the user.
[0058] The first imaging unit 11a and the second imaging unit 11b are configured as a so-called stereo camera, and are respectively held by the holding unit 101 so as to face substantially the same direction as the user's line of sight direction when the information processing apparatus 1 is worn on the user's head. At this time, the first imaging unit 11a is held near the user's right eye, and the second imaging unit 11b is held near the user's left eye. Based on such a configuration, the first imaging unit 11a and the second imaging unit 11b image a subject located on the front side (the user's line of sight direction side) of the information processing apparatus 1, particularly, a real object Ro located in the real space, from different positions. Thereby, the information processing apparatus 1 can acquire an image of a subject located on the front side of the user, and can calculate the distance to the subject based on the parallax between the images captured by the first imaging unit 11a and the second imaging unit 11b respectively. Note that the starting point when measuring the distance to the subject may be set as a position near the user's viewpoint position, such as a position where the user's viewpoint position can be regarded, for example, the positions of the first imaging unit 11a and the second imaging unit 11b, etc.
[0059] Note that the method for measuring the distance to the subject is not limited to the stereo method using the above-described first imaging unit 11a and second imaging unit 11b. As a specific example, distance measurement can also be performed based on methods such as motion parallax, ToF (Time Of Flight), and Structured Light. Here, ToF is a method of projecting light such as infrared rays onto the subject and measuring the time until the projected light is reflected back by the subject for each pixel, and obtaining an image (so-called distance image) including the distance (depth) to the subject based on the measurement result. Structured Light is a method of irradiating a pattern with light such as infrared rays onto the subject and imaging it, and obtaining a distance image including the distance (depth) to the subject based on the change in the pattern obtained from the imaging result. Motion parallax is a method of measuring the distance to the subject based on parallax even in a monocular camera. Specifically, by moving the camera, the subject is imaged from different viewpoints, and the distance to the subject is measured based on the parallax between the captured images. At this time, by recognizing the moving distance and moving direction of the camera by various sensors, it becomes possible to measure the distance to the subject with higher accuracy. Note that the configuration of the imaging unit (for example, a monocular camera, a stereo camera, etc.) may be changed according to the distance measurement method.
[0060] The operation unit 12 is a configuration for receiving operations from the user to the information processing apparatus 1. The operation unit 12 may be configured by an input device such as a touch panel or buttons, for example. The operation unit 12 is held at a predetermined position of the information processing apparatus 1 by the holding unit 101. For example, in the example shown in FIG. 2, the operation unit 12 is held at a position corresponding to the temple of the glasses.
[0061] The information processing apparatus 1 illustrated in FIG. 2 corresponds to an example of a see-through type HMD (Head Mounted Display). The see-through type HMD holds, in front of the user's eyes, a virtual image optical system including a transparent light guide portion or the like using, for example, a half mirror or a transparent light guide plate, and displays an image inside the virtual image optical system. Therefore, a user wearing the see-through type HMD can see the external scenery while viewing the image displayed inside the virtual image optical system. With such a configuration, the see-through type HMD can superimpose an image of a virtual object on an optical image of a real object located in the real space, for example, based on AR technology.
[0062] (1-2. Example of internal configuration of information processing apparatus) FIG. 3 is a block diagram showing an example of the internal configuration of the information processing apparatus 1. As shown in the figure, the information processing apparatus 1 includes the above-described display 10, imaging unit 11, and operation unit 12, and further includes a sensor unit 13, a CPU (Central Processing Unit) 14, a ROM (Read Only Memory) 15, a RAM (Random Access Memory) 16, a GPU (Graphics Processing Unit) 17, an image memory 18, a display controller 19, a recording / reproducing control unit 20, a communication unit 21, and a bus 22. As shown in the figure, each of the imaging unit 11, operation unit 12, sensor unit 13, CPU 14, ROM 15, RAM 16, GPU 17, image memory 18, display controller 19, recording / reproducing control unit 20, and communication unit 21 is connected via the bus 22 and can perform data communication with each other via the bus 22.
[0063] The sensor unit 13 comprehensively shows sensors for detecting the position (position in the real space) and movement of the information processing apparatus 1 itself according to the movement of the head of the user wearing the information processing apparatus 1. Specifically, the sensor unit 13 in this example has an acceleration sensor and an angular velocity sensor (gyro sensor). As the acceleration sensor, a three-axis acceleration sensor is used, and as the angular velocity sensor, it is configured to be able to detect components in the yaw direction, pitch direction, and roll direction respectively. Thereby, changes in the position and orientation of the information processing apparatus 1 itself can be detected.
[0064] Here, the position of the information processing apparatus 1 itself detected based on the detection signal of the sensor unit 13 (hereinafter sometimes referred to as the "sensor signal") can be regarded as the viewpoint position of the user. Also, the orientation (direction) of the information processing apparatus 1 itself detected based on the detection signal of the sensor unit 13 can be regarded as the line-of-sight direction of the user. In this sense, in the following description, the detection of the position of the information processing apparatus 1 based on the sensor signal is referred to as "detection of the viewpoint position", and the detection of the orientation of the information processing apparatus 1 based on the sensor signal is referred to as "detection of the line-of-sight direction".
[0065] The CPU 14 executes various processes according to the programs stored in the ROM 15 or the programs loaded into the RAM 16. The RAM 16 also appropriately stores data and the like necessary for the CPU 14 to execute various processes.
[0066] The GPU 17 performs rendering processing of the virtual object Vo as a 3D (three-dimensional) object. During this rendering, the image memory 18 is used. Specifically, in the image memory 18, a plurality of buffers (buffer areas) 18a used as the frame buffer of the image can be set, and the GPU 17 appropriately uses any one of these buffers 18a as the frame buffer during 3D object rendering. In the present disclosure, the GPU 17 may be regarded as corresponding to the rendering processing unit. Here, an example in which the GPU 17 is configured as a separate processor from the CPU 14 is shown, but there may be a case where the GPU 17 is configured as an integrated processor with the CPU 14.
[0067] The display controller 19 performs output processing of the image (two-dimensional image) obtained by the rendering processing of the GPU 17 to the display 10. The display controller 19 in this example has a function as an image correction processing unit 19a. The function as this image correction processing unit 19a is a function of correcting the image of the virtual object Vo drawn in two dimensions (for example, correction of position and deformation), and the details thereof will be described later again.
[0068] Here, in this example, the processing cycle of the image output by the display controller 19 (the processing cycle of the image correction by the image correction processing unit 19a) is set to be shorter than the frame cycle of the imaging unit 11. For example, while the frame cycle of the imaging unit 11 is 60 Hz, the processing cycle of the display controller 19 is set to 120 Hz. The processing cycle of the object recognition processing by the image recognition processing unit F1 described later coincides with the frame cycle of the imaging unit 11. Therefore, the processing cycle of the display controller 19 is set to be shorter than the processing cycle of the object recognition processing.
[0069] The recording and playback control unit 20 performs recording and playback on a recording medium such as a non-volatile memory. The actual form of the recording and playback control unit 20 can be considered in various ways. For example, the recording and playback control unit 20 may be configured as a flash memory built into the information processing device 1 and its writing / reading circuit, or may be in the form of a card recording and playback unit that performs recording and playback access on a recording medium detachable from the information processing device 1, such as a memory card (portable flash memory, etc.). It can also be realized as an SSD (Solid State Drive), HDD (Hard Disk Drive), etc. in the form built into the information processing device 1.
[0070] The communication unit 21 performs communication processing via a network and inter-device communication. The CPU 14 is enabled to perform data communication with an external device via this communication unit 21.
[0071] FIG. 4 is an explanatory diagram of the functions of the CPU 14 of the information processing device 1. As shown in the figure, the CPU 14 has functions as an image recognition processing unit F1, a drawing control processing unit F2, and an image correction control unit F3.
[0072] The image recognition processing unit F1 performs recognition processing (object recognition processing) of the real object Ro located in the real space based on the captured image obtained by the imaging unit 11. Specifically, in the recognition processing of this example, recognition of the type of the real object Ro and recognition of the position and orientation in the real space are performed. As described above, in this example, the distance to the real object Ro can be calculated based on the parallax information between the stereoscopically captured images. The image recognition processing unit F1 recognizes the position of the real object Ro based on this distance information.
[0073] The rendering control unit F2 performs rendering control of the virtual object Vo. Specifically, it controls the GPU17 so that the virtual object Vo is rendered at the required position and orientation. In this example, the virtual object Vo is to be displayed superimposed on the corresponding real object Ro. Therefore, based on the information on the position and orientation of the real object Ro obtained by the image recognition processing unit F1 in the recognition processing of the real object Ro, the rendering control unit F2 controls the rendering process of the virtual object Vo by the GPU17 so that a display image of the virtual object Vo can be obtained at the position and orientation superimposed on the real object Ro.
[0074] Note that in this example, when rendering the virtual object Vo using the GPU17, a plurality of rendering planes can be used. The rendering control unit F2 in this example performs processing to switch the number of rendering planes to be used and the usage mode of the rendering planes according to the number and type of the virtual object Vo (the virtual object Vo to be displayed to the user) to be rendered. This will be described in detail later.
[0075] The image correction control unit F3 controls the image correction process by the image correction processing unit 19a. By controlling the image correction process by such an image correction processing unit 19a, it is possible to adjust the position and orientation of the image of the rendered virtual object Vo on the display 10. The details of the process executed by the CPU14 as the image correction control unit F3 will be described in detail later.
[0076] <2. Delay Associated with Rendering of Virtual Object> FIG. 5 is an explanatory diagram of the delay associated with the drawing of the virtual object Vo. "Input" in the figure means the input of information necessary for obtaining the object recognition result of the real object Ro. In this example, the image capture by the imaging unit 11 corresponds to this. Therefore, the period of "Input" shown in the figure corresponds to the frame period in the imaging unit 11. Also, "Recognition" in the figure means the object recognition of the real object Ro based on "Input" (in this example, particularly the recognition of the position of the real object Ro). "Drawing" means the drawing of the virtual object Vo superimposed on the recognized real object Ro, and "Output" means the output of the image of the drawn virtual object Vo (output to the display 10). Note that, as understood from the explanations so far, "Drawing" should be performed based on the recognition result of the position and orientation of the real object Ro that is the superimposition target of the virtual object Vo, and it should start after the completion of "Recognition".
[0077] As shown in the figure, under the condition that "Input" is repeated at a predetermined period, for each "Input", "Recognition", "Drawing", and "Output" are performed in order. At this time, the time required from "Input" to "Output" becomes the display delay amount of the virtual object Vo with respect to the real object Ro. Conventionally, since the image of the virtual object Vo obtained by "Drawing" was output as it was, the time required for "Drawing" was directly reflected in the display delay amount of the virtual object Vo. Since it takes a relatively long time to draw the virtual object Vo as a 3D object, it has been difficult to suppress the display delay conventionally.
[0078] <3. Delay suppression method as an embodiment> FIG. 6 is an explanatory diagram of a delay suppression method as an embodiment. In order to suppress the display delay of the virtual object Vo, in this embodiment, instead of outputting an image of the virtual object Vo drawn based on the latest object recognition result as in the prior art, in response to obtaining the latest object recognition result, a method is adopted in which an image of the virtual object Vo that has been drawn based on a past object recognition result is corrected based on the latest object recognition result. In other words, the image obtained in the drawing process of the virtual object Vo based on the recognition result of the object recognition process executed at the first time point is corrected based on the recognition result of the object recognition process executed at the second time point after the first time point.
[0079] The image correction here is performed by the above-described image correction processing unit 19a. Specifically, as the image correction in this case, for example, when the real object Ro to be superimposed moves leftward from the first time point to the second time point, the virtual object Vo drawn based on the object recognition result at the first time point is moved leftward within the drawing frame for image correction. FIG. 7A shows the image thereof. Alternatively, when the real object Ro to be superimposed moves upward from the first time point to the second time point, the virtual object Vo drawn based on the object recognition result at the first time point is moved upward within the drawing frame for image correction (see FIG. 7B). Thus, in this example, correction is performed to change the position of the virtual object Vo within the frame in the vertical and horizontal directions in response to the change in the position of the real object Ro in the vertical and horizontal directions within the plane. In other words, correction is performed to change the position of the virtual object Vo in the vertical and horizontal directions within the plane of the image drawn in the drawing process based on the information on the position of the real object Ro recognized in the object recognition process in the vertical and horizontal directions within the plane.
[0080] Further, when the posture of the real object Ro changes, such as when the real object Ro to be superimposed rotates rightward from the first time point to the second time point, as the image correction of the virtual object Vo drawn based on the object recognition result at the first time point, correction is performed to change the posture of the virtual object Vo within the drawing frame so as to follow the recognized change in the posture of the real object Ro, such as rotating the virtual object Vo rightward within the drawing frame.
[0081] Furthermore, in the image correction in this case, when the real object Ro to be superimposed moves in the forward direction (the direction approaching the user) or the back direction from the first time point to the second time point, for the virtual object Vo drawn based on the object recognition result at the first time point, image correction is performed to increase or decrease its size. In other words, based on the information on the position of the real object Ro recognized in the object recognition process in the depth direction, correction is performed to change the size of the virtual object Vo in the image drawn in the drawing process.
[0082] By performing image correction based on the object recognition result as described above, when the real object Ro to be superimposed on the virtual object Vo is a moving object, it is possible to appropriately follow the changes in the position and posture of the virtual object Vo with respect to the changes in the position and posture of the real object Ro. In addition, when performing such image correction, as shown in FIG. 6, after the latest object recognition result is obtained, it is possible to output the image of the virtual object Vo without going through the drawing process based on the latest object recognition result. Therefore, the display delay amount can be significantly suppressed compared to the case of FIG. 5. At this time, in order to suppress the display delay amount, the image correction based on the latest object recognition result (the recognition result by the second recognition process) is executed before the drawing process (the second drawing process) based on the latest object recognition result is completed.
[0083] For confirmation, it should be noted that since the image correction process by the image correction processing unit 19a is a process for a two-dimensional image, the processing time is significantly shorter compared to the drawing process by the GPU 17. Here, a configuration is exemplified in which the image correction process based on the object recognition result is performed by the display controller 19 provided outside the CPU 14, but the CPU 14 can also execute the image correction process. Alternatively, for at least some functions, the image correction process can adopt a configuration in which the display controller 19 cooperates with the CPU 14 to perform the process.
[0084] Here, between the first point in time and the second point in time described above, the viewpoint position and the line-of-sight direction may change due to the movement of the user's head or the like. Regarding the relative deviation between the real object Ro and the virtual object Vo caused by such changes in the viewpoint position and the line-of-sight direction, it cannot be suppressed only by the image correction based on the object recognition result described above.
[0085] Therefore, in this example, as the image correction using the image correction processing unit 19a, image correction based on the detection signal of the sensor unit 13 is also performed together.
[0086] FIG. 8 is an explanatory diagram of image correction based on the detection signal of the sensor unit 13. In this FIG. 8, the processing timings along the respective time series of the drawing of the virtual object, the sensor input (input of the detection signal of the sensor unit 13) for detecting the viewpoint position and the line-of-sight direction of the user, and the output of the drawn virtual object Vo are schematically shown.
[0087] In FIG. 8, the time points from T1 to T3 indicate the start timing of the drawing of the virtual object Vo, and the time points from T1' to T3' indicate the end timing of the drawing of the virtual object Vo. Further, the frame images FT1 to FT3 show an example of the frame images drawn at the time points from T1 to T3, and schematically show the shape and position of the virtual object Vo in the image. Also, in FIG. 8, the time points from t1 to t4 indicate the output timing of the image of the virtual object Vo to the display 10, and the frame images Ft1 to Ft4 show an example of the frame images output at the time points from t1 to t4, and schematically show the shape and position of the virtual object Vo in the image. Here, as shown in FIG. 8, the sensor input is acquired at a cycle (higher frequency) faster than the cycle (frequency) at which the virtual object Vo is drawn.
[0088] First, at time point T1, the rendering of the virtual object Vo starts, and at time point T1', the rendering is completed, obtaining a frame image FT1. Then, when the output timing of the image, i.e., time point t1, arrives, based on the sensor input immediately before time point t1, the position of the virtual object Vo in the frame image FT1 is corrected, and the corrected image is output as the frame image Ft1. Next, when the output timing of the image, i.e., time point t2, arrives, but at this time, the rendering at time point T2 has not been executed yet. Therefore, based on the sensor input immediately before time point t2, the position of the virtual object Vo in the frame image FT1 is corrected, and the image obtained by this correction is output as the frame image Ft2.
[0089] Next, at time point T2, the rendering of the virtual object Vo starts, and the rendering ends at time point T2', obtaining a frame image FT2. That is, at time points t3 and t4, which are the output timings that arrive after time point T2', frame images Ft3 and Ft4 are output, in which the position of the virtual object Vo in the frame image FT2 is corrected based on the sensor input immediately before each time point. In the illustrated example, the rendering at time point T3 starts after the output timing of time point t4. However, at the output timings after time point T3, unless new rendering is performed after time point T3, a frame image obtained by correcting the position of the virtual object Vo in the frame image FT3 based on the sensor input immediately before is output.
[0090] By performing image correction based on the sensor signal as described above, even if the user's viewpoint position or line-of-sight direction changes between the first time point and the second time point, the position of the virtual object Vo in the post-rendering image can be corrected so as to follow the change. That is, it is possible to suppress the display delay of the virtual object Vo caused by the change in the user's line-of-sight position or line-of-sight direction.
[0091] Here, a pattern of image correction based on the sensor signal as described above will be described with reference to FIG. 9. As shown in FIG. 9, when the movement of the user's head (movement of the viewpoint position or the line-of-sight direction) is in the left or right direction, corrections are made to change the position of the virtual object Vo in the drawing frame to the right or left direction, respectively. When the movement of the user's head is in the downward or upward direction, corrections are made to change the position of the virtual object Vo in the drawing frame to the upward or downward direction, respectively. Further, when the user's head moves forward (i.e., approaches the real object Ro to be superimposed) or backward, corrections are made to increase or decrease the size of the virtual object Vo, respectively. For rotation, image correction is performed to rotate the virtual object Vo in the direction opposite to the movement of the head.
[0092] Furthermore, in this example, it is also possible to perform trapezoidal correction as the image correction in the image correction processing unit 19a. This trapezoidal correction is also performed in a manner corresponding to the movement of the head detected from the sensor input.
[0093] Here, when different virtual objects Vo are superimposed on different real objects Ro, it is conceivable to perform image correction for each virtual object Vo individually based on the above-described object recognition result and sensor signal. When individual image correction should be performed for each virtual object Vo in this way, it is ideal to draw each virtual object Vo using an individual drawing plane and perform image correction for each virtual object Vo on the frame image obtained in each drawing.
[0094] FIG. 10 is an explanatory diagram for image correction after rendering a plurality of virtual objects Vo using individual rendering planes. First, as a premise, in the information processing apparatus 1 of this example, two rendering planes, a first plane and a second plane, can be used as rendering planes. For confirmation, the rendering plane corresponds to the display surface of the display 10 and means a frame in which a 3D object as a virtual object Vo is rendered as two-dimensional image information. One rendering plane corresponds to one buffer 18a in the image memory 18. When there are a plurality of rendering planes, different virtual objects Vo can be rendered on each of them and combined to represent the image information of each virtual object Vo on the display surface. And the display controller 19 of this example can perform image correction processing individually for the first plane and the second plane. In other words, different image correction processing can be performed for each rendering plane.
[0095] Note that FIG. 10 shows an example in which the positions of two virtual objects Vo overlap within the combined frame. In this case, which virtual object Vo is brought to the front is determined based on the position (distance) in the depth direction of the target real object Ro.
[0096] However, rendering a plurality of virtual objects Vo simultaneously leads to an increase in the processing load and is not desirable. For this reason, in this example, control is performed to switch the usage mode of the rendering plane, the update cycle of rendering, etc. according to the number and type of virtual objects Vo to be displayed. Note that this point will be described in detail below.
[0097] <4. Processing Procedure> The flowcharts in FIGS. 11 to 13 show examples of specific processing procedures that the CPU 14 should execute as the above-described rendering control unit F2 and image correction control unit F3. Note that the processing shown in FIGS. 11 to 13 is executed based on a program stored in the ROM 15 by the CPU 14 or a program stored in a storage device that can be read by the recording and playback control unit 20.
[0098] Figure 11 shows the process corresponding to the drawing control unit F2. First, in step S101, the CPU 14 determines whether there is a drawing of the virtual object Vo superimposed on the real object Ro. If there is no drawing of the virtual object Vo, the CPU 14 executes the drawing setting and image correction setting process in step S102 and ends the series of processes shown in Figure 11. In the drawing setting and image correction setting process in step S102, the CPU 14 controls the GPU 17 to use the first plane for the drawing of all virtual objects Vo, and as the control of the image correction process of the virtual object Vo drawn by the first plane, the first correction control is executed. Also, in the drawing setting and image correction setting process in step S102, the CPU 14 sets the second plane as not in use.
[0099] Here, the first correction control means controlling so that image correction based on the above-described sensor signal is performed. That is, according to the processes of steps S101→S102 described above, when the virtual object Vo to be drawn (that is, the display object) is only a virtual object Vo (non-related virtual object) that does not overlap the real object Ro, only image correction based on the sensor signal is performed as the image correction of all virtual objects Vo. Also, at this time, since it is not necessary to perform drawing by dividing the drawing plane for each virtual object Vo, the second plane is set as not in use.
[0100] Examples of the virtual object Vo that does not overlap the real object Ro include a virtual object Vo that should be fixedly arranged at a predetermined position in the AR space.
[0101] Figure 12 shows the process for realizing the first correction control. First, in step S201, the CPU 14 acquires the information on the position and orientation of the head. This is a process of acquiring the information on the position and orientation of the user's head (information on the line-of-sight position and line-of-sight direction) based on the detection signal of the sensor unit 13. Note that, as described above, the acquisition period of the sensor signal is set to be shorter than the drawing period of the virtual object Vo and the image output period to the display 10.
[0102] In step S202 following step S201, the CPU 14 calculates the amount of change in the position and orientation of the head. As can be understood from the previous FIG. 8, as this amount of change, the amount of change from the latest drawing start time to the sensor signal acquisition time immediately before output is calculated.
[0103] Next, in step S203, the CPU 14 issues an image correction instruction for the virtual object Vo according to the calculated amount of change, and finishes the first correction control process shown in FIG. 12. Here, as can be understood from the previous FIG. 9 and the like, in the image correction processing unit 19a, as image correction of the virtual object Vo, it is possible to perform various image corrections such as displacement in each of the up, down, left, and right directions, change in size, attitude change such as rotation, and trapezoidal correction. As the process of step S203, based on the amount of change in position and orientation calculated in step S202, each correction parameter for these displacements in each of the up, down, left, and right directions, change in size, attitude change such as rotation, trapezoidal correction, etc. is calculated, and the calculated each correction parameter is instructed to the image correction processing unit 19a (display controller 19) to execute the process.
[0104] Returning to the description with reference to FIG. 11. When it is determined in step S101 that there is a drawing of the virtual object Vo superimposed on the real object Ro, the CPU 14 proceeds to step S103 and determines whether the number of virtual objects Vo to be drawn is plural. When the number of virtual objects Vo to be drawn is not plural, that is, when the virtual object Vo to be drawn is only one virtual object Vo superimposed on the real object Ro, the CPU 14 executes the drawing setting and image correction setting process of step S104 and finishes the series of processes shown in FIG. 11. In this drawing setting and image correction setting process of step S104, the CPU 14 controls the GPU 17 to use the first plane for drawing the target virtual object Vo, and as control of the image correction process of the virtual object Vo drawn by the first plane, executes the second correction control. Also, in the drawing setting and image correction setting process of step S104, the CPU 14 sets the second plane as not used.
[0105] The second correction control means controlling so that image correction based on both the sensor signal and the object recognition result is performed. That is, according to the above-described processing of steps S103→S104, when there is only one virtual object Vo in which the virtual object Vo to be drawn overlaps the real object Ro, image correction based on both the sensor signal and the object recognition result is performed as the image correction of the virtual object Vo. Also in this case, since it is not necessary to perform drawing by dividing the drawing plane for each virtual object Vo, the second plane is not used.
[0106] FIG. 13 shows the processing for realizing the second correction control. First, in order to perform image correction based on the sensor signal, also in this case, the CPU 14 performs the processing of steps S201 and S202 to calculate the amount of change in the position and orientation of the head. Then, in response to the execution of the processing of step S202, the CPU 14 acquires the recognition result in step S210. That is, the information on the position and orientation of the real object Ro recognized in the recognition process for the corresponding real object Ro is acquired.
[0107] In step S211 following step S210, the CPU 14 issues an image correction instruction for the virtual object Vo according to the calculated amount of change and the recognition result, and ends the second correction control process shown in FIG. 13. As the processing of this step S211, the CPU 14 first obtains the amount of change in the real object Ro from the first time point to the second time point based on the recognition result obtained in step S210. Then, based on such an amount of change in the real object Ro and the amount of change calculated in step S202, correction parameters for each image correction such as displacement in each of the above-described up, down, left, and right directions, change in size, attitude change such as rotation, and trapezoid correction that can be executed by the image correction processing unit 19a are calculated, and the calculated correction parameters are instructed to the image correction processing unit 19a (display controller 19).
[0108] Returning to FIG. 11, when the CPU 14 determines in step S103 that there are a plurality of virtual objects Vo to be drawn, it proceeds to step S105 to determine whether there are a plurality of virtual objects Vo to be superimposed on the real object Ro. When it is determined that there are not a plurality of virtual objects Vo to be superimposed on the real object Ro, that is, when the virtual object Vo to be drawn is only one virtual object Vo superimposed on the real object Ro and one or more virtual objects Vo not superimposed on the real object Ro, the CPU 14 proceeds to step S106 to determine whether the virtual object Vo superimposed on the real object Ro has an animation. The animation mentioned here is assumed to be, for example, an animation that changes at least one of the color, pattern, and shape of the virtual object Vo in response to the occurrence of a predetermined event for the virtual object Vo, such as the user's hand touching (virtual contact) the virtual object Vo in the AR space. Note that the virtual object Vo "having an animation" can be paraphrased as the virtual object Vo "performing an animation".
[0109] In step S106, when it is determined that the virtual object Vo superimposed on the real object Ro does not have an animation, the CPU 14 executes the drawing setting and image correction setting process of step S107 and ends the series of processes shown in FIG. 11. As the drawing setting and image correction setting process of step S107, the CPU 14 controls the GPU 17 to use the first plane for drawing the virtual object Vo superimposed on the real object Ro and perform the drawing at a low update frequency, and as the control of the image correction process of the virtual object Vo drawn by the first plane, executes the second correction control. Also, in the drawing setting and image correction setting process of step S107, the CPU 14 controls the GPU 17 to use the second plane for drawing the other virtual objects Vo, and as the control of the image correction process of the virtual object Vo drawn by the second plane, executes the first correction control.
[0110] In the case of reaching step S107, there are virtual objects Vo that overlap with the real object Ro and virtual objects Vo that do not overlap with the real object Ro mixed together. However, if the second correction control is performed on the latter virtual objects Vo together with the former virtual objects Vo, there is a risk that the latter virtual objects Vo cannot be displayed at appropriate positions. For this reason, the drawing planes used for the virtual objects Vo that overlap with the real object Ro and those that do not overlap are separated, and efforts are made to ensure that each virtual object Vo is displayed at an appropriate position. And at this time, if the drawing for the two drawing planes is executed at their respective normal update frequencies, the processing load increases, which is not desirable. For this reason, the drawing of the virtual objects Vo that overlap with the real object Ro is performed at a lower update frequency than normal. Here, in this example, the normal update frequency of the drawing process is set to 60 Hz, and the low update frequency is set to a lower update frequency such as 30 Hz or the like.
[0111] Note that in step S107, the drawing update frequency of the second plane can also be set to a low update frequency. Regarding this point, the same applies to the drawing of the second plane in steps S108, S111, and S112, which will be described later.
[0112] On the other hand, in step S106, when it is determined that the virtual object Vo that overlaps with the real object Ro has an animation, the CPU 14 executes the drawing setting and image correction setting process of step S108 and finishes the series of processes shown in FIG. 11. As the drawing setting and image correction setting process of step S108, the CPU 14 controls the GPU 17 to use the first plane for drawing the virtual object Vo that overlaps with the real object Ro, and as the control of the image correction process of the virtual object Vo drawn by the first plane, the second correction control is executed. Also, in the drawing setting and image correction setting process of step S108, the CPU 14 controls the GPU 17 to use the second plane for drawing other virtual objects Vo, and as the control of the image correction process of the virtual object Vo drawn by the second plane, the first correction control is executed.
[0113] When the virtual object Vo superimposed on the real object Ro has an animation as described above, the drawing update frequency of the virtual object Vo should not be decreased. By doing so, it is possible to prevent the accuracy of the animation of the virtual object Vo from deteriorating.
[0114] Also, in step S105, when it is determined that there are a plurality of virtual objects Vo superimposed on the real object Ro, the CPU 14 proceeds to step S109 and executes a process of selecting one, that is, a process of selecting one from among the plurality of virtual objects Vo superimposed on the real object Ro.
[0115] Here, in the selection process of step S109, the virtual object Vo is selected based on the magnitude of the movement, the area, etc. of the real object Ro to be superimposed. As a basic concept, the virtual object Vo with a large projection error is selected. Specifically, for each virtual object Vo, the following index value S of the projection error is obtained, and the virtual object Vo with the maximum index value S is selected. In the following formula, the area a is the area of the real object Ro to be superimposed (the area of the surface visible from the user's perspective), and the movement amount m is the movement amount of the real object Ro to be superimposed. Index value S = (1 / area a) × movement amount m
[0116] Also, since a person cannot see the details except for the fixation point, taking the proximity to the fixation point (the reciprocal of the distance between the fixation point and the real object Ro) as α Index value S' = (1 / area a) × movement amount m × α it is also possible to calculate this and select the virtual object Vo with the maximum index value S'. Here, as the fixation point, a position such as the center point of the screen of the display 17, which is predetermined as the object that the user is looking at, may be set. Alternatively, in a configuration that performs user gaze detection, the position estimated from the gaze detection result can also be used.
[0117] In the selection process of step S109, since it is too costly to calculate the area a precisely, a simple model (such as a bounding box, etc.) can also be used as a substitute. Also, since it is not desirable in terms of the user experience for the selected virtual object Vo to switch frequently, it is effective to provide hysteresis. For example, once selected, the index value S (or index value S') is multiplied by a predetermined multiple such as 1.2 times to make it difficult to switch. Also, when prioritizing power consumption, when the index values S (S') of all virtual objects Vo are below a certain value, they may all be drawn on the same plane.
[0118] In step S110 following step S109, the CPU 14 determines whether the selected virtual object Vo has an animation. If the selected virtual object Vo does not have an animation, the CPU 14 executes the drawing setting and image correction setting process of step S111 and finishes the series of processes shown in FIG. 11. As the drawing setting and image correction setting process of step S111, the CPU 14 controls the GPU 17 to select the first plane for drawing the selected virtual object Vo and perform the drawing at a low update frequency, and as the control of the image correction process of the virtual object Vo drawn by the first plane, executes the second correction control. Also, in the drawing setting and image correction setting process of step S111, the CPU 14 controls the GPU 17 to use the second plane for drawing other virtual objects Vo, and as the control of the image correction process of the virtual object Vo drawn by the second plane, executes the first correction control.
[0119] On the other hand, when it is determined that the virtual object Vo selected in step S110 has an animation, the CPU 14 executes the drawing setting and image correction setting process in step S112 and ends the series of processes shown in FIG. 11. As the drawing setting and image correction setting process in step S112, the CPU 14 controls the GPU 17 to be used for drawing the virtual object Vo selected for the first plane, and as control of the image correction process of the virtual object Vo drawn by the first plane, executes second correction control. Also, in the drawing setting and image correction setting process in step S112, the CPU 14 controls the GPU 17 to be used for drawing other virtual objects Vo for the second plane, and as control of the image correction process of the virtual object Vo drawn by the second plane, executes first correction control.
[0120] As described above, in this example, in response to the case where there are only two drawing planes, when there are a plurality of virtual objects Vo superimposed on the real object Ro, one virtual object Vo is selected, and the drawing of the selected virtual object Vo is exclusively performed using a single drawing plane. The "exclusively" mentioned here means that only a single virtual object is drawn by a single drawing plane, and two or more virtual objects Vo are not drawn simultaneously by the single drawing plane.
[0121] By performing such selection of the virtual object Vo, when it becomes impossible to perform image correction based on the object recognition result for all the virtual objects Vo superimposed on the real object Ro due to the relationship between the number of drawing planes and the number of virtual objects Vo superimposed on the real object Ro, it is possible to preferentially perform image correction based on the object recognition result for one virtual object Vo.
[0122] Here, the case where only two drawable planes are available has been exemplified above. However, even when the number of available drawable planes is 3 or more, a virtual object Vo that exclusively uses a drawable plane can be selected in the same way of thinking. For example, assume that the number of drawable planes is 3 and the number of virtual objects Vo that overlap the real object Ro is 3 or more. In this case, the number of virtual objects Vo that can exclusively use a drawable plane can be 2. Therefore, as the selection of virtual objects Vo, two are selected from among 3 or more virtual objects Vo. Generally speaking, when the number of virtual objects Vo (related virtual objects) that overlap the real object Ro is equal to or more than the number of drawable planes, when the number of drawable planes is n (n is a natural number of 2 or more), n - 1 related virtual objects are selected. Then, the selected related virtual objects are drawn so as to exclusively use a drawable plane (that is, one virtual object Vo is drawn on only one drawable plane), and all virtual objects other than the selected related virtual objects among the virtual objects to be displayed are drawn using the remaining one drawable plane. In the present disclosure, the related virtual object may be regarded as a virtual object in which the relative positional relationship with respect to the absolute position or orientation of the real object Ro is fixed. The display position of the related virtual object may be corrected with reference not only to the image recognition result (object recognition result) of the real object Ro but also to the result of self-position estimation described later.
[0123] As a result, when the virtual object Vo to be displayed includes a virtual object Vo that does not overlap with the real object Ro (that is, an irrelevant virtual object for which image correction based on the object recognition result is unnecessary), and the number of relevant virtual objects is n or more with respect to the number n of drawing planes, image correction based on the recognition result of the relevant real object Ro is performed for n - 1 relevant virtual objects, and for the remaining relevant virtual objects, together with the irrelevant virtual objects, image correction according to the viewpoint position and line-of-sight direction of the user can be performed. That is, when it is impossible to perform image correction based on the recognition result of the real object Ro for all relevant virtual objects due to the relationship between the number of drawing planes and the number of relevant virtual objects, image correction according to the recognition result of the real object can be preferentially performed for n - 1 relevant virtual objects. In the present disclosure, the irrelevant virtual object may be regarded as a virtual object Vo whose position and orientation are controlled independently of the absolute position and orientation of a specific real object Ro. In other words, the position and orientation of the irrelevant virtual object are determined without depending on the image recognition result of a specific real object Ro. For example, the display position of the irrelevant virtual object is determined in the absolute coordinate system (three-dimensional coordinate system) of the real space based on the result of self-position estimation described later. Alternatively, the irrelevant virtual object may be a virtual object (for example, a GUI) displayed in a relative coordinate system with the position of the display device as the origin.
[0124] Here, in order to enhance the suppression effect of the display delay of the virtual object Vo, it is desirable to appropriately adjust the phase of the processing timing (phase of the operation clock) between the object recognition processing side and the image output processing side to the display 10.
[0125] FIG. 14 is a diagram for explaining the phase adjustment of the processing timing between the recognition processing side and the output processing side. FIG. 14A shows the processing cycle of the recognition processing and the execution period of the recognition processing within one cycle, and FIGS. 14B and 14C show the processing cycle of the output processing. As described above, the processing cycle of the output processing (that is, the processing cycle of the image correction processing unit 19a: for example, 120 Hz) is set to be shorter than the processing cycle of the recognition processing (for example, 60 Hz).
[0126] In the phase relationship shown as a comparison between FIGS. 14A and 14B, the error time (see the arrow in the figure) from the completion timing of the recognition process until the start of image output is relatively long, and this error time is reflected as the display delay time of the virtual object Vo. On the other hand, in the phase relationship shown as a comparison between FIGS. 14A and 14C, the completion timing of the recognition process and the start timing of image output substantially coincide, and the error time can be suppressed to approximately 0. That is, the display delay suppression effect of the virtual object Vo can be enhanced compared to the case of FIG. 14B.
[0127] <5. Another Example of Reducing the Rendering Processing Load> In the above, an example of reducing the rendering update frequency of at least one rendering plane was given for reducing the rendering processing load. However, as illustrated in FIG. 15 for example, it is also possible to reduce the size of at least one rendering plane to reduce the rendering processing load. In the figure, an example of reducing the size of the first plane is shown when using the first plane and the second plane. Note that the reduction of the rendering plane can be achieved by reducing the size of the buffer 18a (frame buffer) used as the rendering plane. The reduction of the rendering plane here means using a rendering plane whose size is reduced compared to other rendering planes. For the virtual object Vo rendered using the reduced rendering plane, when it is combined with the virtual object Vo rendered on other rendering planes, an enlargement process is performed according to the size of the virtual object Vo rendered on the other rendering plane and then the combination is carried out.
[0128] <6. Regarding the Virtual Object for Occlusion> In the AR system 50, as the virtual object Vo, in addition to the one that superimposes on the real object Ro other than the user as exemplified in FIG. 1, there can also be considered a virtual object that superimposes on a part of the user's body. As an example, when a part of the user's body overlaps with a virtual object as seen from the user's viewpoint position, a virtual object Vo (hereinafter referred to as "virtual object for shielding") that shields the overlapping part of the virtual object can be cited. As an example of this virtual object for shielding, a virtual object Vo that imitates the user's hand (a virtual object Vo that imitates the shape of the hand) can be cited. The virtual object for shielding can be paraphrased as region information that defines a shielding region for other virtual objects Vo.
[0129] For such a virtual object for shielding, image correction based on the object recognition result can also be performed. That is, image correction of the virtual object for shielding is performed based on the object recognition result for the corresponding part of the body. Specifically, in that case, the CPU 14 includes the virtual object for shielding as one of the "virtual objects Vo that superimpose on the real object Ro" and executes the process shown in FIG. 11 above.
[0130] FIG. 16 shows an example in which another virtual object Vo is shielded by a shielding virtual object. In this figure, the shielding virtual object is assumed to be in the shape of a user's hand, and an example is shown in which the first plane is used for drawing the shielding virtual object and the second plane is used for drawing another virtual object Vo. In this case, for the shielding virtual object drawn by the first plane, image correction is performed by the image correction processing unit 19a based on the object recognition result of the user's hand. In the illustrated example, correction is performed such that the shielding virtual object is enlarged in response to the user's hand moving to the front side. On the other hand, for another virtual object Vo drawn by the second plane, image correction is performed by the image correction processing unit 19a based on the object recognition result of the real object Ro corresponding to the other virtual object Vo. Then, the corrected images are combined and output to the display 10. At this time, if the shielding virtual object is located on the front side, the other virtual object Vo located on the back side thereof has the overlapping portion with the shielding virtual object shielded. In the illustrated example, the entire area of the other virtual object Vo overlaps with the shielding virtual object, and in this case, the entire area of the other virtual object Vo is shielded and becomes a non-display state.
[0131] By performing image correction based on the object recognition result for such a shielding virtual object, it is possible to suppress the display delay of the shielding virtual object. That is, it is possible to alleviate the discomfort that may occur to the user because the overlapping portion of the virtual object Vo is not shielded even though a part of the user's body overlaps with the virtual object Vo as viewed from the user's viewpoint position.
[0132] Here, when a plurality of drawing planes can be used, when performing image correction for the shielding virtual object, at least one of the drawing planes can be exclusively used as the drawing plane for the shielding virtual object. Thereby, it becomes possible to preferentially perform image correction based on the object recognition result for the shielding virtual object, and it becomes easier to further alleviate the discomfort that may occur to the user because the overlapping portion between a part of the user's body and the virtual object Vo is not shielded.
[0133] <7. Shadow display> In the display of the virtual object Vo, it is effective for improving the sense of reality to display the shadow (virtual shadow) of the virtual object Vo. Regarding the virtual shadow, it is required to follow the movement of the virtual object Vo. However, in order to suppress the display delay, for the image (hereinafter referred to as "virtual shadow image") on which the virtual shadow is drawn, similar to the image correction of the virtual object Vo, it is conceivable to perform image correction based on the latest object recognition result. However, if the virtual shadow image on which the virtual shadow is drawn is image-corrected based on the object recognition result in this way, there is a possibility that an appropriate shadow expression corresponding to the movement of the object cannot be performed.
[0134] FIG. 17 is an explanatory diagram of the problems when image correction based on the object recognition result is applied to the virtual shadow image. In FIG. 17A, an example of the virtual shadow Vs formed when the virtual object Vo is irradiated with light from the virtual light source Ls is illustrated. When the position of the virtual object Vo moves upward on the paper surface from the state shown in FIG. 17A, if the correction of the virtual shadow image is performed so that the virtual shadow Vs moves in the same direction and amount as the movement direction and amount of the virtual object Vo as shown in FIG. 17B, the correct shadow expression shown in FIG. 17C cannot be realized. As shown in FIG. 17C, in this case, as the virtual object Vo moves upward, the center of the shadow should be shifted to the left direction on the paper surface and the range of the shadow should be widened.
[0135] If image correction corresponding to the movement of the object is performed on the virtual shadow image on which the virtual shadow Vs is drawn in this way, it will not be possible to perform the correct shadow expression. Therefore, hereinafter, an information processing apparatus 1A for suppressing the display delay of the virtual shadow Vs while realizing the correct shadow expression will be described.
[0136] FIG. 18 is a block diagram showing an example of the internal configuration of the information processing apparatus 1A. In the following description, the same parts as those already described will be denoted by the same reference numerals and the description will be omitted. The differences from the information processing apparatus 1 shown in FIG. 3 are that a CPU 14A is provided instead of the CPU 14, and a display controller 19A is provided instead of the display controller 19. The display controller 19A is different from the display controller 19 in that it has an image correction processing unit 19aA instead of the image correction processing unit 19a. The image correction processing unit 19aA is different from the image correction processing unit 19a in that it has a function of performing image correction on the depth image as a shadow map, which will be described later. The CPU 14A is the same as the CPU 14 in terms of hardware configuration, but is different from the CPU 14 in that it performs processing related to the display of the virtual shadow Vs.
[0137] Hereinafter, a specific method for displaying the virtual shadow Vs will be described with reference to FIGS. 19 and 20. In this example, the shadow map method is used for displaying the virtual shadow Vs. The shadow map method is a method of drawing the virtual shadow Vs using a texture called a shadow map that stores the depth value (depth value) from the virtual light source Ls.
[0138] FIG. 19 is an explanatory diagram of the distances d1 and d2 used in the shadow map method. Basically, for the image Pcr having the same position as the viewpoint (drawing viewpoint) Pr when drawing the virtual object Vo as the viewpoint, pixels that become shadows are specified. Hereinafter, the image Pcr will be referred to as the drawing image Pcr. Also, the pixels constituting the drawing image Pcr will be referred to as pixel g1.
[0139] In the shadow map method, in identifying the pixel g1 that becomes a shadow in the drawing image Pcr, information on the distance d1 from each point p1 (indicated by an x mark in the figure) in the three-dimensional space projected onto each pixel g1 of the drawing image Pcr to the virtual light source Ls is used. In the figure, as an example of the point p1, the point p1 projected onto the pixel g1 1 projected onto the pixel g1 1 and the pixel g1 2 projected onto the pixel g1 2 are illustrated. Point p1 1 The distance from point p1 to the virtual light source Ls is distance d1 1 and the distance from point p1 2 to the virtual light source Ls is distance d1 2 as well.
[0140] In the shadow map method, as a shadow map, map information including an image of the virtual object Vo viewed from the position of the virtual light source Ls, specifically, a depth image of the virtual object Vo viewed with the virtual light source Ls as the viewpoint, is generated. Here, the depth image included in the shadow map, that is, the depth image of the virtual object Vo viewed with the virtual light source Ls as the viewpoint, is denoted as the light source viewpoint image Sm. Also, the pixel constituting the light source viewpoint image Sm is denoted as pixel g2. Furthermore, each point (indicated by ▲ in the figure) in the three-dimensional space projected onto each pixel g2 of the light source viewpoint image Sm is denoted as point p2. It can be paraphrased that the light source viewpoint image Sm as a depth image is an image representing the distance from each point p2 to the virtual light source Ls. Hereinafter, the distance from point p2 to the virtual light source Ls is denoted as distance d2.
[0141] In the shadow map, for each pixel g2 of the light source viewpoint image Sm, the corresponding pixel g1 in the drawn image Pcr and the distance d1 of that pixel g1 are mapped. In FIG. 19, the pixel g1 corresponding to the pixel g2 in the light source viewpoint image Sm 1 is pixel g1 1 and it shows that the pixel g1 corresponding to the pixel g2 2 is pixel g1 2 as well. Here, that a certain pixel g1 corresponds to a certain pixel g2 means that the point p2 projected onto the pixel g2 is located on the straight line connecting the point p1 projected onto the pixel g1 and the virtual light source Ls.
[0142] In the shadow map method, for the drawn image Pcr, using the shadow map in which the corresponding pixel g1 and the distance d1 of that pixel g1 are associated for each pixel g2 of the light source viewpoint image Sm in this way, it is determined for each pixel g1 whether it is a shadow part. Specifically, for the target pixel g1, the corresponding pixel g2 in the light source viewpoint image Sm is identified, and whether "d1 > d2" holds for the depth value of this pixel g2, that is, the distance d2, and the distance d1 of the target pixel g1 is used as the determination of whether it is in the shadow or not. For example, in the example in the figure, for the pixel g1 1 from the shadow map, the pixel g2 of the light source viewpoint image Sm 1 is identified as the corresponding pixel g2, and at the same time, the distance d1 1 of the pixel g1 1 (the distance d1 from the point p1 1 to the virtual light source Ls) and the distance d2 1 (the distance d2 from the point p2 1 to the virtual light source Ls) are identified. And since "d1 1 > d2 1 ", for the pixel g1 1 it is determined that it is in the shadow part. On the other hand, for the pixel g1 2 from the shadow map, the pixel g2 of the light source viewpoint image Sm 2 is identified as the corresponding pixel g2, and at the same time, the distance d1 2 of the pixel g1 2 (the distance d1 from the point p1 2 to the virtual light source Ls) and the distance d2 2 (the distance d2 from the point p2 2 to the virtual light source Ls) are identified, and their relationship is "d1 2 = d2 2 ", so for the pixel g1 2 it is determined that it is not in the shadow part.
[0143] Figure 20 is an explanatory diagram of the shadow range. Regarding the correspondence relationship between the pixel g1 and the pixel g2, it is represented by attaching the same value as the numerical value shown in the subscript at the end of the symbol. In the drawn image Pcr, for the pixel g1 5 the projected point p1 5 is the pixel g1 corresponding to the pixel g2 5 . This pixel g2 5is the pixel g2 onto which one end of the upper surface of the virtual object Vo (the surface facing the virtual light source Ls is defined as the upper surface) is projected in the light source viewpoint image Sm. Therefore, for the pixel g1 5 since the distance d1 > d2, it is determined to be in the shadow area. Also, for the pixel g1 6 the projected point p1 6 corresponds to the pixel g2 onto which the approximate center of the upper surface of the virtual object Vo is projected in the light source viewpoint image Sm 6 and for this pixel g1 6 since the distance d1 > d2, it is also in the shadow area. Furthermore, for the pixel g1 7 the projected point p1 7 corresponds to the pixel g2 onto which the other end of the upper surface of the virtual object Vo is projected in the light source viewpoint image Sm 7 and for this pixel g1 7 since the distance d1 > d2, it is also in the shadow area. As can be understood from these points, in the drawn image Pcr, the range from the pixel g1 5 through the pixel g1 6 to the pixel g1 7 is the shadow area due to the virtual object Vo.
[0144] Also, in the drawn image Pcr, the pixel g1 8 the projected point p1 8 corresponds to the pixel g2 onto which the approximate center of the side surface of the virtual object Vo is projected in the light source viewpoint image Sm 8 and for this pixel g1 8 since the distance d1 > d2, it is also in the shadow area.
[0145] For confirmation, FIG. 20 schematically shows the plan view of the light source viewpoint image Sm. Thus, the light source viewpoint image Sm can be represented as an image on which the virtual object Vo is projected.
[0146] Here, as described above, if the virtual shadow image Vs is corrected based on the latest object recognition result in the same way as the image correction of the virtual object Vo, there is a risk that appropriate shadow expression according to the movement of the object cannot be performed. Therefore, in this example, instead of performing image correction based on the latest object recognition result on the virtual shadow image, a method is adopted in which the image correction is performed on the light source viewpoint image Sm used for generating the virtual shadow image in the shadow map method.
[0147] FIG. 21 is an explanatory diagram of the image correction of the light source viewpoint image Sm. Specifically, in FIG. 21, a method of image correction of the light source viewpoint image Sm corresponding to the case where the virtual object Vo moves from the position indicated by the dotted line to the position indicated by the solid line is illustrated.
[0148] Here, the generation of the light source viewpoint image Sm (that is, the generation of the shadow map) is performed based on the position of the real object Ro recognized in the object recognition process at a certain time point. The image correction of the light source viewpoint image Sm here corrects the light source viewpoint image Sm generated based on the position of the real object Ro at a certain time point based on the position of the real object Ro recognized in the object recognition process at a time point later than the certain time point. Since the virtual shadow image is an image of the shadow of the virtual object Vo, in order to realize appropriate shadow expression, the recognition results of the real object Ro used as the reference in the image correction of the light source viewpoint image Sm and the image correction of the virtual object Vo need to be common. In other words, it is necessary to perform the image correction of the light source viewpoint image Sm and the image correction of the virtual object Vo using the recognition result of the real object Ro at the same time point.
[0149] FIG. 22 is a timing chart showing the flow of processing related to the image correction of the light source viewpoint image Sm and the image correction of the virtual object Vo. Regarding the image correction of the virtual object Vo, for example, drawing processing (refer to drawing (object) in the figure) based on the result of the object recognition process at a certain time point represented by time point t1 in the figure is performed, and after the completion of this drawing processing, the virtual object Vo that has been drawn is corrected based on the result of the latest object recognition process (refer to time point 2 in the figure). For the virtual shadow image, since a shadow image adjusted according to the position of the virtual object Vo corrected in this way should be generated, the image correction of the light source viewpoint image Sm uses the object recognition result at time t2, which was used as a reference in the image correction of the virtual object Vo.
[0150] Specifically, in this case, the generation of the shadow map is performed based on the result of the object recognition process at time t1. That is, as the light source viewpoint image Sm, an image based on the position of the real object Ro at time t1 is generated. After the rendering process for the virtual object Vo is completed, in response to obtaining the latest object recognition process result at time t2, the light source viewpoint image Sm is corrected based on this latest object recognition process result.
[0151] Here, an example was given in which the generation process of the shadow map is performed based on the result of the object recognition process at time t1. However, the generation process of the shadow map may be performed based on the result of any object recognition process obtained before the completion of the rendering process of the virtual object Vo.
[0152] Returning to the explanation in Fig. 21. As described above, the light source viewpoint image Sm is generated based on the result of the object recognition process at a certain time (time t1). In the figure, the virtual object Vo at that certain time in the light source viewpoint image Sm is shown by a dashed line. After the rendering process for the virtual object Vo is completed, in response to obtaining the latest object recognition process result at time t2, the moving direction and moving amount of the virtual object Vo from time t1 can be specified. According to the moving direction and moving amount of the virtual object Vo specified in this way, the image area of the virtual object Vo in the light source viewpoint image Sm is corrected. Specifically, in the light source viewpoint image Sm in the figure, the image area of the virtual object Vo indicated by the dotted line is corrected to be the image area indicated by the solid line.
[0153] At this time, the image correction of the light source viewpoint image Sm is performed as at least one of the corrections of the position and size of the image area of the virtual object Vo. In this example, in order to correspond to both the displacement of the virtual object Vo in the direction of the distance d2 and the displacement in the direction parallel to the image plane of the light source viewpoint image Sm, in the image correction of the light source viewpoint image Sm, it is possible to correct both the size and the position of the image area of the virtual object Vo. In the illustrated example, the virtual object Vo approaches the virtual light source Ls side in the direction of the distance d2 and is displaced to the left end side of the light source viewpoint image Sm in the direction parallel to the image plane. Therefore, in the image correction of the light source viewpoint image Sm in this case, correction is performed to increase the size of the image area of the virtual object Vo and displace it to the left side of the image.
[0154] Note that in FIG. 21, the virtual object Vo and the virtual shadow Vs projected onto the drawn image Pcr are schematically shown. Also in the drawn image Pcr, for the virtual object Vo, the one before movement is represented by a dashed line and the one after movement is represented by a solid line, respectively. For the virtual shadow Vs, the one generated for the virtual object Vo before movement is represented by a dashed line and the one generated for the virtual object Vo after movement is represented by a solid line, respectively.
[0155] FIG. 23 is an explanatory diagram showing the relationship between the image correction of the light source viewpoint image Sm and the pixel g1 (corresponding pixel) in the drawn image Pcr mapped to each pixel g2 of the light source viewpoint image Sm in the shadow map. FIG. 23A illustrates the drawn image Pcr and the light source viewpoint image Sm generated based on the object recognition result at a certain point in time (time point t1), and shows the correspondence between the pixel g2 and the pixel g1 in the shadow map. Here, taking the coordinate system of the drawn image Pcr as the xy coordinate system and the coordinate system of the light source viewpoint image Sm as the uv coordinate system, the coordinates of each pixel g1 and g2 are shown together. Specifically, in FIG. 23A, as an example of the pixel g2, the pixel g2 1 , g2 2 , g2 3 are shown, and the pixels g1 1 , g1 2 , g1 3 of the drawn image Pcr corresponding to these pixels g2 are illustrated. As shown in the figure, the pixels g2 1 , g2 2 , g2 3Let the coordinates be (u1, v1), (u2, v2), and (u3, v3) respectively, and the pixels be g1 1 , g1 2 , g1 3 The coordinates of are (x1, y1), (x2, y2), and (x3, y3) respectively.
[0156] In FIG. 23B, an example of a drawn image Pcr and a light source viewpoint image Sm that are image-corrected based on an object recognition result obtained at a time point t2 after the time point t1 is illustrated. Specifically, as the image correction (2D correction) of the light source viewpoint image Sm here, in accordance with the displacement of the virtual object Vo shown as the transition from FIG. 23A to FIG. 23B, while expanding the image area of the virtual object Vo, it is performed in a manner of shifting the position downward. At this time, no correction is performed on the correspondence relationship between the pixel g2 and the pixel g1. That is, for example, for the pixel g2 1 , the pixel g1 1 corresponds, and for the pixel g2 2 , the pixel g1 2 corresponds, and for the pixel g2 3 , the pixel g1 3 corresponds. The mapping information with each pixel g1 on the side of the drawn image Pcr is maintained without correction.
[0157] As described above, in this example, in order to suppress the display delay of the virtual shadow Vs, a method of performing image correction of the light source viewpoint image Sm based on the object recognition result is adopted. Thereby, compared with the case of performing image correction of the virtual shadow image based on the object recognition result (see FIG. 17B), the accuracy of the expression of the virtual shadow Vs can be improved.
[0158] Here, in FIG. 22, the amount of display delay for the virtual shadow Vs is represented by double-headed arrows indicated as "delay" in the figure. From this delay amount, it can be seen that for the virtual shadow Vs as well, display delay can be suppressed in the same manner as in the case of the virtual object Vo.
[0159] An example of a specific processing procedure to be executed to implement the shadow display method described above will be described with reference to the flowchart of FIG. 24. Note that in FIG. 24, as an example of the processing procedure, the processing procedure executed by the CPU 14A shown in FIG. 18 is illustrated.
[0160] First, in step S301, the CPU 14A waits for the start of rendering of the virtual object Vo, and in response to the start of rendering of the virtual object Vo, executes the generation process of the shadow map. This generation process of the shadow map is performed based on the result of the same object recognition process as the rendering process of the virtual object Vo whose start was confirmed in step S301. Specifically, as this generation process of the shadow map, the CPU 14A generates the light source viewpoint image Sm based on the result of the object recognition process, and for each point p1 in the three-dimensional space projected onto each pixel g1 of the rendering image Pcr based on the result of the object recognition process, calculates the distance d1 respectively. In addition to this, the CPU 14A specifies the corresponding pixel g1 in the rendering image Pcr for each pixel g2 of the light source viewpoint image Sm, and performs the process of associating the coordinate information of the corresponding pixel g1 with the distance d1 for each pixel g2. Thereby, the shadow map is generated.
[0161] In response to performing the generation process of the shadow map in step S302, the CPU 14A waits until the rendering of the virtual object Vo is completed in step S303, and when the rendering of the virtual object Vo is completed, proceeds to step S304 and waits until the latest object recognition result is obtained.
[0162] In response to obtaining the latest object recognition result in step S304, the CPU 14A proceeds to step S305 and performs correction control of the shadow map based on the object recognition result. Specifically, for the light source viewpoint image Sm obtained in the generation process of the shadow map in step S302, the CPU 14A causes the image correction processing unit 19aA in the display controller 19A to execute image correction. At this time, as the image correction, as described above, in accordance with the movement of the virtual object Vo (the movement from time point t1 to time point t2) specified from the latest object recognition result, it is executed in such a manner that at least either the position or the size of the image area of the virtual object Vo in the light source viewpoint image Sm is changed. Specifically, in this example, as described above, both the position and the size of the image area of the virtual object Vo can be corrected.
[0163] In step S306 following step S305, the CPU 14A performs a process of generating a shadow image based on the corrected shadow map. That is, a virtual shadow image is generated based on the shadow map including the light source viewpoint image Sm corrected by the correction control in step S305. As described above, for generating a virtual shadow image based on a shadow map, for each pixel g1 in the drawing image Pcr, its distance d1 and the distance d2 of the corresponding pixel g2 in the light source viewpoint image Sm are specified, and it is determined whether "d1 > d2" for these distances d1 and d2. Then, for the pixel g1 determined to be "d1 > d2", a virtual shadow image is generated by performing shadow drawing.
[0164] In step S307 following step S306, the CPU 14A performs a process of synthesizing the corrected virtual object image and the shadow image. That is, a process is performed to cause the display controller 19A to synthesize the drawing image of the virtual object Vo for which image correction has been performed with reference to FIGS. 6 to 13 and the virtual shadow image generated in step S306. Then, in step S308 following step S307, as an output process of the synthesized image, the CPU 14A performs a process of causing the display controller 19A to output the image synthesized in step S307 to the display 10. In response to executing the process of step S309, the CPU 14A finishes the series of processes shown in FIG. 24.
[0165] In the above, regarding the image correction of the light source viewpoint image Sm, an example of changing the size and position of the image area of the virtual object Vo has been given. However, in addition to the change in size (i.e., scaling) and the change in position, for example, deformation, rotation, etc. can also be considered.
[0166] Also, in the above, regarding the image correction of the light source viewpoint image Sm, an example of performing correction based on the object recognition result has been described. However, image correction based on the detection signal of the sensor unit 13 that detects the viewpoint position and line-of-sight direction of the user can also be performed.
[0167] <8. Modification Example> Here, the present embodiment is not limited to the specific examples exemplified above, and various modification examples can be considered. For example, in the above, an example of superimposing and displaying the virtual object Vo on the real object Ro was given, but it is not essential to superimpose the virtual object Vo on the real object Ro. For example, even without superimposition, a case where the virtual object Vo is displayed so as to maintain a predetermined positional relationship with the real object Ro can be considered. The present technology can be widely and preferably applied to cases where the virtual object Vo is displayed in association with the real object Ro, such as superimposing and displaying the virtual object Vo on the real object Ro or displaying it so as to maintain a predetermined positional relationship.
[0168] Also, in the above, a configuration was exemplified in which the imaging unit 11 for obtaining an imaging image for performing object recognition, the sensor unit 13 for detecting information on the user's line-of-sight position and line-of-sight direction, the display 10 for performing image display for allowing the user to recognize the AR space, and the correction control unit (CPU 14) for controlling image correction for the image on which the virtual object Vo is drawn are provided in the same device as the information processing device 1. However, a configuration in which the imaging unit 11, the sensor unit 13, and the display 10 are provided in a head-mounted device and the correction control unit is provided in a device different from the head-mounted device can also be adopted.
[0169] Also, in the above, an example of a head-mounted display device (HMD) was exemplified as a see-through type HMD, but in addition to this, a video see-through type HMD and a retinal projection type HMD can also be mentioned.
[0170] When a video see-through type HMD is worn on a user's head or face, it is worn so as to cover the user's eyes, and a display unit such as a display is held in front of the user's eyes. Further, the video see-through type HMD has an imaging unit for imaging the surrounding scenery, and an image of the scenery in front of the user imaged by the imaging unit is displayed on the display unit. With such a configuration, although it is difficult for a user wearing a video see-through type HMD to directly bring the external scenery into the field of view, it becomes possible to check the external scenery by the image displayed on the display unit. Also, at this time, the video see-through type HMD may superimpose a virtual object on the image of the external scenery, for example, according to the recognition result of at least either the position or the orientation of the video see-through type HMD based on AR technology.
[0171] A retinal projection type HMD has a projection unit held in front of the user's eyes, and the image is projected from the projection unit toward the user's eyes so that the image is superimposed on the external scenery. More specifically, in the retinal projection type HMD, the image is directly projected from the projection unit onto the retina of the user's eyes, and the image is formed on the retina. With such a configuration, even in the case of a myopic or hyperopic user, it becomes possible to view a clearer video. Also, a user wearing a retinal projection type HMD can bring the external scenery into the field of view even while viewing the image projected from the projection unit. With such a configuration, the retinal projection type HMD can also superimpose an image of a virtual object on the optical image of a real object located in the real space, for example, according to the recognition result of at least either the position or the orientation of the retinal projection type HMD based on AR technology.
[0172] In addition, in the above, an example of providing the sensor unit 13 as a configuration for estimating the viewpoint position and line-of-sight direction of the user has been given. However, the viewpoint position and line-of-sight direction of the user can also be estimated by the following method. For example, the information processing apparatus 1 images a marker or the like with a known size presented on the real object Ro in the real space by an imaging unit such as a camera provided in itself. Then, the information processing apparatus 1 estimates at least either the relative position or the relative orientation of itself with respect to the marker (and thus the real object Ro on which the marker is presented) by analyzing the captured image. Specifically, it is possible to estimate the relative direction of the imaging unit (and thus the information processing apparatus 1 including the imaging unit) with respect to the marker according to the orientation of the marker imaged in the image (for example, the orientation of the pattern of the marker). Also, when the size of the marker is known, it is possible to estimate the distance between the marker and the imaging unit (that is, the information processing apparatus 1 including the imaging unit) according to the size of the marker in the image. More specifically, when imaging the marker from a farther distance, the marker will be imaged smaller. Also, at this time, the range in the real space imaged in the image can be estimated based on the angle of view of the imaging unit. By utilizing the above characteristics, it is possible to calculate the distance between the marker and the imaging unit in reverse according to the size of the marker imaged in the image (in other words, the ratio occupied by the marker within the angle of view). With the above configuration, the information processing apparatus 1 can estimate its relative position and orientation with respect to the marker. Consequently, it becomes possible to estimate the viewpoint position and line-of-sight direction of the user.
[0173] In addition, a technique called so-called SLAM (simultaneous localization and mapping) may be used for estimating the self-position of the information processing apparatus 1. SLAM is a technique that performs self-position estimation and creation of an environmental map in parallel by using an imaging unit such as a camera, various sensors, an encoder, or the like. As a more specific example, in SLAM (particularly, Visual SLAM), based on a moving image captured by the imaging unit, the three-dimensional shape of the captured scene (or subject) is sequentially restored. Then, by associating the restoration result of the captured scene with the detection result of the position and orientation of the imaging unit, creation of a map of the surrounding environment and estimation of the position and orientation of the imaging unit (and thus, the information processing apparatus 1) in the environment are performed. Regarding the position and orientation of the imaging unit, for example, by providing various sensors such as an acceleration sensor and an angular velocity sensor in the information processing apparatus 1, it is possible to estimate, as information indicating a relative change, based on the detection result of the sensor. Of course, if the position and orientation of the imaging unit can be estimated, the method is not necessarily limited to only the method based on the detection results of various sensors such as an acceleration sensor and an angular velocity sensor.
[0174] Under the configuration as described above, for example, the estimation result of the relative position and orientation of the information processing apparatus 1 with respect to the marker based on the imaging result of a known marker by the imaging unit may be used for the initialization process and position correction in the above-described SLAM. With such a configuration, even in a situation where the marker is not included within the field of view angle of the imaging unit, the information processing apparatus 1 can estimate its own position and orientation with respect to the marker (and thus, the real object Ro to which the marker is presented) by self-position estimation based on SLAM that has received the results of the previously executed initialization and position correction.
[0175] In the above, it is premised on estimating the user's line-of-sight direction from the posture of the information processing apparatus 1 (head-mounted device), but a configuration may be adopted in which the user's line-of-sight direction is detected based on an imaging image or the like of the user's eyes.
[0176] Note that the target of suppressing display delay by image correction is not limited to the virtual object Vo displayed in association with the real object Ro. For example, in an AR game or the like, for a virtual object Vo such as an avatar of another user as an opponent player, the position data thereof in the AR space is received via a network, and the information processing apparatus 1 displays the virtual object Vo at a position according to the received position data. In this case, display delay suppression by image correction may be achieved for the virtual object Vo to be displayed. The image correction in this case is performed not based on the recognition result of the real object Ro, but based on the amount of change in the position indicated by the position data received via the network.
[0177] Also, the image correction may be performed in units of tiles (segment units) instead of in units of planes. Furthermore, as correction for suppressing the display delay of the virtual object Vo, correction within the drawing process may be considered instead of correction for the image after drawing. For example, full-fledged rendering and simple rendering that can be performed in real time are separated. At this time, each virtual object Vo is rendered as a billboard in the front-stage full-fledged rendering, and only the synthesis of the billboards is performed in the back-stage simple rendering. Alternatively, as correction for suppressing the display delay of the virtual object Vo, a method of replacing with a matrix based on the latest object recognition result immediately before drawing by the GPU may also be considered.
[0178] Also, when the virtual object Vo superimposed on the real object Ro has an animation, information specifying the animation may be instructed to the image correction processing unit 19a. Specifically, there may also be an animation in which the size and color change according to the object recognition result. For example, the brightness changes when twisted. A configuration may also be adopted in which such a change in the virtual object Vo is realized by image correction by the image correction processing unit 19a.
[0179] Also, when the virtual object Vo is a human face, image correction corresponding to mesh deformation can also be performed. For example, assuming that the object recognition result is a face landmark, when image correction should be performed based on this, the landmark information for the rendering result is instructed to the image correction processing unit 19a to execute image correction based on the landmark.
[0180] <9. Program and Storage Medium> As described above, the information processing apparatus (the same as 1) as an embodiment has been described. However, the program of the embodiment is a program that causes a computer device such as a CPU to execute the processing as the information processing apparatus 1.
[0181] The program of the embodiment is a program readable by a computer device. Based on a captured image including a real object, a first recognition process regarding the position and orientation of the real object at a first time point is performed, and the rendering processing unit is controlled to perform a first rendering process for a related virtual object associated with the real object based on the first recognition process. At a second time point after the first time point, based on a captured image including the real object, a second recognition process regarding the position and orientation of the real object is performed, and the rendering processing unit is controlled to perform a second rendering process for a related virtual object associated with the real object based on the second recognition process. Before the second rendering process is completed, based on the result of the second recognition process, a process of correcting a first image of the related virtual object obtained by the completion of the first rendering process is caused to be executed by the computer device. That is, this program corresponds to a program that causes a computer device to execute the processing described with reference to FIGS. 11 to 13 and the like, for example.
[0182] Such a program can be pre-stored in a computer-readable storage medium, such as a ROM, SSD (Solid State Drive), HDD (Hard Disk Drive), etc. Alternatively, it can also be temporarily or permanently stored (recorded) in a removable storage medium such as a semiconductor memory, memory card, optical disk, magneto-optical disk, magnetic disk, etc. Such a removable storage medium can be provided as so-called packaged software. Also, such a program can be installed from a removable storage medium into a personal computer or the like, or downloaded from a download site to a required information processing device such as a smartphone via a network such as a LAN (Local Area Network) or the Internet.
[0183] <10. Summary of Embodiments> As described above, the information processing apparatus (the same as 1, 1A) as an embodiment includes an image recognition processing unit (the same as F1) that performs a first recognition process regarding the position and orientation of a real object at a first time point and a second recognition process regarding the position and orientation of the real object at a second time point after the first time point based on a captured image including the real object, a drawing control unit (the same as F2) that controls a drawing processing unit (GPU17) to perform a first drawing process for a related virtual object associated with the real object based on the first recognition process and a second drawing process for a related virtual object associated with the real object based on the second recognition process, and a correction control unit (image correction control unit F3) that corrects a virtual object image, which is an image of the related virtual object obtained by the completion of the first drawing process, based on the result of the second recognition process before the completion of the second drawing process.
[0184] By performing image correction of related virtual objects based on the recognition results of the position and orientation of the real object as described above, when the position or orientation of the real object changes, it becomes possible to change the position and orientation of the related virtual objects following the change. And according to the above configuration, for the image of the related virtual object, as long as the latest recognition result (the recognition result of the second recognition process) is obtained, without waiting for the completion of the drawing process (the second drawing process) based on the latest recognition result, it can be immediately output as an image obtained by correcting the image obtained by the drawing process (the first drawing process) based on the past recognition result. Therefore, regarding the image of the virtual object displayed in association with the real object, it is possible to suppress display delay, alleviate the user's sense of discomfort, and enhance the immersion in the AR space.
[0185] Also, in the information processing apparatus as an embodiment, the correction control unit performs correction to change the position of the related virtual object in the in-plane direction of up, down, left, and right of the virtual object image based on the information of the position of the real object recognized by the image recognition processing unit (see FIG. 7).
[0186] Thereby, when the real object moves in the in-plane direction of up, down, left, and right, it becomes possible to perform image correction to change the position of the related virtual object in the in-plane direction of up, down, left, and right according to the movement. Therefore, it is possible to suppress the display delay with respect to the movement of the real object in the in-plane direction of up, down, left, and right.
[0187] Furthermore, in the information processing apparatus as an embodiment, the correction control unit performs the above correction to change the size of the related virtual object for the virtual object image based on the information of the position of the real object recognized by the image recognition processing unit in the depth direction.
[0188] Accordingly, when the real object moves in the depth direction, for example, when the real object approaches the user's viewpoint, the image of the related virtual object can be greatly changed, or conversely, when the real object moves away from the viewpoint, the image of the related virtual object can be made smaller. Thus, it is possible to change the size of the related virtual object according to the position of the real object in the depth direction. Therefore, it is possible to suppress the display delay with respect to the movement of the real object in the depth direction.
[0189] Furthermore, in the information processing apparatus according to the embodiment, the correction control unit performs correction to change the position or orientation of the related virtual object according to a change in the viewpoint position or the line-of-sight direction of the user.
[0190] Accordingly, it is possible to suppress the display delay caused by a change in the viewpoint position or the line-of-sight direction when the user moves the head or the like. Therefore, it is possible to realize an AR system that allows the user to move the head and the line of sight, and in terms of not restricting the free movement of the user's body, it is possible to enhance the sense of immersion in the AR space.
[0191] Also, in the information processing apparatus according to the embodiment, when the correction control unit selects one or a plurality of related virtual objects to be corrected from among a plurality of related virtual objects each associated with a different real object, the related virtual object of the real object with a large movement is preferentially selected (see step S109 in FIG. 11).
[0192] Accordingly, it is possible to prevent the image correction from being blindly executed for the virtual object associated with the real object with a small movement or no movement. Therefore, it is possible to reduce the processing load in suppressing the display delay.
[0193] Furthermore, in the information processing apparatus according to the embodiment, the processing cycle of the correction is set to be shorter than the processing cycle of the image recognition processing unit (see FIG. 14).
[0194] As a result, it becomes possible to shorten the delay time from the point in time when the recognition result of the real object is obtained until the image correction of the virtual object starts. Therefore, the effect of suppressing the display delay of the virtual object can be enhanced. Also, since the correction processing cycle is short, the virtual object can be displayed smoothly.
[0195] Furthermore, in the information processing apparatus according to the embodiment, the drawing control unit controls the drawing processing unit to draw the related virtual object and the non-related virtual object, which is a virtual object independent of the image recognition processing of the real object, on different drawing planes in a plurality of drawing planes (see FIG. 11).
[0196] As a result, it becomes possible to perform appropriate image correction according to whether the virtual object is a related virtual object or not, such as performing image correction according to the user's viewpoint position and line-of-sight direction for the non-related virtual object, and performing image correction according to the position and posture of the associated real object and the viewpoint position and line-of-sight direction for the related virtual object. Therefore, it is possible to appropriately suppress the display delay of the virtual object.
[0197] Also, in the information processing apparatus according to the embodiment, when the number of related virtual objects is greater than or equal to the number of a plurality of drawing planes, and the number of the plurality of drawing planes is n (n is a natural number), the drawing control unit selects n - 1 related virtual objects, exclusively draws the selected related virtual objects on at least one drawing plane, and controls the drawing processing unit to draw the unselected related virtual objects and the non-related virtual objects on the remaining one drawing plane (see FIG. 11).
[0198] As a result, when the virtual object includes an irrelevant virtual object for which image correction based on the object recognition result is unnecessary as a virtual object, and the number of relevant virtual objects is n or more with respect to the number n of drawing planes, image correction based on the recognition result of the relevant real object is performed for n - 1 relevant virtual objects, and for the remaining relevant virtual objects, image correction based on the user's viewpoint position and line-of-sight direction is performed together with the irrelevant virtual objects. That is, when it is impossible to perform image correction based on the recognition result of the real object for all relevant virtual objects due to the relationship between the number of drawing planes and the number of relevant virtual objects, image correction based on the recognition result of the real object is preferentially performed for n - 1 relevant virtual objects. Therefore, it is possible to suppress an appropriate display delay according to the relationship between the number of drawing planes and the number of relevant virtual objects.
[0199] Furthermore, in the information processing apparatus according to the embodiment, the drawing control unit makes the selection using a selection criterion in which the possibility of selection increases as the amount of movement of the real object increases (see step S109 in FIG. 11).
[0200] As a result, it is possible to preferentially select a relevant virtual object with a large amount of movement and a high possibility of perceiving display delay as an object of image correction based on the recognition result of the real object. Therefore, when it is possible to perform image correction based on the recognition result of the real object for only a part of the relevant virtual objects, it is possible to appropriately select the relevant virtual objects to be the object of the image correction.
[0201] Furthermore, in the information processing apparatus according to the embodiment, the drawing control unit makes the selection using a selection criterion in which the possibility of selection increases as the area of the real object decreases.
[0202] When a related virtual object is superimposed on a real object, even if the amount of movement of the real object is large, if the area of the real object is large, the ratio of the area of the position error generation part of the related virtual object to the area of the real object may be small, and in such a case, it becomes difficult to perceive the display delay. On the other hand, even if the amount of movement of the real object is small, if the area of the real object is small, the ratio may be large, and in such a case, it becomes easy to perceive the display delay. Therefore, according to the above configuration, considering such a ratio of the area of the position error generation part of the virtual object to the area of the real object, it is possible to appropriately select a related virtual object to be the target of image correction based on the object recognition result.
[0203] Also, in the information processing apparatus as an embodiment, the drawing control unit makes the selection using a selection criterion in which the possibility of selection increases as the distance between the user's gaze point and the real object becomes shorter.
[0204] Thereby, it becomes possible to select, as the target of image correction based on the recognition result of the real object, a related virtual object that is displayed near the user's gaze point and has a high possibility of being perceived as a display delay. Therefore, when only a part of the related virtual objects can be subjected to image correction based on the recognition result of the real object, it is possible to appropriately select the related virtual object to be the target of the image correction.
[0205] Furthermore, in the information processing apparatus as an embodiment, the drawing control unit controls the drawing processing unit so as to reduce the update frequency of the drawing plane for drawing an unrelated virtual object independent of the image recognition processing of the real object among a plurality of drawing planes to be lower than the update frequency of the drawing plane for drawing a related virtual object (see FIG. 11).
[0206] Thereby, it is possible to prevent the drawing of all drawing planes from being performed at a high update frequency. Therefore, it is possible to reduce the processing load and the power consumption.
[0207] Furthermore, in the information processing apparatus according to an embodiment, when the related virtual object is a related virtual object that performs animation, the drawing control unit controls the drawing processing unit to reduce the drawing update frequency of the related virtual object compared to the case where the related virtual object does not perform animation.
[0208] Thereby, when it is necessary to use a plurality of drawing planes, if the related virtual object to be drawn does not perform animation, the drawing of the related virtual object is performed at a low update frequency, and if it performs animation, the drawing of the related virtual object is performed at a high update frequency. Therefore, it is possible to achieve both the reduction of the processing load and the reduction of power consumption by reducing the drawing update frequency of at least one drawing plane, and the prevention of the reduction of the reproducibility of the animation of the related virtual object.
[0209] Also, in the information processing apparatus according to an embodiment, when drawing processing is performed on a plurality of drawing planes, the drawing control unit controls the drawing processing unit to use at least one drawing plane having a smaller size than the other drawing planes (see FIG. 15).
[0210] Thereby, when it is necessary to use a plurality of drawing planes, it is possible to reduce the processing load of the drawing processing. Therefore, it is possible to reduce the processing load and the power consumption in order to reduce the display delay of the virtual object.
[0211] Furthermore, in the information processing apparatus according to an embodiment, when a part of the user's body overlaps a virtual object as viewed from the user's viewpoint position, the correction control unit performs the correction on a shielding virtual object that is a virtual object that shields the overlapping part of the virtual object (see FIG. 16).
[0212] This makes it possible to suppress the display delay for the virtual object for shielding. Therefore, it is possible to alleviate the discomfort that may occur because the overlapping part of the virtual object is not shielded even though a part of the user's body overlaps the virtual object as seen from the user's viewpoint position. By alleviating such discomfort, it is possible to enhance the sense of immersion in the AR space.
[0213] Furthermore, in the information processing apparatus as an embodiment, the virtual object for shielding is a virtual object imitating the user's hand.
[0214] This makes it possible to suppress the display delay for the virtual object for shielding that imitates the user's hand. Therefore, it is possible to alleviate the discomfort that may occur because the overlapping part of the virtual object is not shielded even though the user's hand overlaps the virtual object as seen from the user's viewpoint position. By alleviating such discomfort, it is possible to enhance the sense of immersion in the AR space.
[0215] Also, in the information processing apparatus as an embodiment, the drawing control unit controls the drawing processing unit to exclusively use at least one of the plurality of drawing planes that can be used by the drawing processing unit for the virtual object for shielding.
[0216] This makes it possible to preferentially perform image correction based on the object recognition result for the virtual object for shielding. Therefore, it is possible to more easily alleviate the discomfort that may occur because the overlapping part between a part of the user's body and the virtual object is not shielded, and it is possible to further improve the sense of immersion in the AR space.
[0217] Furthermore, in the information processing apparatus as an embodiment, before the completion of the first drawing process, based on the result of the first recognition process, a light source viewpoint image (the same as Sm) which is an image of the related virtual object viewed from the position of a virtual light source (the same as Ls) that illuminates the related virtual object is generated, and control is performed so that the generated light source viewpoint image is corrected based on the result of the second recognition process before the completion of the second drawing process. Further, a virtual shadow image generation unit (for example, CPU 14A) that generates a virtual shadow image which is an image of a virtual shadow for the virtual related object based on the corrected light source viewpoint image is provided.
[0218] As a result, for the light source viewpoint image used for generating the virtual shadow image, even when the target real object moves, it is possible to immediately correct and use the image generated based on the past recognition result (the result of the first recognition process) based on the latest recognition result (the result of the second recognition process). When improving the realism by displaying the shadow (virtual shadow) of the related virtual object, it is possible to suppress the display delay of the shadow. Therefore, it is possible to alleviate the discomfort of the user caused by the display delay of the shadow, and enhance the sense of immersion in the AR space.
[0219] Furthermore, in the information processing apparatus as an embodiment, the virtual shadow image generation unit, before the completion of the first drawing process, based on the result of the first recognition process, calculates the distance (the same as d1) from each point (point p1) in the three-dimensional space projected onto each pixel (pixel g1) of the drawing image by the drawing process unit to the virtual light source as the distance between the drawing-side light sources, respectively. As the light source viewpoint image, a depth image as a shadow map by the shadow map method is generated. As the correction of the light source viewpoint image, a process of changing the position or size of the image area of the related virtual object in the shadow map based on the result of the second recognition process is performed, and a virtual shadow image is generated based on the corrected shadow map and the distance between the drawing-side light sources.
[0220] That is, in the generation of the virtual shadow image by the shadow map method, correction is performed to change the position or size of the image area of the real object in the shadow map generated based on the result of the first recognition process based on the latest object recognition result (the result of the second recognition process). Accordingly, when improving the sense of reality by displaying the shadow of an associated virtual object, it becomes possible to suppress the display delay of the shadow, and by reducing the discomfort of the user caused by the display delay of the shadow, the sense of immersion in the AR space can be enhanced.
[0221] In addition, the control method as an embodiment performs a first recognition process regarding the position and orientation of a real object at a first time point based on a captured image including the real object, and controls a drawing processing unit to perform a first drawing process on an associated virtual object associated with the real object based on the first recognition process. At a second time point after the first time point, a second recognition process regarding the position and orientation of the real object is performed based on a captured image including the real object, and the drawing processing unit is controlled to perform a second drawing process on an associated virtual object associated with the real object based on the second recognition process. Before the second drawing process is completed, based on the result of the second recognition process, it is a control method for correcting a first image of the associated virtual object obtained by the completion of the first drawing process. Also by such a control method as an embodiment, the same operations and effects as those of the information processing apparatus as the above-described embodiment can be obtained.
[0222] In addition, the program of the embodiment is a program readable by a computer device, which performs a first recognition process regarding the position and orientation of a real object at a first time point based on a captured image including the real object, and controls a drawing processing unit to perform a first drawing process on an associated virtual object associated with the real object based on the first recognition process. At a second time point after the first time point, a second recognition process regarding the position and orientation of the real object is performed based on a captured image including the real object, and the drawing processing unit is controlled to perform a second drawing process on an associated virtual object associated with the real object based on the second recognition process. Before the second drawing process is completed, based on the result of the second recognition process, it is a program for causing a computer device to execute a process of correcting a first image of the associated virtual object obtained by the completion of the first drawing process. Furthermore, the storage medium of the embodiment is a storage medium storing the program as the above-described embodiment. Such a program and storage medium can realize the information processing apparatus as the above-described embodiment.
[0223] Note that the effects described in this specification are merely illustrative and not limiting, and there may be other effects.
[0224] <10. The present technology> Note that the present technology can also adopt the following configurations. (1) An image recognition processing unit that performs a first recognition process related to the position and orientation of the real object at a first time point and a second recognition process related to the position and orientation of the real object at a second time point after the first time point, based on a captured image including the real object; A drawing control unit that controls a drawing processing unit to perform a first drawing process on a related virtual object associated with the real object based on the first recognition process and a second drawing process on a related virtual object associated with the real object based on the second recognition process; A correction control unit that corrects a first image of the related virtual object obtained by completion of the first drawing process based on the result of the second recognition process before completion of the second drawing process. An information processing apparatus comprising an information processing apparatus. (2) The correction control unit performs the correction of changing the position of the related virtual object in the in-plane direction of up, down, left, and right of the related virtual object in the first image based on the information on the position of the real object recognized by the image recognition processing unit in the in-plane direction of up, down, left, and right of the real object. The information processing apparatus according to (1) above. (3) The correction control unit performs the correction of changing the size of the related virtual object in the first image based on the information on the position of the real object recognized by the image recognition processing unit in the depth direction of the real object. The information processing apparatus according to (1) or (2) above. (4) The correction control unit performs the correction of changing the position or orientation of the related virtual object according to a change in the viewpoint position or line-of-sight direction of the user. The information processing apparatus according to any one of (1) to (3) above. (5) The correction control unit When selecting one or more associated virtual objects to be the target of the correction from among a plurality of associated virtual objects each associated with a different real object, preferentially selects the associated virtual object of the real object with a large movement The information processing apparatus according to any one of (1) to (4) above. (6) The processing cycle of the correction is set to be shorter than the processing cycle of the image recognition processing unit The information processing apparatus according to any one of (1) to (5) above. (7) The drawing control unit Controls the drawing processing unit to draw the associated virtual object and a non-associated virtual object, which is a virtual object independent of the image recognition processing of the real object, on different drawing planes in a plurality of drawing planes The information processing apparatus according to any one of (1) to (6) above. (8) When the number of the associated virtual objects is equal to or more than the number of the plurality of drawing planes, where the number of the plurality of drawing planes is n (n is a natural number), The drawing control unit Selects n - 1 of the associated virtual objects, exclusively draws the selected associated virtual objects on at least one drawing plane, and controls the drawing processing unit to draw the unselected associated virtual objects and the non-associated virtual objects on the remaining one drawing plane The information processing apparatus according to (7) above. (9) The drawing control unit Performs the selection using a selection criterion in which the greater the movement amount of the real object, the higher the possibility of selection The information processing apparatus according to (8) above. (10) The drawing control unit Performs the selection using a selection criterion in which the smaller the area of the real object, the higher the possibility of selection The information processing apparatus according to (9) above. (11) The drawing control unit makes the selection using a selection criterion such that the possibility of selection increases as the distance between the user's gaze point and the real object becomes shorter The information processing apparatus according to (9) or (10) above (12) The drawing control unit controls the drawing processing unit so as to reduce the update frequency of the drawing plane for drawing an unrelated virtual object independent of the image recognition processing of the real object among a plurality of drawing planes, compared to the update frequency of the drawing plane for drawing the related virtual object The information processing apparatus according to any one of (1) to (11) above (13) The drawing control unit when the related virtual object is a related virtual object that performs animation, controls the drawing processing unit so as to reduce the drawing update frequency of the related virtual object compared to the case where the related virtual object does not perform animation The information processing apparatus according to any one of (1) to (12) above (14) The drawing control unit when drawing processing is performed for a plurality of drawing planes, controls the drawing processing unit to use at least one drawing plane that is smaller in size than other drawing planes The information processing apparatus according to any one of (1) to (13) above (15) The correction control unit performs the correction for a shielding virtual object, which is a virtual object that shields a part of the user's body overlapping the virtual object when viewed from the user's viewpoint position The information processing apparatus according to any one of (1) to (14) above (16) The shielding virtual object is a virtual object imitating the user's hand The information processing apparatus according to (15) above (17) The drawing control unit Control the drawing processing unit so as to exclusively use at least one of the plurality of drawing planes that can be used by the drawing processing unit as the virtual object for shielding. The information processing apparatus according to (15) or (16) above. (18) Before completion of the first drawing process, based on the result of the first recognition process, generate a light source viewpoint image which is an image of the related virtual object as seen from the position of a virtual light source that illuminates the related virtual object. Perform control so that the generated light source viewpoint image is corrected based on the result of the second recognition process before completion of the second drawing process. Further include a virtual shadow image generation unit that generates a virtual shadow image which is an image of a virtual shadow for the virtual related object based on the corrected light source viewpoint image. The information processing apparatus according to any one of (1) to (17) above. (19) The virtual shadow image generation unit Before completion of the first drawing process, based on the result of the first recognition process, calculate, for each point in the three-dimensional space projected onto each pixel of the drawing image by the drawing processing unit, the distance from each point to the virtual light source as the drawing-side light source distance respectively. Generate, as the light source viewpoint image, a depth image as a shadow map by the shadow map method. As correction of the light source viewpoint image, perform a process of changing the position or size of the image area of the related virtual object in the shadow map based on the result of the second recognition process. Generate the virtual shadow image based on the corrected shadow map and the drawing-side light source distance. The information processing apparatus according to (18) above. (20) Based on a captured image including a real object, perform a first recognition process regarding the position and orientation of the real object at a first time point. Control the drawing processing unit to perform a first drawing process for a related virtual object associated with the real object based on the first recognition process. At a second time point after the first time point, perform a second recognition process regarding the position and orientation of the real object based on the captured image including the real object, control the drawing processing unit to perform a second drawing process on the related virtual object associated with the real object based on the second recognition process, before the completion of the second drawing process, correct the first image of the related virtual object obtained at the completion of the first drawing process based on the result of the second recognition process Control method. (21) A storage medium storing a program readable by a computer device, perform a first recognition process regarding the position and orientation of the real object at a first time point based on a captured image including the real object, control the drawing processing unit to perform a first drawing process on the related virtual object associated with the real object based on the first recognition process, at a second time point after the first time point, perform a second recognition process regarding the position and orientation of the real object based on the captured image including the real object, control the drawing processing unit to perform a second drawing process on the related virtual object associated with the real object based on the second recognition process, a program that causes a computer device to execute a process of correcting the first image of the related virtual object obtained at the completion of the first drawing process based on the result of the second recognition process before the completion of the second drawing process Storage medium.
Explanation of Signs
[0225] 1, 1A Information processing device 10 Output unit 11 Imaging unit 11a First imaging unit 11b Second imaging unit 12 Operation unit 13 Sensor unit 14, 14A CPU 15 ROM 16 RAM 17 GPU 18 Image memory 18a Buffer 19, 19A Display controller 19a, 19aA Image correction processing unit 20 Recording and playback control unit 21 Communication unit 22 Bus 100a, 100b Lenses 101 Holding unit F1 Image recognition processing unit F2 Drawing control unit F3 Image correction control unit 50 AR system Ro (Ro1, Ro2, Ro3) Real object Vo (Vo2, Vo3) Virtual object Ls Virtual light source Vs Virtual shadow Pr Viewpoint (drawing viewpoint) Pcr Drawn image g1, g2 Pixels Sm Light source viewpoint image
Claims
1. An image recognition processing unit that performs first recognition processing related to the position and orientation of the real object at a first time point and second recognition processing related to the position and orientation of the real object at a second time point after the first time point, based on a captured image including the real object; A drawing control unit that controls a drawing processing unit to perform a first drawing process of drawing the related virtual object according to the position and orientation corresponding to the position and orientation of the real object recognized in the first recognition process, and a second drawing process of drawing the related virtual object according to the position and orientation corresponding to the position and orientation of the real object recognized in the second recognition process, as drawing processing for the related virtual object displayed in association with the real object; A correction control unit that corrects a virtual object image, which is an image of the related virtual object obtained by completion of the first drawing process, based on the result of the second recognition process before the second drawing process is completed while the optical image or image of the real object is visible to the user. The information processing apparatus includes: The correction control unit: As the correction, performs correction to make the position of the related virtual object follow the movement of the real object recognized by the second recognition process, and correction to make the orientation of the related virtual object follow the change in the orientation of the real object recognized by the second recognition process An information processing apparatus.
2. The correction control unit: Based on the information on the position of the real object in the in-plane direction of up, down, left, and right recognized by the image recognition processing unit, performs the correction to change the position of the related virtual object in the in-plane direction of up, down, left, and right in the virtual object image to follow the position of the real object in the in-plane direction of up, down, left, and right recognized. The information processing apparatus according to claim 1. The information processing apparatus according to claim 1.
3. The correction control unit: Based on the information on the position of the real object in the depth direction recognized by the image recognition processing unit, performs the correction to change the size of the related virtual object in the virtual object image to a size corresponding to the position of the real object in the depth direction recognized. The information processing apparatus according to claim 1. The information processing apparatus according to claim 1.
4. The correction control unit: Based on the information on the viewpoint position or line-of-sight direction of the user detected by a sensor, performs the correction to change the position or orientation of the related virtual object to the position or orientation corresponding to the detected viewpoint position or line-of-sight direction of the user. The information processing apparatus according to claim 1. The information processing apparatus according to claim 1.
5. The correction control unit: When there are a plurality of associated virtual objects that are each associated with a different real object and displayed, one or more associated virtual objects to be the subject of the correction are selected from among the plurality of associated virtual objects according to a priority based on the size of the real object that is the subject of the association. The information processing apparatus according to claim 1.
6. The processing cycle of the correction is set to be shorter than the processing cycle of the image recognition processing unit. The information processing apparatus according to claim 1.
7. The drawing control unit controls the drawing processing unit to draw the associated virtual object and a non-associated virtual object, which is a virtual object not associated with and displayed with the real object, on different drawing planes among a plurality of drawing planes. The information processing apparatus according to claim 1.
8. When the number of the associated virtual objects is equal to or more than the number of the plurality of drawing planes, with the number of the plurality of drawing planes being n (n is a natural number), The drawing control unit selects n - 1 of the associated virtual objects, exclusively draws the selected associated virtual objects on at least one drawing plane, and controls the drawing processing unit to draw the unselected associated virtual objects and the non-associated virtual objects on the remaining one drawing plane. The information processing apparatus according to claim 7.
9. The drawing control unit performs the selection using a selection criterion such that the more the movement amount of the real object that is the subject of the association of an associated virtual object is, the more likely the associated virtual object is to be selected. The information processing apparatus according to claim 8.
10. The drawing control unit performs the selection using a selection criterion such that the smaller the area of the real object that is the subject of the association of an associated virtual object is, the more likely the associated virtual object is to be selected. The information processing apparatus according to claim 9.
11. The drawing control unit performs the selection using a selection criterion such that the shorter the distance between the user's gaze point and the real object that is the subject of the association of an associated virtual object is, the more likely the associated virtual object is to be selected. The information processing apparatus according to claim 9.
12. The drawing control unit controls the drawing processing unit to reduce the update frequency of the drawing plane for drawing a non-associated virtual object, which is a virtual object not associated with and displayed with the real object, among the plurality of drawing planes, compared to the update frequency of the drawing plane for drawing the associated virtual objects. The information processing apparatus according to claim 1.
13. The drawing control unit When the related virtual object is a related virtual object that performs animation, the drawing processing unit is controlled to reduce the drawing update frequency of the related virtual object compared to the case where the related virtual object does not perform animation. The information processing apparatus according to claim 1.
14. The drawing control unit When drawing processing is performed on a plurality of drawing planes, the drawing processing unit is controlled to use, as at least one of the plurality of drawing planes, a drawing plane having a smaller size than other drawing planes. The information processing apparatus according to claim 1.
15. The correction control unit When a part of the user's body is located in front of the virtual object as viewed from the user's viewpoint position, the correction is performed using, as the related virtual object, a shielding virtual object that shields a part of the virtual object that overlaps with the part of the body. The information processing apparatus according to claim 1.
16. The shielding virtual object is a virtual object imitating the user's hand. The information processing apparatus according to claim 15.
17. The drawing control unit The drawing processing unit is controlled to use at least one of the plurality of drawing planes that can be used by the drawing processing unit for drawing only the shielding virtual object. The information processing apparatus according to claim 15.
18. Before completion of the first drawing process, based on the result of the first recognition process, a light source viewpoint image, which is an image of the related virtual object viewed from the position of a virtual light source that illuminates the related virtual object, is generated. Control is performed so that the generated light source viewpoint image is corrected based on the result of the second recognition process before completion of the second drawing process. The apparatus further includes a virtual shadow image generation unit that generates a virtual shadow image, which is an image of a virtual shadow for the related virtual object, by the shadow map method based on the corrected light source viewpoint image. The virtual shadow image generation unit As correction of the light source viewpoint image, correction is performed to make the position of the related virtual object in the light source viewpoint image follow the movement of the real object recognized by the second recognition process. The information processing apparatus according to claim 1.
19. Based on a captured image including a real object, a first recognition process regarding the position and orientation of the real object at a first point in time is performed. As drawing processing for the related virtual object that is displayed in association with the real object, control the drawing processing unit to perform first drawing processing for drawing the related virtual object in a position and orientation corresponding to the position and orientation of the real object recognized in the first recognition processing. At a second time point after the first time point, perform second recognition processing regarding the position and orientation of the real object based on the captured image including the real object. As drawing processing for the related virtual object that is displayed in association with the real object, control the drawing processing unit to perform second drawing processing for drawing the related virtual object in a position and orientation corresponding to the position and orientation of the real object recognized in the second recognition processing. In a state where the optical image or image of the real object is visible to the user, before the second drawing processing is completed, correct the first image of the related virtual object obtained by the completion of the first drawing processing based on the result of the second recognition processing. As the correction, perform correction to cause the position of the related virtual object to follow the movement of the real object recognized by the second recognition processing, and correction to cause the orientation of the related virtual object to follow the change in the orientation of the real object recognized by the second recognition processing. An information processing method.
Citation Information
Patent Citations
Image generation apparatus and image generation method
JP2015095045A
Head mounted display and method for controlling the same
JP2018106157A
Information processing device, information processing method, computer program, and image processing system
WO2016002318A1
Image processing device and image generation method
WO2017086263A1
Information processing device, information processing method, and program
WO2017183346A1